Data query method, electronic equipment and device

By pre-forming the first index and the second index in the database, the problem of low query efficiency during fuzzy queries in databases is solved, and an efficient data query method in which a target string of any length can be queried based on the index is realized.

CN120067407APending Publication Date: 2025-05-30NEW H3C TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510221246.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When performing fuzzy queries in the database, a single query takes a long time, is low in query efficiency, and consumes high CPU and memory resources. Especially when the length of the target string to be queried is less than the preset length, it is impossible to effectively use the index to query.

Method used

A data query method is provided, by pre-forming the first index and the second index, respectively, for the case where the length of the target character string is less than and not less than the preset length. The first index is generated based on the split record data and includes a plurality of character groups, each character group containing continuous characters with a length smaller than a preset length. The second index is generated based on the complete record data.

Benefits of technology

It improves the efficiency of data query, reduces the waste of CPU and memory resources of the database, and ensures that target strings of any length can be queried based on the index.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067407A_ABST
    Figure CN120067407A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data query method, electronic equipment and a data query device, and relates to the technical field of data processing. The method comprises the following steps: judging whether the length of a to-be-queried target character string is smaller than a preset length or not; if it is judged that the length of the to-be-queried target character string is smaller than the preset length, a first target index matched with the target character string is queried in first indexes generated in advance, and record data corresponding to the first target index stored in a database is determined, the first index corresponding to each stored record data is generated based on character groups contained in the record data, and each character group contains continuous characters with the length smaller than the preset length in the record data. By adopting the technical scheme provided by the embodiment of the invention, the data query efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a data query method, an electronic device, and a device. Background Art

[0002] In many industries, a large amount of data is generated every moment, thus forming a storage requirement for massive data, and these massive data are usually stored based on a database. With the development of information systems in various industries, there is also a query requirement for these massive data.

[0003] When querying a form in a database, fuzzy query is usually used. When querying and filtering, it is necessary to traverse and scan all the data in the table. When the amount of data stored in the database is large, the time consumed for a single query is long, the query efficiency is low, and it requires consuming a high amount of Central Processing Unit (CPU) resources and memory resources of the database. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a data query method, an electronic device, and a device to improve the efficiency of data query. The specific technical solutions are as follows:

[0005] In a first aspect, the embodiments of this application provide a data query method, and the method includes:

[0006] Judge whether the length of the target string to be queried is less than a preset length;

[0007] If it is determined that the length of the target string to be queried is less than the preset length, then in the pre-generated first index, query the first target index that matches the target string, and determine the record data corresponding to the first target index stored in the database, where the first index corresponding to each record data is generated based on the split record data, and the split record data contains multiple character groups, and each character group contains consecutive characters in the record data and the length is less than the preset length.

[0008] In an embodiment of this application, the method further includes:

[0009] If it is determined that the length of the target string to be queried is not less than the preset length, then in the pre-generated second index, query the second target index that matches the target string, and determine the record data corresponding to the second target index stored in the database, where the second index corresponding to each record data is generated based on the record data.

[0010] In an embodiment of this application, the method further includes:

[0011] If the recorded data stored in the database changes, then for the changed recorded data, a first index and a second index corresponding to the recorded data are generated.

[0012] In one embodiment of the present application, the method further includes:

[0013] In the process of generating the first index, specified characters in the character group are removed and / or duplicate character groups are removed.

[0014] In a second aspect, an embodiment of the present application provides an electronic device, which includes:

[0015] A processor;

[0016] A transceiver;

[0017] A machine-readable storage medium storing machine-executable instructions that can be executed by the processor, and the machine-executable instructions cause the processor to execute the following steps:

[0018] Determine whether the length of the target string to be queried is less than a preset length;

[0019] If it is determined that the length of the target string to be queried is less than the preset length, then in the pre-generated first index, query a first target index that matches the target string, and determine the recorded data corresponding to the first target index stored in the database, where the first index corresponding to each stored recorded data is generated based on the split recorded data, and the split recorded data contains multiple character groups, and each character group contains consecutive characters in the recorded data whose length is less than the preset length.

[0020] In one embodiment of the present application, the machine-executable instructions further cause the processor to execute the following steps:

[0021] If it is determined that the length of the target string to be queried is not less than the preset length, then in the pre-generated second index, query a second target index that matches the target string, and determine the recorded data corresponding to the second target index stored in the database, where the second index corresponding to each recorded data is generated based on the recorded data.

[0022] In one embodiment of the present application, for each recorded data, the machine-executable instructions further cause the processor to generate the first index corresponding to the recorded data based on the following method:

[0023] Extract character groups from the recorded data;

[0024] Construct a first index representing the correspondence between each character group and the position of the recorded data in the database.

[0025] In one embodiment of the present application, the machine - executable instructions further cause the processor to perform the following steps:

[0026] If the recorded data stored in the database changes, then for the changed recorded data, generate a first index and a second index corresponding to the recorded data.

[0027] In one embodiment of the present application, the machine - executable instructions further cause the processor to perform the following steps:

[0028] In the process of generating the first index, remove the specified characters in the character group and / or remove duplicate character groups.

[0029] In a third aspect, an embodiment of the present application provides a data query device, and the device includes

[0030] A length judgment module, configured to judge whether the length of a target string to be queried is less than a preset length;

[0031] A first query module, configured to, when the length judgment module determines that the length of the target string to be queried is less than the preset length, query a first target index matching the target string in a pre - generated first index, and determine the recorded data corresponding to the first target index stored in the database, where the first index corresponding to each recorded data is generated based on the split recorded data, the split recorded data contains multiple character groups, and each character group contains consecutive characters in the recorded data and the length of each character group is less than the preset length.

[0032] In one embodiment of the present application, the device further includes:

[0033] A second query module, configured to, when the length judgment module determines that the length of the target string to be queried is not less than the preset length, query a second target index matching the target string in a pre - generated second index, and determine the recorded data corresponding to the second target index stored in the database, where the second index corresponding to each recorded data is generated based on the recorded data.

[0034] In one embodiment of the present application, for each recorded data, the first index corresponding to the recorded data is generated based on a first index module:

[0035] The first index generation module is configured to extract character groups from the recorded data;

[0036] Construct a first index representing the correspondence between each character group and the position of the recorded data in the database.

[0037] In one embodiment of the present application, the device further includes:

[0038] An index update module, configured to, if the record data stored in the database changes, generate a first index and a second index corresponding to the changed record data for the changed record data.

[0039] In one embodiment of the present application, the apparatus further includes:

[0040] A character removal module, configured to remove specified characters in a character group and / or remove duplicate character groups during the process of generating the first index.

[0041] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the data query method described in any one of the above is implemented.

[0042] An embodiment of the present application further provides a computer program product including instructions, which when running on a computer, causes the computer to execute the data query method described in any one of the above.

[0043] Advantages of the embodiments of the present application:

[0044] In the technical solution provided by the embodiment of the present application, a first index is pre-constructed for each record data in the database table stored in the database. When performing data query, the first index is used to avoid traversing and scanning all the data in the table, improving the query efficiency and at the same time reducing the waste of CPU resources and memory resources of the database.

[0045] In addition, in the technical solution provided by the embodiment of the present application, the first index is pre-generated based on the split record data, and the split record data includes multiple character groups, and each character group includes consecutive characters in the record data with a length less than the preset length. When the length of the target string to be queried is less than the preset length, the first index can be used to query the record data that matches the target string. That is, by adopting the technical solution provided by the embodiment of the present application, the length of the target string used for data query is not restricted during data query. When the length of the target string is less than the preset length, the index can also be used for data query, ensuring that target strings of any length can be based on the index for data query, further improving the query efficiency of data query.

[0046] Of course, it is not necessary for any product or method implementing the present application to achieve all the above advantages at the same time. Description of the Drawings

[0047] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other embodiments can also be obtained based on these drawings.

[0048] Figure 1 Schematic flowchart of the first data query method provided by an embodiment of the present application;

[0049] Figure 2 Schematic flowchart of constructing a first index for record data provided by an embodiment of the present application;

[0050] Figure 3 Schematic flowchart of constructing a second index for record data provided by an embodiment of the present application;

[0051] Figure 4 Schematic flowchart of the second data query method provided by an embodiment of the present application;

[0052] Figure 5 Schematic flowchart of the third data query method provided by an embodiment of the present application;

[0053] Figure 6 Schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0054] Figure 7 Schematic structural diagram of a data query device provided by an embodiment of the present application. Detailed implementation manners

[0055] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.

[0056] In many industries, a large amount of data is generated every moment, thus forming a storage requirement for massive data. These massive data are usually stored based on a database. With the development of information systems in various industries, there is also a query requirement for these massive data.

[0057] When querying a form in a database, fuzzy query is usually used. When querying and filtering, it is necessary to traverse and scan all the data in the table. When the amount of data stored in the database is large, a single query takes a long time, the query efficiency is low, and it requires consuming a high amount of CPU resources and memory resources of the database.

[0058] In the related art, although an index is built for the record data stored in the database, there is a limit on the length of the string to be queried during data query. When the length of the string is less than the preset length, data query cannot be performed based on the index.

[0059] Taking the PostgreSQL (abbreviated as PG) database as an example, PG is an open-source relational database management system originally developed by the Computer Science Department of the University of California, Berkeley. Both PG and MySQL are popular open-source relational database management systems. In the PG database, a Generalized Inverted Index (GIN) can be built for the record data in the database table based on the pg_trgm plugin. In this way, the index can be used to improve the query efficiency during data query. However, when the length of the query string is less than 3, the index cannot be used, which still results in low data query efficiency.

[0060] For easy understanding, the reason why the gin index cannot be used when the length of the query string is less than 3 when building the gin index based on the pg_trgm plugin in the PG database will be briefly described below.

[0061] The pg_trgm plugin provides a data type called trgm for storing character triples and also provides a set of functions for calculating the similarity between two strings. By converting the string to the trgm type, the gin index can be used to accelerate data query. Since the pg_trgm plugin converts the record data in the database to the trgm type and builds the gin index, when performing forward and backward fuzzy queries, when the length of the input query string is less than 3, the corresponding character triples cannot be matched, and thus the index cannot be used for query. It is necessary to traverse and scan all the data in the table to complete the data query, resulting in low data query efficiency.

[0062] To solve the above technical problems, the embodiments of the present application provide a data query method, an electronic device, and a device.

[0063] See Figure 1 , which is a schematic flowchart of the first data query method provided by the embodiments of the present application. This method can be applied to electronic devices such as mobile terminals and servers. For easy description, the electronic device is used as the execution subject below, which is not restrictive. This method includes step S101-step S102.

[0064] S101, determine whether the length of the target string to be queried is less than the preset length.

[0065] The above-mentioned target string to be queried can be input by the user based on an electronic device. Specifically, the database system where the database is located can provide an interactive interface to the user based on the electronic device. The interactive interface provides a control for the user to input the target string, and the user can input the target string through this control. The electronic device can obtain the target string input by the user.

[0066] The above-mentioned target string to be queried can also be pre-set by the user, and the electronic device can directly obtain the target string.

[0067] Or the above-mentioned target string to be queried can also be read by the electronic device from a preset document.

[0068] The value of the above-mentioned preset length generally needs to be greater than or equal to 3, and the specific value can be set based on actual requirements.

[0069] In the embodiment of the present application, after the electronic device obtains the target string to be queried, it determines whether the length of the target string to be queried is less than the preset length. If it is determined that the length of the target string to be queried is less than the preset length, then step S102 is executed.

[0070] S102, in the first index pre-generated, query the first target index that matches the target string, and determine the record data corresponding to the first target index stored in the database.

[0071] Among them, the first index corresponding to each stored record data is generated based on the character groups included in the record data. Each character group includes consecutive characters in the record data with a length less than the preset length. For example, when the preset length is 3, each character group can include a single character in the record data, or two consecutive characters. Taking the record data as "abc" as an example, the character groups included in this record data can be "a", "b", "c", "ab", and "bc".

[0072] The first index can be an inverted index or a forward index. The present application does not limit the specific type of the first index. When constructing the first index for the character groups included in the record data, it can be constructed based on the plug-ins in the existing database.

[0073] It should be noted that in the related art, when constructing an index for each record data in the database, the complete record data is used as the attribute value, or multiple adjacent strings with a preset length extracted from the record data are used as the attribute value. This results in that when the length of the target string to be queried is less than the preset length, the corresponding attribute value cannot be matched, that is, data query cannot be performed based on the constructed index. Therefore, in the embodiments of the present application, a first index is constructed for each record data in the database to ensure that target strings of any length can use the index for data query, so as to improve the data query efficiency.

[0074] Specifically, in the related art, when generating an index for the record data stored in the database, the complete record data is used as the attribute value, or multiple adjacent strings with a preset length extracted from the record data are used as the attribute value. This results in that when the length of the target string to be queried is less than the preset length, the corresponding attribute value cannot be matched, that is, data query cannot be performed based on the index. Therefore, in the embodiments of the present application, for each record data stored in the database, the consecutive characters with a length less than the preset length in the record data are used as a character group, and a first index is constructed for the character group. At this time, the attribute of the attribute value in the first index can be a text attribute. In addition, due to some special mechanisms for index construction, no index is constructed for single-character or double-character strings. Therefore, in the embodiments of the present application, the character group can be stored in some data types, such as array, JSONB data, etc., and a first index is created for each character group. At this time, the attribute of the attribute value in the first index is an array, JSONB, etc. attribute. In this way, regardless of the length of the target string to be queried, data query can be performed based on the index, improving the efficiency of data query.

[0075] In the embodiments of the present application, after the electronic device receives the target string and determines that the length of the target string is less than the preset length, it can further determine whether the attribute of the target string is consistent with the attribute of the attribute value in the first index pre-constructed for the record data in the database. If they are consistent, based on the first index, the first target index matching the target string is queried, and then based on the first target index, the record data corresponding to the first target index stored in the database is determined. If they are inconsistent, the attribute of the target string can be converted into the attribute of the attribute value in the first index, and then data query is performed based on the first index. In addition, the address information corresponding to the record data, that is, the position of the record data in the database, can also be determined based on the first target index. When performing data query, it can be a data query for a specified column or multiple columns of data, or a data query for the entire table data.

[0076] In the technical solution provided by the embodiment of the present application, a first index is pre-constructed for each record data in the database table stored in the database. When performing data query, the first index is used to avoid traversing and scanning all the data in the table, which improves the query efficiency and reduces the waste of CPU resources and memory resources of the database at the same time.

[0077] In addition, in the technical solution provided by the embodiment of the present application, the first index is pre-generated based on the split record data, and the split record data contains multiple character groups. Each character group contains consecutive characters in the record data with a length less than the preset length. When the length of the target string to be queried is less than the preset length, the first index can be used to query the record data that matches the target string. That is, by adopting the technical solution provided by the embodiment of the present application, the length of the target string used for data query is not restricted during data query. Whether the length of the target string is long or short, the index can be used for data query, ensuring that target strings of any length can be queried based on the index, further improving the query efficiency of data query.

[0078] In an embodiment of the present application, when the electronic device obtains the target string to be queried, it determines whether the length of the target string to be queried is less than the preset length. If the determination result is no, step S103 can be executed.

[0079] S103, in the pre-generated second index, query the second target index that matches the target string, and determine the record data corresponding to the second target index stored in the database.

[0080] Among them, the second index corresponding to each record data is generated based on the record data.

[0081] The above-mentioned second index is pre-generated, and a second index will be pre-constructed for each record data in the database. The construction method of the second index can be based on the plug-ins in the existing database. The type of the above-mentioned second index can be an inverted index. An inverted index is pre-constructed for the record data stored in the database, and using the inverted index for data query can further improve the data query efficiency. When performing fuzzy query on a large amount of string text data, by utilizing the database characteristics of the inverted index, the query result can be quickly obtained, avoiding time-consuming operations such as full table scan, and improving the query efficiency.

[0082] Specifically, when the second index is an inverted index, the second index may include an attribute value and the address information of each record data having the attribute value, so as to determine the address of the record data according to the attribute value. The attribute value may specifically be a text attribute. In an embodiment of the present application, the attribute value may be the complete record data, and the second index may record the correspondence between the record data and the address information of the record data. The address information may specifically be the serial number of the record data in the database table. For example, if the record data is "abcd" and the serial number of the record data in the database table is 3, the form of the second index corresponding to the record index may be {"abcd", 3}. In another embodiment of the present application, the attribute value may be multiple adjacent strings of a preset length extracted from the record data, and the second index may record the correspondence between each string and the address information of the record data containing the string. For example, if the record data is "abcd", the serial number of the record data in the database table is 3, and the preset length is 3, the form of the second index corresponding to the record data may be {"abc", 3} and {"bcd", 3}.

[0083] In some other embodiments of the present application, the type of the second index may also be a forward index, and the present application does not specifically limit the type of the second index.

[0084] In the embodiment of the present application, after receiving the target string, the electronic device determines that the length of the target string is not less than the preset length, and then based on the second index, queries the second target index that matches the target string, and further based on the second target index, determines the record data corresponding to the second target index stored in the database. In addition, the address information corresponding to the record data, that is, the position of the record data in the database, can also be determined based on the second target index. When performing data query, it may be a data query on a specified column or multiple columns of data, or a data query on the entire table data.

[0085] In an embodiment of the present application, the electronic device may directly use the target string for matching, query the second target index that matches the target string from the second index, and determine the record data corresponding to the second target index stored in the database.

[0086] For example, a database table included in the database includes three fields: account number, name, and email, and each field includes multiple record data. Refer to Table 1, which is a schematic diagram of a database table provided by the embodiment of the present application.

[0087] Table 1

[0088] Account Name Email admin Zhang San admin@abc.com abdemocd Li Si abdemocd@abc.com …… …… ……

[0089] The user query request obtained by the electronic device is the record data in which the value of the field "account number" contains "demo", that is, the target string is "demo". Then, based on the second index pre-constructed for each record data in the "account number" field of the database table, the second target index is determined. Furthermore, the record data corresponding to the second target index can be determined: "abdemocd", and the address information of "abdemocd" in the database table.

[0090] In an embodiment of the present application, when the attribute value is multiple strings with a preset length extracted from the record data, and the second index records the correspondence between each string and the address information of the record data containing the string, after obtaining the target string, the electronic device can split the target string into consecutive strings with a preset length, and then use these strings for matching to query the second target index that matches all the strings obtained by splitting the target string from the second index, and determine the record data corresponding to the second target index stored in the database.

[0091] Continuing to refer to the example in Table 1, the user query request obtained by the electronic device is the record data in the database table whose value contains "demo", that is, the target string is "demo". If the preset length is 3, the electronic device splits the target string "demo" to obtain the strings to be queried as [dem, emo]. Then, based on the second index pre-constructed for each record data in the database, the second target index is determined. Furthermore, the record data corresponding to the second target index can be determined, that is, the record data in the database table that contains both "dem" and "emo" is determined: "abdemocd" and "abdemocd@abc.com", and the address information of "abdemocd" and "abdemocd@abc.com" in the database table.

[0092] As can be seen from the above embodiments, in the embodiments of the present application, in addition to pre-building a first index for each record data in the database table stored in the database, a second index is also built for each record data. When performing data query, the first index or the second index is used for data query, without traversing and scanning all the data in the table, which improves the query efficiency and reduces the waste of CPU resources and memory resources of the database. In addition, in the technical solution provided by the embodiments of the present application, since the first index is pre-generated based on the split record data, and the split record data contains multiple character groups, and each character group contains consecutive characters in the record data with a length less than the preset length. When the length of the target string to be queried is less than the preset length, the first index can be used to query the record data that matches the target string. The second index is pre-generated based on the record data stored in the database. When the length of the target string to be queried is not less than the preset length, the second index can be directly used to query the record data that matches the target string. That is, by adopting the technical solution provided by the embodiments of the present application, the length of the target string used for data query is not limited during data query. Whether the length of the target string is long or short, the index can be used for data query, ensuring that the target string of any length can be based on the index for data query, further improving the query efficiency of data query.

[0093] The following briefly describes the method of building the first index for the record data in the database.

[0094] See Figure 2 , which is a schematic flow chart of building a first index for record data provided by an embodiment of the present application. The method includes steps S201 - S202.

[0095] S201, extract character groups from the record data.

[0096] In the embodiments of the present application, for each record data in the database, the electronic device can extract multiple character groups from the record data. Specifically, the electronic device can continue to split the record data to obtain multiple character groups, and each character group contains consecutive characters in the record data with a length less than the preset length.

[0097] For example, when the record data is "abc" and the preset length is 3, the electronic device can extract the record data to obtain character groups: "a", "b", "c", "ab", and "bc".

[0098] S202, build a first index representing the correspondence between each character group and the position of the record data in the database.

[0099] In the embodiments of the present application, after the electronic device extracts character groups from the recorded data, it can retrieve the position of the recorded data in the database, and then construct a first index for each character group based on the correspondence between each character group and the position of the recorded data in the database.

[0100] Continuing with the above example, the address information can specifically be the serial number of the recorded data in the database table. When the electronic device retrieves that the serial number of the recorded data "abc" in the database table is 3, the corresponding first index form of this recorded data can be {"a", 3}, {"b", 3}, {"c", 3}, {"ab", 3}, and {"bc", 3}.

[0101] As can be seen from the above embodiments, in the embodiments of the present application, a first index is pre-created for the recorded data in the database. In this way, data queries can be performed based on the index, improving the efficiency of data queries. In addition, the first index is constructed based on the character groups split from the recorded data. In this way, when the length of the target string to be queried is less than the preset length, data queries can also be performed based on the index, further improving the data query efficiency.

[0102] See Figure 3 , which is a schematic flowchart of a process for constructing a second index for recorded data provided by the embodiments of the present application. The method includes steps S301 - S303.

[0103] S301, splice a specified character of the first length at the head end of the recorded data, and splice a specified character of the second length at the tail end of the recorded data to obtain the target recorded data.

[0104] Wherein, the sum of the first length and the second length is the preset length. The above-mentioned specified character can be pre-set and has no real meaning. For example, the specified character can be a space.

[0105] In an embodiment of the present application, the second index can be an index constructed for each string by splitting the recorded data into strings of a preset length. Since fuzzy queries may be performed before and after during data queries, in order to query all recorded data associated with the target string, when the electronic device splits the recorded data into strings of a preset length, it can splice a specified character of the first length at the head end of the recorded data, and splice a specified character of the second length at the tail end of the recorded data to obtain the target recorded data.

[0106] For example, if the preset length is 3, the first length can be 2, the second length is 1, the specified character is a space, and the recorded data is "abcd", then splice a specified character of the first length at the head end of the recorded data, and splice a specified character of the second length at the tail end of the recorded data, and the obtained target recorded data is "..abcd.", where "." represents a space.

[0107] S302. Extract a string composed of a preset number of adjacent characters from the target record data.

[0108] In an embodiment of the present application, after obtaining the target record data, the electronic device may extract a string composed of a preset number of adjacent characters from the target record data.

[0109] Continuing with the above example, the target record data obtained by the electronic device is "..abcd.", and the strings of adjacent length 3 extracted from the target record data are {"..a", ".ab", "abc", "bcd", "cd."}, where "." represents a space.

[0110] S303. Construct a second index representing the correspondence between the extracted string and the position of the record data in the database.

[0111] In an embodiment of the present application, after the electronic device extracts multiple strings, it retrieves the position of the record data in the database. Then, based on the correspondence between the extracted strings and the position of the record data in the database, a second index is constructed for these multiple strings. Since these multiple strings are extracted from the same record data, the address information in the second index constructed for these strings is the same.

[0112] Specifically, taking the PG database as an example, the PG database provides the pg_trgm plugin. The pg_trgm plugin introduces the Trigram concept. A Trigram is a string composed of three consecutive characters taken from a string. In the pg_trgm plugin, the length of the Trigram extracted from the record data is 3. For a Trigram with a length less than 3, it will be filled with space prefixes and suffixes to obtain the final Trigram, and generally, it includes two space prefixes and one space suffix. The pg_trgm plugin provides the GIN index operator class, and the gin_trgm_ops index operator in the pg_trgm plugin can be used to create a GIN index. The working principle of the gin_trgm_ops index operator is: convert the text data into Trigrams and use the GIN index structure to save the Trigrams, that is, construct a second index for the strings extracted from the record data.

[0113] As can be seen from the above embodiments, in the embodiments of the present application, a second index is pre-created for the record data in the database. In this way, data queries can be based on the index, improving the efficiency of data queries. In addition, a string composed of a preset number of adjacent characters is extracted from the target record data, and a second index is constructed for each of the extracted strings. In this way, whether the target string to be queried covers all index strings or only covers a subset of a preset length, data can be quickly queried, further improving the data query efficiency.

[0114] See Figure 4 , which is a schematic flowchart of the second data query method provided by the embodiments of the present application. This method includes steps S401 - S404. Compared with the foregoing embodiments, step S401 is the same as step S101 above, and step S404 is the same as step S103 above, so they will not be elaborated. Steps S402 - S403 are an implementable manner of step S102.

[0115] S402, convert the attribute of the target string into the attribute of a character group.

[0116] In an embodiment of the present application, when an electronic device performs a data query based on the first index, since in the first index constructed for the character group, the attribute of the character group may be of array type or JSONB type, while the attribute of the target string is generally of text type, and the two attributes may be inconsistent, so the attribute of the target string is converted into the attribute of the character group to ensure that the index can be used normally during data query. For converting the attribute of the target string into the attribute of the character group, it can be converted based on some type conversion functions provided by the database itself.

[0117] S403, in the pre-generated first index, query the first target index that matches the target string after attribute conversion, and determine the record data corresponding to the first target index stored in the database.

[0118] In an embodiment of the present application, after the electronic device converts the attribute of the target string into the attribute of the character group, it queries the first target index that matches the target character group after attribute conversion in the pre-generated first index. In this way, the attribute of the target string after attribute conversion is the same as the attribute of the character group in the first index, and thus the corresponding target first index can be determined from the first index. Furthermore, based on the first target index, the address information corresponding to the record data, that is, the location of the record data in the database, can also be determined.

[0119] As can be seen from the above embodiments, in the embodiments of the present application, when querying data using the first index, the attributes of the target string are converted into the attributes of the character group. In this way, the attributes of the string after attribute conversion are consistent with the attributes of the character group in the first index, so as to ensure that when the target string to be queried is less than the preset length during data query, the first index can be used normally.

[0120] See Figure 5 , which is a schematic flowchart of the third data query method provided by the embodiments of the present application. This method includes steps S501 - step S504. Compared with the foregoing embodiments, steps S501 - step S503 are the same as steps S101 - step S103, and this method further includes step S504.

[0121] S504, if the record data stored in the database changes, then for the changed record data, generate the corresponding first index and second index of the record data.

[0122] In an embodiment of the present application, since the record data stored in the database is constantly changing. For example, the record data has changed or a new record data has been added. Therefore, in order to ensure the accuracy of the index in the database, when the electronic device detects that the record data stored in the database has changed, for the changed record data, generate the corresponding first index and second index of the record data.

[0123] Specifically, the electronic device can create a trigger in the database. When the record data in the database table increases or the record data is modified, it automatically triggers a predefined action: for the changed record data, generate the corresponding first index and second index of the record data.

[0124] As can be seen from the above embodiments, in the embodiments of the present application, when it is detected that the record data stored in the database has changed, for the changed record data, generate the corresponding first index and second index of the record data, so as to be able to realize the automatic update of the index and ensure the accuracy of data query.

[0125] In an embodiment of the present application, during the process of generating the first index, the electronic device can remove the specified characters in the character group and / or remove the duplicate character groups.

[0126] The above - mentioned specified characters can be pre - set characters without real meaning. For example, the specified characters can be symbols such as spaces.

[0127] In the process of pre - generating the first index, for each record data stored in the database, consecutive characters in the record data that are less than the preset length are used as a character group, and a first index is constructed for this character group. There may be some characters without real meaning in the character group, that is, specified characters. When querying using the target string, due to the existence of these specified characters in the character group, the corresponding target index may not be matched. Therefore, in an embodiment of the present application, by removing the specified characters in the character group, while ensuring the accuracy of data query during the construction of the first index, the memory resources of the database can also be saved.

[0128] When taking consecutive characters in the record data that are less than the preset length as a character group, there may be duplicate character groups. For example, if the record data is "app", the obtained character groups are {"a", "p", "p", "ap", "pp"}, and the character group "p" is a duplicate character group. Before constructing the first index for these character groups, the duplicate character groups will be removed, and only one of the duplicate character groups needs to be retained. If there are duplicate character groups and a first index is created for each character group, it will not speed up the data query rate but will instead waste the memory resources of the database. Therefore, in an embodiment of the present application, the electronic device will remove the duplicate character groups and then construct the first index for the character groups after deduplication.

[0129] In an embodiment of the present application, during the process of generating the first index, the electronic device can both remove the specified characters in the character group and remove the duplicate character groups. In another embodiment of the present application, during the process of generating the first index, the electronic device can only remove the specified characters in the character group. In another embodiment of the present application, during the process of generating the first index, the electronic device can only remove the duplicate character groups.

[0130] As can be seen from the above embodiments, in the embodiments of the present application, during the process of generating the first index, by removing the specified characters in the character group and / or removing the duplicate character groups, while improving the accuracy of data query, the memory resources of the database are saved.

[0131] Based on the same inventive concept, the embodiments of the present application also provide an electronic device.

[0132] See Figure 6 , which is a schematic structural diagram of an electronic device provided by the embodiments of the present application. The electronic device includes:

[0133] Processor 601;

[0134] Transceiver 604;

[0135] A machine-readable storage medium 602 stores machine-executable instructions that can be executed by the processor 601. The machine-executable instructions cause the processor 601 to perform the following steps:

[0136] Determine whether the length of the target string to be queried is less than a preset length;

[0137] If it is determined that the length of the target string to be queried is less than the preset length, then in the pre-generated first index, query the first target index that matches the target string, and determine the record data corresponding to the first target index stored in the database. Each record data corresponds to a first index generated based on the split record data. After splitting, the record data contains multiple character groups, and each character group contains consecutive characters in the record data and with a length less than the preset length.

[0138] As Figure 6 shown, the network device may further include a communication bus 603. Communication between the processor 601, the machine-readable storage medium 602, and the transceiver 604 is completed through the communication bus 603. The communication bus 603 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus 603 can be divided into an address bus, a data bus, a control bus, etc.

[0139] The transceiver 604 can be a wireless communication module. Under the control of the processor 601, the transceiver 604 exchanges data with other devices.

[0140] The machine-readable storage medium 602 may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Additionally, the machine-readable storage medium 602 may also be at least one storage device located far from the aforementioned processor.

[0141] The processor 601 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0142] In the technical solution provided by the embodiments of the present application, a first index is pre-constructed for each record data in the database table stored in the database. When performing data query, the first index is used to avoid traversing and scanning all the data in the table, which improves the query efficiency and reduces the waste of CPU resources and memory resources of the database at the same time.

[0143] In addition, in the technical solution provided by the embodiments of the present application, the first index is pre-generated based on the split record data, and the split record data contains multiple character groups. Each character group contains consecutive characters in the record data with a length less than the preset length. When the length of the target string to be queried is less than the preset length, the first index can be used to query the record data that matches the target string. That is, by adopting the technical solution provided by the embodiments of the present application, the length of the target string used for data query is not limited during data query. When the length of the target string is less than the preset length, the index can also be used for data query, ensuring that target strings of any length can be queried based on the index, and further improving the query efficiency of data query.

[0144] In one embodiment of the present application, the machine-executable instructions further cause the processor to perform the following steps:

[0145] If it is determined that the length of the target string to be queried is not less than the preset length, then in the pre-generated second index, query the second target index that matches the target string, and determine the record data corresponding to the second target index stored in the database, where the second index corresponding to each record data is generated based on the record data.

[0146] In one embodiment of the present application, for each record data, the machine-executable instructions further cause the processor to generate the first index corresponding to the record data based on the following method:

[0147] Extract character groups from the record data;

[0148] Construct a first index representing the correspondence between each character group and the position of the record data in the database.

[0149] In one embodiment of the present application, the machine-executable instructions further cause the processor to perform the following steps:

[0150] If the record data stored in the database changes, then for the changed record data, generate a first index and a second index corresponding to the record data.

[0151] As can be seen from the above embodiments, in the embodiments of the present application, when it is detected that the record data stored in the database changes, then for the changed record data, generate a first index and a second index corresponding to the record data, so as to be able to realize automatic update of the index and ensure the accuracy of data query.

[0152] In one embodiment of the present application, the machine-executable instructions further cause the processor to perform the following steps:

[0153] During the process of generating the first index, remove the specified characters in the character group and / or remove duplicate character groups.

[0154] As can be seen from the above embodiments, in the embodiments of the present application, during the process of generating the first index, by removing the specified characters in the character group and / or removing duplicate character groups, the memory resources of the database are saved.

[0155] Based on the same inventive concept, the embodiments of the present application further provide a data query device.

[0156] See Figure 7 , which is a schematic structural diagram of a data query device provided by the embodiments of the present application. The device includes:

[0157] A length judgment module 701, configured to judge whether the length of the target string to be queried is less than a preset length;

[0158] A first query module 702, configured to, when the length judgment module determines that the length of the target string to be queried is less than the preset length, query a first target index matching the target string in a pre-generated first index, and determine the record data corresponding to the first target index stored in the database, wherein the first index corresponding to each stored record data is generated based on the split record data, and the split record data contains multiple character groups, and each character group contains consecutive characters in the record data whose length is less than the preset length.

[0159] In the technical solution provided by the embodiment of the present application, a first index is pre-constructed for each record data in the database table stored in the database. When performing data query, the first index is used to avoid traversing and scanning all the data in the table, which improves the query efficiency and at the same time reduces the waste of CPU resources and memory resources of the database.

[0160] In addition, in the technical solution provided by the embodiment of the present application, the first index is pre-generated based on the split record data, and the split record data contains multiple character groups. Each character group contains consecutive characters in the record data with a length less than the preset length. When the length of the target string to be queried is less than the preset length, the first index can be used to query the record data that matches the target string. That is, by adopting the technical solution provided by the embodiment of the present application, the length of the target string used for data query is not restricted during data query. When the length of the target string is less than the preset length, the index can also be used for data query, ensuring that target strings of any length can be queried based on the index, further improving the query efficiency of data query.

[0161] In one embodiment of the present application, the device further includes:

[0162] A second query module, configured to, when the length determination module determines that the length of the target string to be queried is not less than the preset length, query a second target index that matches the target string in the pre-generated second index, and determine the record data corresponding to the second target index stored in the database, where the second index corresponding to each record data is generated based on the complete record data.

[0163] In one embodiment of the present application, for each record data, the first index corresponding to the record data is generated based on a first index module:

[0164] The first index generation module is configured to extract character groups from the record data;

[0165] Construct a first index representing the correspondence between each character group and the position of the record data in the database.

[0166] In one embodiment of the present application, the device further includes:

[0167] An index update module, configured to, if the record data stored in the database changes, generate the first index and the second index corresponding to the changed record data.

[0168] In one embodiment of the present application, the device further includes:

[0169] A character removal module, configured to remove specified characters in a character group and / or remove duplicate character groups during the process of generating the first index.

[0170] In another embodiment provided by the present application, a computer-readable storage medium is further provided. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of any of the above data query methods are implemented.

[0171] In another embodiment provided by the present application, a computer program product including instructions is further provided. When it runs on a computer, the computer is caused to execute any of the data query methods in the above embodiments.

[0172] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more available media integrated. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)).

[0173] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0174] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of electronic devices and devices, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments.

[0175] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are included in the protection scope of the present application.

Claims

1. A data query method, characterized in that: The method comprises: Determine whether the length of the target character string to be queried is less than a preset length; If it is determined that the length of the target string to be queried is less than the preset length, a first target index matching the target string is queried in a pre-generated first index, and the record data corresponding to the first target index stored in the database is determined, wherein the first index corresponding to each record data is generated based on the split record data, and the split record data contains multiple character groups, each character group contains continuous characters in the record data and whose length is less than the preset length.

2. The method according to claim 1, characterized in that The method further comprises: If it is determined that the length of the target character string to be queried is not less than the preset length, a second target index matching the target character string is queried in the pre-generated second index, and the record data corresponding to the second target index stored in the database is determined, wherein the second index corresponding to each record data is generated based on the record data.

3. The method according to claim 1, characterized in that For each record data, the first index corresponding to the record data is generated based on the following method: extracting character groups from the record data; A first index is constructed to indicate the corresponding relationship between each character group and the position of the record data in the database.

4. The method according to any one of claims 1 to 3, characterized in that The method further comprises: If the record data stored in the database changes, a first index and a second index corresponding to the record data are generated for the changed record data.

5. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: In the process of generating the first index, designated characters in the character group are removed and / or repeated character groups are removed.

6. An electronic device, characterized in that: The electronic device comprises: processor; Transceiver; A machine-readable storage medium storing machine-executable instructions that can be executed by the processor, the machine-executable instructions causing the processor to perform the following steps: Determine whether the length of the target character string to be queried is less than a preset length; If it is determined that the length of the target string to be queried is less than the preset length, a first target index matching the target string is queried in a pre-generated first index, and the record data corresponding to the first target index stored in the database is determined, wherein the first index corresponding to each record data is generated based on the split record data, and the split record data contains multiple character groups, each character group contains continuous characters in the record data and whose length is less than the preset length.

7. The electronic device according to claim 6, characterized in that: The machine executable instructions further cause the processor to perform the following steps: If it is determined that the length of the target character string to be queried is not less than the preset length, a second target index matching the target character string is queried in the pre-generated second index, and the record data corresponding to the second target index stored in the database is determined, wherein the second index corresponding to each record data is generated based on the record data.

8. The electronic device according to claim 6, characterized in that: For each record data, the machine executable instructions further cause the processor to generate a first index corresponding to the record data based on the following method: extracting character groups from the record data; A first index is constructed to indicate the corresponding relationship between each character group and the position of the record data in the database.

9. A data query device, characterized in that: The device comprises: A length determination module is used to determine whether the length of the target character string to be queried is less than a preset length; A first query module is used to query a first target index matching the target string in a pre-generated first index when the length judgment module determines that the length of the target string to be queried is less than a preset length, and determine the record data corresponding to the first target index stored in the database, wherein the first index corresponding to each record data is generated based on the split record data, the split record data contains multiple character groups, and each character group contains continuous characters in the record data whose length is less than the preset length.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 1 to 5 are implemented.