Data processing method and device and storage medium

By determining the processing order based on the number of triples associated with the target graph data according to the data query conditions in large-scale graph data processing, the high cost and data inconsistency problems caused by additional index building in the prior art are solved, and efficient data query and processing are achieved.

CN120910316APending Publication Date: 2025-11-07BEIJING XINGYUN DIGITAL TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511052258.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-07

Smart Images

  • Figure CN120910316A_ABST
    Figure CN120910316A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device and a storage medium, and the method comprises the steps: responding to a data processing request for target graph data, determining a data processing type and a plurality of data query conditions corresponding to the data processing request, the target graph data comprises triples larger than a preset number; determining a processing sequence of the plurality of data query conditions according to the number of triples associated with entries corresponding to each data query condition in the target graph data; based on the processing sequence of the multiple data query conditions, performing query processing on each data query condition in the target graph data in sequence, and determining data query results corresponding to the multiple data query conditions; and determining a data processing result for the data processing request according to the data processing type and the data query results corresponding to the plurality of data query conditions.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the field of communication, and in particular, to a data processing method, device and storage medium. BACKGROUND

[0002] Knowledge graph technology can organize and store information through complex relationships between entities and represent structured data in the form of a graph, so knowledge graph technology is increasingly widely used.

[0003] In actual applications, graph data constructed based on knowledge graph technology and the like can be stored through a relational database, but when the data scale of the graph data is large and the retrieval conditions are many, data processing through the relational database has the problem of low data retrieval performance, and in order to improve the data retrieval performance for large-scale data, an index needs to be additionally established, which leads to high data storage and operation and maintenance costs, and problems such as data inconsistency. SUMMARY

[0004] Embodiments of the present application provide a solution to improve the data processing efficiency for large-scale graph data.

[0005] In a first aspect, a data processing method is provided, the method comprising: in response to a data processing request for target graph data, determining a data processing type corresponding to the data processing request and a plurality of data query conditions, the target graph data containing more than a predetermined number of triples; determining a processing order of the plurality of data query conditions according to the number of triples associated with each data query condition in the target graph data; based on the processing order of the plurality of data query conditions, sequentially querying and processing each data query condition in the target graph data to determine data query results corresponding to the plurality of data query conditions; and determining a data processing result for the data processing request according to the data processing type and the data query results corresponding to the plurality of data query conditions.

[0006] In a second aspect, a model training apparatus is provided, which comprises: a condition determining module configured to determine, in response to a data processing request for target graph data, a data processing type corresponding to the data processing request and a plurality of data query conditions, the target graph data containing more than a predetermined number of triples; a sequence determining module configured to determine a processing sequence of the plurality of data query conditions according to a number of triples associated with each data query condition in the target graph data; a data query module configured to sequentially query and process each data query condition in the target graph data based on the processing sequence of the plurality of data query conditions, and determine data query results corresponding to the plurality of data query conditions; and a result determining module configured to determine a data processing result for the data processing request according to the data processing type and the data query results corresponding to the plurality of data query conditions.

[0007] In a third aspect, a data processing device is provided, which comprises a processor and a memory, the memory storing programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of model training according to the first aspect.

[0008] In a fourth aspect, a readable storage medium is provided, which stores programs or instructions executable by a processor, and the programs or instructions, when executed by the processor, implement the steps of model training according to the first aspect.

[0009] The embodiments of the present application adopt the following technical solutions: In response to a data processing request for target graph data, a data processing type corresponding to the data processing request and a plurality of data query conditions are determined, wherein the target graph data contains more than a predetermined number of triples, a processing sequence of the plurality of data query conditions is determined according to a number of triples associated with each data query condition in the target graph data, each data query condition in the target graph data is sequentially queried and processed based on the processing sequence of the plurality of data query conditions, data query results corresponding to the plurality of data query conditions are determined, and a data processing result for the data processing request is determined according to the data processing type and the data query results corresponding to the plurality of data query conditions.

[0010] The above at least one technical solution adopted by the embodiments of the present application can achieve the following beneficial effects: On the one hand, by determining the processing order of the plurality of data query conditions, the query processing can be directly performed in the target graph data, which can avoid the problems of high data storage and operation and maintenance costs and data inconsistency existing in the data query requiring additional index establishment, on the other hand, since the number of triples associated with different retrieval conditions may not be the same, in the case of large-scale graph data and multiple retrieval conditions, the number of triples associated with the word corresponding to the data query condition in the target graph data is determined to determine the processing order of the plurality of data query conditions, and the data query efficiency can be improved by sequentially querying according to the processing order, thereby improving the data processing efficiency for large-scale graph data. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1 is a schematic flow chart of a data processing method according to an embodiment of the present application; Figure 2 is a schematic flow chart of a sentence parsing process according to an embodiment of the present application; Figure 3 is a schematic flow chart of a processing order determination process according to an embodiment of the present application; Figure 4 is a structural schematic diagram of a parse verification tree according to an embodiment of the present application; Figure 5 is a schematic flow chart of a processing order determination process according to another embodiment of the present application; Figure 6 is a schematic flow chart of a query operation processing process according to an embodiment of the present application; Figure 7 is a schematic flow chart of a deletion operation processing process according to an embodiment of the present application; Figure 8 is a schematic flow chart of a data query result determination process according to an embodiment of the present application; Figure 9 is a structural schematic diagram of a data processing system according to an embodiment of the present application; Figure 10 is a schematic diagram of a data processing flow of a data processing system according to an embodiment of the present application; Figure 11 is a structural schematic diagram of a model training apparatus according to an embodiment of the present application; Figure 12 is a structural schematic diagram of a data processing device according to an embodiment of the present application. DETAILED DESCRIPTION

[0012] The embodiments of the present specification provide a data processing method, device and storage medium.

[0013] In order for those skilled in the technical field to better understand the technical solutions in the specification, the technical solutions in the specification will be clearly and completely described below in combination with the drawings in the specification. Obviously, the described embodiments are only some of the embodiments of the specification, not all. Based on the embodiments in the specification, all other embodiments obtained by those of ordinary skill in the art without creative labor should fall within the scope of protection of the specification.

[0014] The application concept of the present application is as follows: knowledge graph technology can organize and store information through complex relationships between entities and represent structured data in the form of a graph, so knowledge graph technology is increasingly widely used. In actual application, graph data can be stored through a relational database or a graph database and the like, but when the data scale of graph data is large and the retrieval conditions are more, in order to ensure the data retrieval performance of the above-mentioned database, an index needs to be additionally established, which leads to high data storage and operation and maintenance costs, and data inconsistency and the like. Therefore, the embodiments of the specification provide a technical solution that can solve the above-mentioned problems, in which, in response to a data processing request for target graph data, the data processing type corresponding to the data processing request and a plurality of data query conditions are determined, wherein the target graph data contains more than a predetermined number of triples, the processing order of the plurality of data query conditions is determined according to the number of triples associated with the term corresponding to each data query condition in the target graph data, each data query condition in the target graph data is sequentially queried and processed based on the processing order of the plurality of data query conditions, the data query results corresponding to the plurality of data query conditions are determined, and the data processing result for the data processing request is determined according to the data processing type and the data query results corresponding to the plurality of data query conditions. In this way, on the one hand, by determining the processing order of the plurality of data query conditions, the query processing is directly performed in the target graph data, which can avoid the problems of high data storage and operation and maintenance costs and data inconsistency existing in the data query requiring an additional index, and on the other hand, since the number of triples associated with different retrieval conditions may not be the same, in the case that the data scale of graph data is large and the retrieval conditions are more, the processing order of the plurality of data query conditions is determined according to the number of triples associated with the term corresponding to each data query condition in the target graph data, so that the sequential query is performed according to the processing order, which can improve the data query efficiency, thereby improving the data processing efficiency for large-scale graph data. For details, see the following content.

[0015] In an embodiment, as Figure 1As shown, the embodiment of the present specification provides a data processing method, the execution subject of the method can be a server, wherein the server can be an independent server or a server cluster composed of multiple servers. The method can specifically include the following steps: In S102, in response to a data processing request for target graph data, the data processing type and multiple data query conditions corresponding to the data processing request are determined.

[0016] Among them, the target graph data contains more than a predetermined number of triples. The target graph data can be graph structure data containing more than a predetermined number of triples constructed based on data corresponding to a preset scene using knowledge graph technology, etc. For example, the target graph data can be graph structure data constructed based on supply chain data containing complex logical relationships in the supply chain field, or the target graph data can also be graph structure data constructed based on complex relationship data in the fields of data mining, community analysis, and blockchain, etc. The predetermined number can be 1 million, 5 million, etc. The triple can include an object, a relationship (or attribute), and an object value (or attribute value). The storage form of the triple can be (S, P, O), S is the object, P is the relationship or attribute, and O is the object value or attribute value. S and P can be stored in the form of a string. The storage form of O can be related to the data type. For example, if the data type of O is a string, O can be stored in the form of a string. If the data type of O is a date type, O can be stored in the form of a date type. If the data type of O is a floating-point type, O can be stored in the form of a floating-point type. If the data type of O is an integer type, O can be stored in the form of an integer type. In addition, for S, P, and O of the string type, fuzzy search and keyword search can be supported. For O of the date type, range search and size comparison of time can be supported. For O of the number type (i.e., integer type and floating-point type), size comparison can be supported. The data processing type can include data query type, data deletion type, etc.

[0017] In implementation, in response to a data processing request for target graph data, the server can obtain a data processing statement corresponding to the data processing request, and determine the data processing type and multiple data query conditions corresponding to the data processing request according to the data processing statement.

[0018] For example, the server can perform parsing processing on the data processing statement through a pre-trained parsing model (such as a fine-tuned large language model, a model constructed based on a preset deep learning algorithm, etc.) to determine the data processing type and multiple data query conditions corresponding to the data processing request.

[0019] Alternatively, the server can also parse the data processing statement by manual parsing (such as sending the data processing statement to a third party for parsing) to determine the data processing type and the plurality of data query conditions corresponding to the data processing request.

[0020] In addition, the above-mentioned method for determining the data processing type and the plurality of data query conditions corresponding to the data processing request is an optional and implementable determination method. In actual application scenarios, there can be various different determination methods, which are not specifically limited by the embodiments of the present specification.

[0021] In S104, the processing order of the plurality of data query conditions is determined according to the number of triples associated with the term corresponding to each data query condition in the target graph data.

[0022] In implementation, the server can determine the term corresponding to each data query condition. For example, the data query condition can include a to-be-queried triple, and the server can determine the term corresponding to each data query condition according to the to-be-queried triple included in each data query condition.

[0023] For example, the server can take the object, attribute, object value (or attribute value) included in the to-be-queried triple as the term corresponding to each data query condition, or the server can perform semantic analysis on the to-be-queried triple and determine the term corresponding to each data query condition according to the semantic analysis result. Specifically, the server can perform semantic analysis on the to-be-queried triple by using a large language model, and determine the semantic analysis result as the term corresponding to each data query condition.

[0024] After determining the term corresponding to each data query condition, the server can determine the association degree between each triple in the target graph data and the term corresponding to each data query condition, so as to determine the number of triples associated with the term corresponding to each data query condition in the target graph data according to the association degree.

[0025] For example, taking the term corresponding to the data query condition as including an object, an attribute, and an object value (or an attribute value), the server can determine the association degree between each triple in the target graph data and the term corresponding to each data query condition according to the similarity between the object included in the data query condition and the object included in each triple in the target graph data, the similarity between the attribute included in the data query condition and the attribute included in each triple in the target graph data, and the similarity between the object value (or the attribute value) included in the data query condition and the object value (or the attribute value) included in each triple in the target graph data.

[0026] Or, taking the word corresponding to the data query condition as an example, the server can use the large language model to determine the word corresponding to each triple in the target graph data, and determine the association degree between each triple in the target graph data and the word corresponding to each data query condition according to the similarity between the word corresponding to each triple in the target graph data and the word corresponding to each data query condition.

[0027] Then, the server can determine the number of triples associated with the word corresponding to each data query condition in the target graph data according to the association degree and the preset association degree threshold, such as the server can determine the number of triples associated with the word corresponding to each data query condition in the target graph data according to the triple with an association degree greater than the preset association degree threshold.

[0028] After determining the number of triples associated with the word corresponding to each data query condition in the target graph data, the server can determine the processing order of the plurality of data query conditions according to the number, such as the server can sort the number from small to large to obtain the processing order of the plurality of data query conditions, that is, the query priority of the data query condition with a smaller number of associated triples is higher.

[0029] In S106, based on the processing order of the plurality of data query conditions, each data query condition in the target graph data is processed in sequence to determine the data query result corresponding to the plurality of data query conditions.

[0030] In implementation, since the query priority of the data query condition with a smaller number of associated triples is higher, based on the processing order, the data query condition with a smaller number of associated triples can be processed in the target graph data first, and then after the end of the first query, the other unprocessed data query conditions can be updated according to the current query result, and the processing order of the plurality of data query conditions can be updated according to the number of triples associated with the word corresponding to the updated data query condition in the target graph data, so that by following the greedy principle of "putting data query conditions with more dependent data to the end as much as possible", the data query processing can improve the data query efficiency.

[0031] In S108, the data processing result for the data processing request is determined according to the data processing type and the data query result corresponding to the plurality of data query conditions.

[0032] In implementation, the server can perform different data processing according to the data query result corresponding to the plurality of data query conditions according to the different data processing types, for example, taking the data processing type as the data deletion type as an example, the server can delete the triple in the target graph data corresponding to the data query result corresponding to the plurality of data query conditions.

[0033] The embodiment of the present specification provides a data processing method, in response to a data processing request for target graph data, determining a data processing type corresponding to the data processing request and a plurality of data query conditions, wherein the target graph data contains more than a predetermined number of triples, determining the processing order of the plurality of data query conditions according to the number of triples associated with each data query condition in the target graph data, based on the processing order of the plurality of data query conditions, sequentially querying each data query condition in the target graph data, determining the data query result corresponding to the plurality of data query conditions, and determining the data processing result for the data processing request according to the data processing type and the data query result corresponding to the plurality of data query conditions. In this way, on the one hand, by determining the processing order of the plurality of data query conditions, the query processing is directly performed in the target graph data, which can avoid the problems of high data storage and operation and maintenance costs and data inconsistency existing in the need to establish additional indexes for data query, on the other hand, since the number of triples associated with different search conditions may not be the same, therefore, in the case of large-scale graph data and multiple search conditions, the processing order of the plurality of data query conditions is determined by the number of triples associated with the data query condition in the target graph data, and the sequential query is performed according to the processing order, which can improve the data query efficiency, thereby improving the data processing efficiency for large-scale graph data.

[0034] In practical applications, the specific processing method of determining the data processing type corresponding to the data processing request and the plurality of data query conditions in the above step S102 can be various, and the following provides an optional processing method, as shown in the following steps S1022-S1024. Figure 2

[0035] In S1022, the data processing statement corresponding to the data processing request is obtained.

[0036] The data processing statement can be a SparQL (SparQL Protocol and RDF Query Language) statement supporting the SparQL protocol.

[0037] In S1024, based on the pre-constructed regular expression, the data processing statement is parsed and processed to obtain the data processing type corresponding to the data processing request and the plurality of data query conditions.

[0038] ​In actual applications, in step S1024, the data processing request corresponding data processing type and the specific processing mode of the plurality of data query conditions can be variously obtained by analyzing the data processing statement based on the pre-constructed regular expression. An optional processing mode is provided below, which can specifically include the following steps A1-A4.

[0039] In A1, the first clause and the second clause corresponding to the data processing request are determined based on the pre-constructed regular expression and the preset keyword.

[0040] The first clause is a select clause or a delete clause, and the second clause is a where clause.

[0041] In implementation, the preset keyword can include select, delete, where, etc. The server can perform keyword matching processing on the data processing statement based on the pre-constructed regular expression and the preset keyword, so as to split the data processing statement into a head clause (i.e. the first clause) for determining the data processing type and a condition clause (i.e. the second clause) for determining the data query condition according to the matched keyword.

[0042] For example, assuming that the data processing statement is "delete data where {condition triple list}", the data processing statement can be split into a delete clause and a where clause based on the pre-constructed regular expression and the preset keyword, i.e. the first clause corresponding to the data processing request can be "delete data", and the second clause can be "where {condition triple list}".

[0043] Alternatively, assuming that the data processing statement is "select?a,?c FROM kg_answer where {?a friend?c.?a hr_post_label "IT software development engineer".?c employee_name "Zhang San"}", the data processing statement can be split into a select clause and a where clause based on the pre-constructed regular expression and the preset keyword, i.e. the first clause corresponding to the data processing request can be "select?a,?c FROM kg_answer", and the second clause can be "where {?a friend?c.?a hr_post_label "IT software development engineer".?c employee_name "Zhang San"}".

[0044] In A2, the data processing type corresponding to the data processing request is determined according to the first clause.

[0045] In implementation, for example, if the first clause is a select clause, the data processing type corresponding to the data processing request can be a data query type, and if the first clause is a delete clause, the data processing type corresponding to the data processing request can be a data deletion type.

[0046] In A3, based on the pre-constructed regular expression and the preset triple format, the plurality of conditional triples contained in the second clause is determined.

[0047] The preset triple format can be the format of the triples of the target graph data, that is, the preset triple format can be (S, P, O).

[0048] In implementation, the server can perform format matching processing on the second clause based on the pre-constructed regular expression and the preset triple format, to determine the plurality of conditional triples contained in the second clause. For example, assuming that the preset triple format is (S, P, O), and the second clause is "where {?a friend?c.?a hr_post_label "IT software development engineer".?c employee_name "Zhang San"}", then the triple matching the preset triple format in the second clause can be filtered by the regular expression, that is, the plurality of conditional triples obtained can be "?a friend?c", "?a hr_post_label "IT software development engineer", and "?c employee_name "Zhang San".

[0049] In A4, according to the plurality of conditional triples, the plurality of data query conditions corresponding to the data processing request is determined.

[0050] In implementation, the server can determine the conditional triple as the data query condition.

[0051] In actual application, in step S104, the specific processing manner of determining the processing order of the plurality of data query conditions according to the number of triples associated with the keyword corresponding to each data query condition in the target graph data can be various, and an optional processing manner is provided as follows. Figure 3 As shown in the figure, the specific processing can include the following steps S1042-S1048.

[0052] In S1042, according to the pre-constructed regular expression, the plurality of non-variable values in each data query condition is determined, and the plurality of keywords corresponding to each data query condition is determined.

[0053] In implementation, the server can disassemble each data query condition according to a pre-constructed regular expression, obtain parameters such as objects, strings, numerical lengths, and string variables corresponding to each data query condition, and then filter out non-variable values in the disassembled parameters through the regular expression. The server can determine the filtered multiple non-variable values as the multiple terms corresponding to the data query condition.

[0054] In this way, the server can adopt regular matching and tree-shaped verification to parse the data query statement layer by layer to determine the multiple terms corresponding to each data query condition.

[0055] Specifically, taking a data processing statement as a SparQL statement as an example, the server can construct a regular expression R according to the syntax of the SparQL statement, take anchoring the preset keywords such as select, where, and delete as a feature, form an analysis syntax tree as shown in Figure 4 , and parse the SparQL statement according to the analysis syntax tree. The specific parsing process is as follows: first, the server can remove the redundant spaces, blank lines, tab characters, and other preset characters in the SparQL statement, and then cut the user-input SparQL statement according to the regular expression R and the preset keywords to form the select clause, the where clause, and the delete clause. In this way, the SparQL statement can be divided into the select clause and the where clause, or the delete clause and the where clause, forming the root node and the first-level node of the analysis verification tree.

[0056] After the first clause and the second clause are parsed, the server can continue to disassemble the first clause and the second clause according to the regular expression R to form the next-level nodes of the analysis verification tree until the leaf nodes, wherein the leaf nodes can include string constants (such as strings wrapped in double quotes), variables (such as strings starting with “?”), objects (strings not wrapped in quotes), and constants (such as strings wrapped in double quotes followed by ^^ [int|float|date] strings, such as date, numerical type, and other strings).

[0057] In addition, the server can also perform regular verification on the data processing statement according to the analysis result of each layer. If the first clause and / or the second clause are not parsed at the first layer, the server can determine that the regular verification of the data processing statement fails, and at this time, the server can feed back the error of the SparQL statement.

[0058] In addition, since there is no insert and delete operation in the SparQL statement, the server can define the syntax of the delete and insert in advance. The delete syntax can be: delete where {condition triple}, and the insert syntax can be: insert data values {instance triple}.

[0059] In addition, the data processing type corresponding to the data processing request received by the server can also include a data insertion type, i.e., the first clause corresponding to the data processing request can also include an insert clause, and the second clause can also include a values clause. The server can determine the instance triple contained in the second clause according to the pre-constructed regular expression and the preset triple format, and then concatenate the instance triple to be inserted into a statement for a full-text search operation, and then execute the statement to insert the instance triple to be inserted into the target graph data.

[0060] In S1044, the triple in the triple of the target graph data corresponding to each term in the data query condition and having a preset association relationship is determined.

[0061] In actual application, the specific processing mode of the triple in the triple of the target graph data corresponding to each term in the data query condition and having a preset association relationship in the above step S1044 can be various, and an optional processing mode is provided below, which can include the following steps B1-B3.

[0062] In B1, the target term in the term corresponding to the data query condition and related to the object value or the attribute value is determined.

[0063] In implementation, taking the case that the term corresponding to the data query condition includes one or more of S, P, and O of the condition triple as an example, if there is a term corresponding to O in the term corresponding to the data query condition, the term can be determined as the target term.

[0064] In B2, the sub-database in the index database corresponding to the data query condition is determined according to the target term in the term corresponding to the data query condition and the data type of the target term.

[0065] The object value or the attribute value contained in the triple of the target graph data stored in the sub-database of the index database has the same data type as the data type of the target term.

[0066] In implementation, before determining the sub-database corresponding to the data query condition in the index database according to the target term and the data type of the target term in the term corresponding to the data query condition, the server can store the target graph data in the index database, so that, since the target graph data itself is indexed data due to the storage of the target graph data under the full-text index structure, the target graph data can be directly stored in the index database based on the storage architecture of the index, without the need to store the index data corresponding to the target graph data, thereby avoiding the consistency problem of data storage caused by the separation of graph data and index data, realizing the unified storage of data and index, reducing the operation and maintenance cost and storage cost, and supporting the SparQL statement at the same time, and being more consistent with the storage standard of the knowledge graph.

[0067] When storing the target graph data, a plurality of index structures can be defined in advance, for example, taking the data type of O as an example, including string type, date type, floating point type and integer type, each index can include three fields, namely S, P and O. Among them, the index types of S and P are the same, and both support keyword type and text type.

[0068] The server can create corresponding sub-databases in the index database according to the type of O, for example, the server can create sub-database 1 according to the string type, create sub-database 2 according to the date type, create sub-database 3 according to the floating point type, and create sub-database 4 according to the integer type, wherein the sub-database 1 can support keyword type and text type, the sub-database 2 can support date type, the sub-database 3 can support floating point type, and the sub-database 4 can support integer type.

[0069] The server can store the triples contained in the target graph data into the sub-databases of the index database according to the data type of O in the triples contained in the target graph data.

[0070] In addition, the above is an example of indexing classification of the data type of O including string type, date type, floating point type and integer type, and in actual application, the data type of O can be various, which is not limited in the embodiments of the present application.

[0071] After storing the triples contained in the target graph data in the index database, the server can determine the sub-database corresponding to the data query condition in the index database according to the target term and the data type of the target term in the term corresponding to the data query condition.

[0072] In B3, in each sub-database corresponding to the data query condition, the triples having a preset association relationship with the target term are determined.

[0073] In implementation, for example, assuming that the terms corresponding to the data query condition include S and O, and the data type of O is date type, then the sub-database corresponding to the data query condition in the index database is the sub-database 2 supporting date type, and the server can determine the triplets having the preset association relationship with the target term according to O and S in the sub-database 2.

[0074] In which, since different sub-databases are created according to the data type of O and support the retrieval type corresponding to the data type of O, such as the sub-database 1 supporting the retrieval of keyword type and text type, and the sub-database 2 supporting the retrieval of date type, therefore, in the query scenarios such as range query, size comparison query, string fuzzy query, time query, etc., the fast triplet search can be realized through the corresponding sub-database, and the data search efficiency is improved.

[0075] In addition, if the terms corresponding to the data query condition do not include the target terms related to the object value or the attribute value, that is, the terms corresponding to the data query condition are S and / or P, then the server can query the associated triplets according to the terms in the index database.

[0076] In S1046, the number of triplets corresponding to each data query condition is determined according to the number of triplets associated with each term corresponding to each data query condition.

[0077] In implementation, the server can determine the minimum number of triplets associated with each term corresponding to the data query condition as the number of triplets corresponding to the data query condition. For example, assuming that the number of triplets associated with term 1 corresponding to the data query condition is 5000, the number of triplets associated with term 2 is 700, and the number of triplets associated with term 3 is 6, then the number of triplets corresponding to the data query condition can be 6.

[0078] In S1048, the processing order of the multiple data query conditions is determined according to the number of triplets corresponding to each data query condition.

[0079] In implementation, the processing order of the multiple data query conditions can be sorted from small to large according to the number of triplets corresponding to each data query condition.

[0080] In actual application, the specific processing manner of determining the processing order of the multiple data query conditions according to the number of triplets associated with each term corresponding to each data query condition in the target graph data in the above step S104 can be various, and an optional processing manner is provided as follows. Figure 5 As shown in the figure, the specific processing can include the following steps S10410-S10412.

[0081] In S10410, it is determined that the query type to which each data query condition corresponds.

[0082] The query type can include an object query type and an attribute query type.

[0083] In S10412, the processing order of the plurality of data query conditions is determined according to the query type to which each data query condition corresponds and the number of triples associated with each data query condition in the target graph data.

[0084] In implementation, since the attribute repetition rate is higher than the object repetition rate, the query priority of the object query type can be higher than the query priority of the attribute query type. For example, the server can first divide the plurality of data query conditions into two categories according to the query type to which each data query condition corresponds, and then arrange the data query conditions in each category in ascending order based on the number of triples associated with each data query condition in the target graph data. Then, the sorted data query conditions corresponding to the two categories are sorted according to the query type to which each data query condition corresponds, and the rule that the query priority of the object query type is higher than the query priority of the attribute query type, to obtain the processing order of the plurality of data query conditions.

[0085] For example, assume that the data query conditions include data query condition 1, data query condition 2, data query condition 3, and data query condition 4, wherein the query type to which the data query condition 1 and the data query condition 3 correspond is the object query type, the query type to which the data query condition 2 and the data query condition 4 correspond is the attribute query type, and the number of triples associated with the data query condition 1, the data query condition 2, the data query condition 3, and the data query condition 4 in the target graph data is 100, 200, 300, and 400, respectively.

[0086] Then, the server can sort the data query type 1 and the data query type 3 according to the number of triples associated with the data query type 1 and the data query type 3 in the target graph data, and sort the data query type 2 and the data query type 4 according to the number of triples associated with the data query type 2 and the data query type 4 in the target graph data. Finally, the processing order obtained by aggregating the sorting is data query condition 1, data query condition 3, data query condition 2, and data query condition 4.

[0087] In addition, in order to prevent memory overflow and ensure performance, each condition query can be limited within a threshold, such as 10,000.

[0088] In actual applications, in step S108, the specific processing manner of the data processing result for the data processing request can be various according to the data processing type and the data query result corresponding to the plurality of data query conditions. The data processing type can include a data query type corresponding to a select clause. Accordingly, an optional processing manner is provided as follows. Figure 6 As shown in FIG. 8, the specific processing can include the following steps S1082-S1086.

[0089] In S1082, the data query results corresponding to the plurality of data query conditions are spliced to obtain a target query result.

[0090] In S1084, the data filtering condition corresponding to the select clause is determined.

[0091] In implementation, the data filtering condition corresponding to the select clause can be determined according to the variable identifier in the select clause. For example, assuming that the select clause is “select?a” and the variable identifier in the clause is a, the corresponding data filtering condition can be: filtering out the object a.

[0092] In S1086, the target query result is filtered according to the data filtering condition to obtain the data processing result for the data processing request.

[0093] In implementation, since the target query result can contain redundant information, the target query result can be filtered according to the data filtering condition corresponding to the select clause to obtain the data processing result for the data processing request.

[0094] In actual applications, in step S108, the specific processing manner of the data processing result for the data processing request can be various according to the data processing type and the data query result corresponding to the plurality of data query conditions. The data processing type can include a data query type corresponding to a select clause. Accordingly, an optional processing manner is provided as follows. Figure 7 As shown in FIG. 8, the specific processing can include the following steps S1082-S1086.

[0095] In S1088, the data query results corresponding to the plurality of data query conditions are spliced to obtain a target query result.

[0096] In S10810, the triple corresponding to the target query result in the target graph data is deleted.

[0097] In implementation, the server can pre-store the word entries corresponding to each triple in the target graph data, so as to quickly determine the number of triples in the target graph data associated with the word entry corresponding to the data query condition according to the pre-stored word entries. For example, the server can determine the number of triples in the target graph data associated with each word entry according to the pre-stored word entries, and obtain the word entries related to the word entry corresponding to the data query condition from the pre-stored word entries, so as to determine the number of triples in the target graph data associated with the word entry corresponding to the data query condition according to the number of triples in the target graph data associated with the obtained word entries.

[0098] Therefore, after the server deletes the triples in the target graph data corresponding to the target query result, the server can update the number of triples in the target graph data associated with the pre-stored word entries, so that the number of triples in the target graph data associated with the word entry corresponding to the data query condition can be quickly and accurately determined according to the updated number in the next data processing.

[0099] Similarly, after the data insertion processing is performed, the server can also update the number of triples in the target graph data associated with the pre-stored word entries according to the target graph data after the insertion.

[0100] In actual application, in step S106, the specific processing manner of the data query results corresponding to the plurality of data query conditions can be various based on the processing order of the plurality of data query conditions, and an optional processing manner is provided as follows. Figure 8 As shown in the figure, the specific processing can include the following steps S1062-S10610.

[0101] In S1062, the first query condition in the plurality of data query conditions is determined according to the processing order of the plurality of data query conditions.

[0102] In S1064, the first data query result is obtained by performing query processing on the target data according to the first query condition.

[0103] In S1066, the data query condition related to the first query condition in the data query condition other than the first query condition is updated according to the first data query result, and the updated data query condition is obtained.

[0104] In S1068, the processing order of the plurality of data query conditions is updated according to the number of triples in the target graph data associated with the word entry corresponding to the updated data query condition, and the updated processing order of the plurality of data query conditions is obtained.

[0105] In S10610, based on the processing order of the plurality of data query conditions, the updated data query conditions are sequentially queried in the target graph data to obtain data query results corresponding to the plurality of data query conditions.

[0106] In implementation, for example, it is assumed that the data processing statement is "select?a,?c from kg_answer where{?a friend?c.?a hr_post_label "IT software development engineer".?c employee_name "Zhang San"}.".

[0107] According to the above regular tree (i.e., the parsing verification tree), the select clause and the where clause are extracted. The first clause obtained after the regular extraction processing, i.e., the select clause, can be "select?a?c". The second clause extracted, i.e., the where clause, can be "where {?a friend?c.?a hr_post_label "software development engineer".?c employee_name "Zhang San"}." Further extracting the data filtering condition in the select clause, i.e., the variable list

?a,?c

[0108] The server can traverse the three data query conditions in the data query list of the where clause, and take out the number of triples associated with the non-variable objects or constants (i.e., non-variable values) in each node (i.e., data query condition) in the target graph data.

[0109] The non-variable values corresponding to each data query condition obtained by regular extraction of the three data query conditions by the pre-constructed regular expression R can be: friend in data query condition 1, hr_post_label and IT software development engineer in data query condition 2, and employee_name and Zhang San in data query condition 3.

[0110] Then, the server can determine the number of triples associated with the target graph data for each non-variable value (i.e., term) in the data query condition, for example, the number of triples associated with the target graph data for "friend" can be 5000, the number of triples associated with the target graph data for "IT software development engineer" can be 700, the number of triples associated with the target graph data for "hr_post_label" can be 6000, the number of triples associated with the target graph data for "IT software development engineer" can be 700, the number of triples associated with the target graph data for "employee_name" can be 60000, and the number of triples associated with the target graph data for "Zhang San" can be 6.

[0111] The server can take the minimum number of triples associated with the target graph data for each term corresponding to the data query condition as the number of triples corresponding to the data query condition, i.e., the number of triples corresponding to the data query condition 1 can be 5000, the number of triples corresponding to the data query condition 2 can be 700, and the number of triples corresponding to the data query condition 3 can be 6. The server can sort the three data query conditions in ascending order according to the number of triples corresponding to the data query conditions to obtain the processing order of the multiple data query conditions.

[0112] According to the processing order of the above three data query conditions (i.e., data query condition 3, data query condition 2, and data query condition 1), the data query condition 3 (i.e.,?c employee_name "Zhang San") can be determined as the first query condition, and the execution of the first query condition can obtain the first data query result. Assuming that the first data query result is R1 (assuming that the length of R1 is 6), if R1 is empty, the query is ended directly. If R1 is not empty, then the remaining data query conditions related to the data query condition 3 can be updated according to the first query result R1 to obtain the updated data query conditions.

[0113] For example, the server can fill the first query result R1 into the variable?c of "Zhang San" in the employee_name, and the number of triples corresponding to?c in all data query conditions becomes the number of R1. The server can reorder the updated data query conditions (i.e., the data query conditions that have not performed the query operation). Then, the number of triples associated with the terms corresponding to the remaining two data query conditions is: Count ["hr_post_label"] = 60000, Count ["IT software development engineer"] = 700; Count ["friend"] = 5000, Count ["?c"] = 6. That is, the number of triples corresponding to data query condition 2 is 700, and the number of triples corresponding to data query condition 1 is 6. According to the sorting from small to large, the processing order of the two data query conditions is determined as: data query condition 1, data query condition 2, that is, the query plan priority is changed to?c employee_name "Zhang San";?a hr_post_label "IT software development engineer".

[0114] The server can then execute the first data query condition (i.e., data query condition 1) according to the updated processing order, obtain the execution result R2, and can fill R2 into other data query conditions related to data query condition 1 according to the above processing process, continue to update and reorder the data query conditions, until all data query conditions are executed, and obtain the final query result R3 (here, 3 data query conditions, so the query result has 3 stages, and the final query result is R3).

[0115] The server can perform a full connection operation on the results of the query in the order of R3, R2, R1 in reverse order to obtain the final result of the query (i.e., the target query result), and can obtain the fields?a and?c that need to be retained according to the variables in the select clause, and filter the target query result obtained according to?a and?c, that is, only the results corresponding to?a and?c in the target query result are retained.

[0116] In this way, the above data processing method can construct a data processing system as shown in Figure 9 The data processing system can include a query input module, a statement parsing module, a knowledge graph storage module, a knowledge graph management module, and a query result output module.

[0117] The query input module can be used to receive a data processing statement corresponding to a data processing request for a target graph data (such as a standard language SparQL statement for querying a knowledge graph), and can verify the correctness of the input data processing statement.

[0118] The sentence parsing module can be used to anchor keywords and query patterns in the data query statement, and formulate a regular expression R for parsing the data processing statement to decompose the query task.

[0119] Since different data types support different retrieval methods, in the knowledge graph storage module, an index corresponding to each data type (such as string, integer, date, etc.) can be constructed based on full-text retrieval according to the data type of O in the triples in the target graph data. For example, for string type, fuzzy search and keyword search need to be supported when building the index, for date type, range query and size comparison support are needed, and for numerical type (such as floating point type and integer type), size comparison support is needed. After building the index based on this method, the triples of the target graph data can be stored in the corresponding sub-database of the index database according to the data type of O in the triples, to meet the knowledge graph storage requirements such as deletion, change, and insertion.

[0120] In the knowledge graph management module, the server can first divide the ontology part and the instance part of the target graph data into two modules based on the use scenario of the target graph data, but when storing internally, both the ontology part and the instance part can be stored in the form of triples, and the ontology part and the instance part can be associated by setting system keywords (such as "IS"). Since the background storage is based on full-text retrieval, there is no need to manage the index additionally. During user use, the server can design an optimized query logic based on the query task decomposed by parsing the SparQL statement, find the data in the knowledge graph storage module based on the logic, and finally transmit it to the user.

[0121] The query result output module can be used to format the output of the found data. First, the server can determine the processing order of multiple data query conditions according to the number of triples associated with each data query condition in the target graph data, and the more triples associated with the data query condition, the later the execution order. At the same time, in order to prevent memory overflow and ensure performance, each condition query can be limited within a threshold (such as 10,000). In addition, since the probability of attribute repetition is higher than that of entity repetition, the query priority of entity nodes (i.e. the query priority of the word corresponding to the object type) can be higher than that of attribute nodes (i.e. the query priority of the word corresponding to the attribute query type). When deleting and inserting, the server can update the number of triples associated with the corresponding instance in the target graph data in real time, which provides support for determining the number of triples associated with each data query condition. In addition, the update operation can be processed by the method of deleting first and then inserting.

[0122] Specifically, the data processing relationship between each module can be as shown in Figure 10 The knowledge graph storage module can be used to establish a full-text retrieval index suitable for knowledge graph storage. Before knowledge graph storage and query, the storage structure of the knowledge graph can be defined, and then how to construct the storage form based on the full-text retrieval can be determined. The index form of the full-text retrieval can be referred to the description in the embodiments of the present specification, and will not be described here.

[0123] The query input module can be used to check the rationality of the SparQL statement. Regular matching and tree-shaped checking can be adopted to check the correctness of the SparQL statement. The statement checking manner of SparQL can also be referred to the description in the embodiments of the present specification, and will not be described here.

[0124] The statement analysis module can be used to establish an analyzer for analyzing the standard query language (SparQL) of the knowledge graph, to analyze the SparQL statement and generate an execution plan. In the case that the SparQL statement passes the regular check, a correct SparQL statement and corresponding select clause, where clause, insert data clause and delete clause can be obtained.

[0125] The variable in the select clause can be used to filter the object returned by the data query statement (i.e., the variable in the select clause can be used to determine the data filtering condition corresponding to the select clause). The constant and variable in the where clause can constitute a query condition list L={L1, L2,..., Ln} in the form of a triple. n is the number of data query conditions. The values clause corresponding to the insert clause can be used to determine the list of triple to be inserted. The delete clause can be used to determine the data processing type as a data deletion type, to represent that the delete operation needs to be performed.

[0126] The server can determine the operation type of a data processing statement by whether it contains select, insert, or delete clauses. For example, if a data processing statement contains a select clause, it is a query operation, meaning the corresponding data processing type is a data query. If a data processing statement contains an insert clause, it is a data insertion operation. If a data processing statement contains a delete clause, it is a data deletion operation. Furthermore, a single data processing statement can contain any one of the select, insert, or delete clauses, but it cannot contain two or more of these clauses simultaneously. If two or more of these clauses are present, it indicates a syntax error in the data processing statement.

[0127] The knowledge graph management module can be used to support the insertion, deletion, modification, and query processing of knowledge graphs. Specifically, when inserting triples, when the statement parsing module parses a data processing statement containing an insert clause, it can determine the triples to be inserted based on the values ​​clause. The server can then concatenate the triples to be inserted into a statement for full-text search operation, execute the statement, and return the number of inserted data records. At the same time, the server can update the number of triples associated with the terms corresponding to the inserted data records in the target graph data.

[0128] When deleting a triple, if the statement parsing module parses a data processing statement containing a delete clause, it can construct data query conditions based on the where clause in the data processing statement, concatenate the statement for the full-text search deletion operation, and then execute the statement. At the same time, it can return the deleted data records and update the number of triples associated with the terms corresponding to the deleted records in the target graph data.

[0129] When modifying triples, you can delete them first and then insert them.

[0130] When querying triples, the query conditions in the query condition list FL can be traversed in processing order. A full-text search query statement is constructed based on the query conditions, and the query results are retrieved. These results are then added to the next query condition, and this process is repeated until the final result is obtained. The specific algorithm flow is as follows:

[0131] The knowledge graph query optimization is used for optimizing the query efficiency of the knowledge graph. In the query process, the number of triples corresponding to each data query condition can be different. Therefore, a greedy principle of "trying to put the condition with a large amount of dependent data to the last execution" can be followed to design a query efficiency optimization algorithm. The specific method is as follows: Based on the data query condition list FL, the number of triples associated with the target graph data corresponding to each data query condition is obtained, and the number is reordered to obtain the condition query list FL={f1, f2,..., fn}. If the data query condition corresponding to the word item and the triples contained in the target graph data do not have a preset corresponding relationship, the number of triples corresponding to the data query condition is 0. At the same time, the data query condition with all variable word items can be arranged at the end of the processing order. If a data query condition contains one or more objects or constants, the maximum value of the number of associated triples can be taken as the number of triples corresponding to the data query condition.

[0132] The data query statement FL list is executed for query. The query logic includes updating the number of triples C associated with each data query condition at each condition query, and early query pruning. The specific algorithm is as follows:

[0133] The query result output module can output the final data processing result according to the data processing type corresponding to the first clause. For example, in the case of data query processing, the query result R can contain a large amount of redundant information. Therefore, the corresponding data filtering condition can be determined according to the variable identifier of the select clause, the query result R can be filtered according to the data filtering condition, and the filtering result can be returned. In this way, the content corresponding to the target graph data can be output in the format of the select clause. For the three operations of deletion, modification and insertion, the success or failure of the operation can be returned as a state code.

[0134] In this way, through the high-performance index mechanism, the storage efficiency and query performance of the data can be improved, and distributed storage and processing of the data can be implemented to ensure the consistency and security of the data and reduce the maintenance cost of the system. On the basis of the storage mode, the alignment and query performance optimization of the SparQL statement can be implemented to make it more suitable for the application scenarios of the knowledge graph, provide strong backend support for the application of the knowledge graph, and play a greater role in the scenarios of search engines, recommendation systems, intelligent question answering and the like.

[0135] The application scenarios of the embodiments of the present specification can include scenarios such as an intelligent customer service system scenario combining voice recognition technology and natural language processing technology, and a scenario of providing intelligent and personalized solutions combining machine learning and data analysis technology.

[0136] In addition, a full-text search can be performed by using an ElasticSearch search engine, and the corresponding index configuration under E can be as shown in Table 1.

[0137] Table 1

[0138] After the above configuration is completed, data processing can be performed by using ES. For example, in a data deletion scenario, a data processing statement can be: delete data values{condition triple list}, the server can split the data processing statement into a first clause (i.e., a delete clause) and a second clause (i.e., a where clause) according to a pre-constructed regular expression. Further, according to the where clause, a data query condition list containing multiple data query conditions, i.e.,

condition triple list

[0139] The server can splice an ES deletion statement according to the data query condition list, execute the deletion statement, complete the deletion operation, and update the number of triples associated with the word in the target graph data after the deletion is completed.

[0140] For example, in a data insertion scenario, a data processing statement can be: insert data{to-be-inserted triple 1. to-be-inserted triple 2. to-be-inserted triple 3}, the server can split the data processing statement into an insert clause and an insertion clause according to a pre-constructed regular expression, wherein the insert clause is “insert data”, and the insertion clause is: to-be-inserted triple 1. to-be-inserted triple 2. to-be-inserted triple 3.

[0141] The server can determine a triple list

to-be-inserted triple 1. to-be-inserted triple 2. to-be-inserted triple 3

[0142] The server can execute the batch insertion statement in ES, thereby completing the data insertion operation, and updating the number of triples associated with the word in the target graph data after the insertion is completed.

[0143] In this way, compared with the storage manner of the graph database combined with the ES, the storage logic of the embodiments of the present specification is simpler through the storage manner of the RDB combined with the ES and the graph database and the like, and the integrated storage of the ontology, the instance and the index can be realized only through the full-text search, the operation and maintenance cost is reduced, and the business logic is simplified. In addition, the full-text search system based on the ES in the embodiments of the present specification can support large-scale distributed graph data storage and query, and has high data processing efficiency.

[0144] The embodiments of the present specification provide a data processing method, in response to a data processing request for target graph data, determining a data processing type corresponding to the data processing request and a plurality of data query conditions, wherein the target graph data contains more than a predetermined number of triples, determining the processing order of the plurality of data query conditions according to the number of triples associated with each data query condition in the target graph data, based on the processing order of the plurality of data query conditions, sequentially querying and processing each data query condition in the target graph data, determining the data query result corresponding to the plurality of data query conditions, and determining the data processing result for the data processing request according to the data processing type and the data query result corresponding to the plurality of data query conditions. In this way, on the one hand, by determining the processing order of the plurality of data query conditions, the query processing is directly performed in the target graph data, which can avoid the problems of high data storage and operation and maintenance cost and data inconsistency existing in the additional index for data query, on the other hand, since the number of triples associated with different search conditions may not be the same, in the case of large-scale graph data and multiple search conditions, the processing order of the plurality of data query conditions is determined according to the number of triples associated with each data query condition in the target graph data, and the data query is sequentially performed according to the processing order, which can improve the data query efficiency, thereby improving the data processing efficiency for large-scale graph data.

[0145] In another embodiment, based on the same idea as the data processing method provided by the embodiments of the present specification, the embodiments of the present specification also provide a model training device, as shown in Figure 5 .

[0146] The model training device comprises a condition determination module 1101, an order determination module 1102, a data query module 1103 and a result determination module 1104, wherein: The condition determination module 1101 is configured to, in response to a data processing request for target graph data, determine a data processing type corresponding to the data processing request and a plurality of data query conditions, wherein the target graph data contains more than a predetermined number of triples. The sequence determining module 1102 is configured to determine the processing sequence of the multiple data query conditions according to the number of triples associated with each of the data query conditions in the target graph data. The data query module 1103 is configured to sequentially perform query processing on each of the data query conditions in the target graph data based on the processing sequence of the multiple data query conditions, and determine data query results corresponding to the multiple data query conditions. The result determining module 1104 is configured to determine a data processing result for the data processing request according to the data processing type and the data query results corresponding to the multiple data query conditions.

[0147] In the embodiments of the present specification, the condition determining module 1101 is configured to: obtain a data processing statement corresponding to the data processing request; perform parsing processing on the data processing statement based on a pre-constructed regular expression, to obtain a data processing type and multiple data query conditions corresponding to the data processing request.

[0148] In the embodiments of the present specification, the condition determining module 1101 is configured to: determine a first clause and a second clause corresponding to the data processing request based on the pre-constructed regular expression and a preset keyword; the first clause is a select clause or a delete clause, and the second clause is a where clause; determine the data processing type corresponding to the data processing request according to the first clause; determine multiple conditional triples included in the second clause based on the pre-constructed regular expression and a preset triple format; determine the multiple data query conditions corresponding to the data processing request according to the multiple conditional triples.

[0149] In the embodiments of the present specification, the sequence determining module 1102 is configured to: determine multiple non-variable values in each of the data query conditions according to the pre-constructed regular expression, to determine multiple terms corresponding to each of the data query conditions; determine triples in the triples of the target graph data that have a preset association relationship with each of the terms corresponding to the data query conditions; determine the number of triples corresponding to each of the data query conditions according to the number of triples associated with each of the terms corresponding to each of the data query conditions; determine the processing sequence of the multiple data query conditions according to the number of triples corresponding to each of the data query conditions.

[0150] In an embodiment of the present specification, the order determining module 1102 is configured to: determine a target term related to the object value or the attribute value in the term corresponding to the data query condition; determine a sub-database corresponding to the data query condition in the index database according to the target term in the term corresponding to the data query condition and the data type of the target term; wherein the data type of the object value or the attribute value contained in the triple of the target graph data stored in the sub-database of the index database is the same as the data type of the target term; determine the triple having the preset association relationship with the target term in each sub-database corresponding to the data query condition.

[0151] In an embodiment of the present specification, the order determining module 1102 is configured to: determine a query type to which each term corresponding to the data query condition belongs, the query type including an object query type and an attribute query type; determine the processing order of the plurality of data query conditions according to the query type to which each term corresponding to the data query condition belongs and the number of triples associated with the term in the target graph data.

[0152] In an embodiment of the present specification, the data processing type includes a data query type corresponding to a select clause, and the result determining module 1104 is configured to: perform splicing processing on the data query results corresponding to the plurality of data query conditions to obtain a target query result; determine a data filtering condition corresponding to the select clause; perform filtering processing on the target query result according to the data filtering condition to obtain a data processing result for the data processing request.

[0153] In an embodiment of the present specification, the data processing type includes a data deletion type corresponding to a delete clause, and the result determining module 1104 is configured to: perform splicing processing on the data query results corresponding to the plurality of data query conditions to obtain a target query result; delete the triple corresponding to the target query result in the target graph data.

[0154] In an embodiment of the present specification, the data query module 1103 is configured to: determine a first query condition in the plurality of data query conditions according to the processing order of the plurality of data query conditions; According to the first query condition, query processing is performed on the target data to obtain a first data query result; According to the first data query result, a data query condition related to the first query condition in the data query condition except the first query condition is updated to obtain an updated data query condition; According to the number of triples associated with the updated data query condition in the target graph data, the processing order of the plurality of data query conditions is updated to obtain an updated processing order of the plurality of data query conditions; Based on the updated processing order of the plurality of data query conditions, the updated data query condition in the target graph data is continuously processed in sequence to obtain a data query result corresponding to the plurality of data query conditions.

[0155] The embodiment of the present specification provides a model training device, in response to a data processing request for target graph data, determining a data processing type corresponding to the data processing request and a plurality of data query conditions, wherein the target graph data contains more than a predetermined number of triples, determining the processing order of the plurality of data query conditions according to the number of triples associated with the term corresponding to each data query condition in the target graph data, based on the processing order of the plurality of data query conditions, sequentially querying each data query condition in the target graph data, determining the data query result corresponding to the plurality of data query conditions, and determining the data processing result for the data processing request according to the data processing type and the data query result corresponding to the plurality of data query conditions. In this way, on the one hand, by determining the processing order of the plurality of data query conditions, the query processing is directly performed in the target graph data, which can avoid the high data storage and operation and maintenance costs and data inconsistency problems existing in the additional index establishment for data query, on the other hand, since the number of triples associated with different search conditions may not be the same, therefore, in the case of large-scale graph data and multiple search conditions, the processing order of the plurality of data query conditions is determined according to the number of triples associated with the term corresponding to the data query condition in the target graph data, and the query is sequentially performed according to the processing order, which can improve the data query efficiency, thereby improving the data processing efficiency for large-scale graph data.

[0156] In another embodiment, based on the same idea, the present specification also provides a data processing device, as shown in Figure 12 .

[0157] The data processing device can have a large difference due to different configurations or performances, and can include one or more processors 1201 and memories 1202, and one or more storage applications or data can be stored in the memories 1202. Among them, the memory 1202 can be temporary storage or persistent storage. The application stored in the memory 1202 can include one or more modules (not shown in the figure), and each module can include a series of computer executable instructions in the data processing device. Further, the processor 1201 can be configured to communicate with the memory 1202 and execute a series of computer executable instructions in the memory 1202 on the data processing device. The data processing device can also include one or more power supplies 1203, one or more wired or wireless network interfaces 1204, one or more input / output interfaces 1206, and one or more keyboards 1206.

[0158] In particular, in the embodiment, the data processing device includes a memory and one or more programs, wherein one or more programs are stored in the memory, and one or more programs can include one or more modules, and each module can include a series of computer executable instructions in the data processing device, and the one or more programs executed by the one or more processors include computer executable instructions for: In response to a data processing request for target graph data, determining a data processing type and a plurality of data query conditions corresponding to the data processing request, the target graph data containing more than a predetermined number of triples; According to the number of triples associated with each data query condition in the target graph data, determining the processing order of the plurality of data query conditions; Based on the processing order of the plurality of data query conditions, sequentially querying and processing each data query condition in the target graph data to determine the data query result corresponding to the plurality of data query conditions; According to the data processing type and the data query result corresponding to the plurality of data query conditions, determining the data processing result for the data processing request.

[0159] Optionally, the determination of the data processing type and the plurality of data query conditions corresponding to the data processing request comprises: Obtaining a data processing statement corresponding to the data processing request; Based on a pre-constructed regular expression, the data processing statement is parsed and processed to obtain the data processing type and the plurality of data query conditions corresponding to the data processing request.

[0160] Optionally, the data processing request is parsed based on the pre-constructed regular expression to obtain a data processing type and a plurality of data query conditions corresponding to the data processing request, comprising: determining a first clause and a second clause corresponding to the data processing request based on the pre-constructed regular expression and a preset keyword; wherein the first clause is a select clause or a delete clause, and the second clause is a where clause; determining a data processing type corresponding to the data processing request according to the first clause; determining a plurality of conditional triples contained in the second clause based on the pre-constructed regular expression and a preset triple format; determining a plurality of data query conditions corresponding to the data processing request according to the plurality of conditional triples.

[0161] Optionally, the processing order of the plurality of data query conditions is determined according to the number of triples associated with each term corresponding to each data query condition in the target graph data, comprising: determining a plurality of terms corresponding to each data query condition according to a plurality of non-variable values in each data query condition determined based on the pre-constructed regular expression; determining triples in the triples of the target graph data that have a preset association relationship with each term corresponding to the data query condition; determining the number of triples corresponding to each data query condition according to the number of triples associated with each term corresponding to each data query condition; determining the processing order of the plurality of data query conditions according to the number of triples corresponding to each data query condition.

[0162] Optionally, the determination of the triples in the triples of the target graph data that have a preset association relationship with each term corresponding to the data query condition comprises: determining a target term related to an object value or an attribute value in the term corresponding to the data query condition; determining a sub-database corresponding to the data query condition in an index database according to the target term in the term corresponding to the data query condition and the data type of the target term; wherein the data type of the object value or the attribute value contained in the triples of the target graph data stored in the sub-database of the index database is the same as the data type of the target term; determining triples having the preset association relationship with the target term in each sub-database corresponding to the data query condition.

[0163] Optionally, the processing order of the plurality of data query conditions is determined according to the number of triples associated with the term corresponding to each of the data query conditions in the target graph data. determining a query type to which the term corresponding to each of the data query conditions belongs, the query type including an object query type and an attribute query type; determining the processing order of the plurality of data query conditions according to the query type to which the term corresponding to each of the data query conditions belongs and the number of triples associated with the term corresponding to each of the data query conditions in the target graph data.

[0164] Optionally, the data processing type includes a data query type corresponding to a select clause, and the data processing result for the data processing request is determined according to the data processing type and the data query results corresponding to the plurality of data query conditions, including: performing splicing processing on the data query results corresponding to the plurality of data query conditions to obtain a target query result; determining a data filtering condition corresponding to the select clause; performing filtering processing on the target query result according to the data filtering condition to obtain the data processing result for the data processing request.

[0165] Optionally, the data processing type includes a data deletion type corresponding to a delete clause, and the data processing result for the data processing request is determined according to the data processing type and the data query results corresponding to the plurality of data query conditions, including: performing splicing processing on the data query results corresponding to the plurality of data query conditions to obtain a target query result; deleting triples corresponding to the target query result in the target graph data.

[0166] Optionally, the data query results corresponding to the plurality of data query conditions are determined by sequentially performing query processing on each of the data query conditions in the target graph data according to the processing order of the plurality of data query conditions, including: determining a first query condition in the plurality of data query conditions according to the processing order of the plurality of data query conditions; performing query processing on the target data according to the first query condition to obtain a first data query result; updating a data query condition related to the first query condition among the data query conditions other than the first query condition according to the first data query result to obtain an updated data query condition; According to the number of triples associated with the updated data query condition in the target graph data, the processing order of the plurality of data query conditions is updated to obtain an updated processing order of the plurality of data query conditions. Based on the updated processing order of the plurality of data query conditions, the updated data query condition in the target graph data is sequentially queried to obtain a data query result corresponding to the plurality of data query conditions.

[0167] The embodiment of the present specification provides a data processing device, in response to a data processing request for target graph data, determining a data processing type corresponding to the data processing request and a plurality of data query conditions, wherein the target graph data contains more than a predetermined number of triples, determining the processing order of the plurality of data query conditions according to the number of triples associated with the term corresponding to each data query condition in the target graph data, sequentially querying each data query condition in the target graph data based on the processing order of the plurality of data query conditions, determining the data query result corresponding to the plurality of data query conditions, and determining the data processing result for the data processing request according to the data processing type and the data query result corresponding to the plurality of data query conditions. In this way, on the one hand, by determining the processing order of the plurality of data query conditions, the query processing is directly performed in the target graph data, which can avoid the high data storage and operation and maintenance costs and data inconsistency problems existing in the additional index for data query, on the other hand, since the number of triples associated with different search conditions may not be the same, in the case of large-scale graph data and multiple search conditions, the processing order of the plurality of data query conditions is determined according to the number of triples associated with the term corresponding to the data query condition in the target graph data, and the data query is sequentially performed according to the processing order, which can improve the data query efficiency, thereby improving the data processing efficiency for large-scale graph data.

[0168] Further, based on the above Figures 1 to 10 The one or more embodiments of the present specification also provide a storage medium for storing computer executable instruction information, in a specific embodiment, the storage medium can be a U disk, an optical disk, a hard disk, etc. The computer executable instruction information stored in the storage medium can implement the following processes when executed by a processor: In response to a data processing request for target graph data, determining a data processing type corresponding to the data processing request and a plurality of data query conditions, wherein the target graph data contains more than a predetermined number of triples; According to the number of triples associated with the term corresponding to each data query condition in the target graph data, the processing order of the plurality of data query conditions is determined; According to the processing sequence of the plurality of data query conditions, the data query conditions are sequentially queried in the target graph data to determine data query results corresponding to the plurality of data query conditions. According to the data processing type and the data query results corresponding to the plurality of data query conditions, a data processing result for the data processing request is determined.

[0169] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts of each of the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, for the above-mentioned storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0170] The embodiment of the specification provides a computer readable storage medium, in response to a data processing request for target graph data, determining a data processing type corresponding to the data processing request and a plurality of data query conditions, wherein the target graph data contains more than a predetermined number of triples, determining the processing sequence of the plurality of data query conditions according to the number of triples associated with each data query condition in the target graph data, based on the processing sequence of the plurality of data query conditions, sequentially querying each data query condition in the target graph data to determine the data query results corresponding to the plurality of data query conditions, and determining the data processing result for the data processing request according to the data processing type and the data query results corresponding to the plurality of data query conditions. In this way, on the one hand, by determining the processing sequence of the plurality of data query conditions, the query processing is directly performed in the target graph data, which can avoid the high data storage and operation and maintenance costs and data inconsistency problems caused by the need to establish additional indexes for data query, on the other hand, since the number of triples associated with different retrieval conditions may not be the same, in the case of large-scale graph data and multiple retrieval conditions, the processing sequence of the plurality of data query conditions is determined according to the number of triples associated with each data query condition in the target graph data, and the sequential query is performed according to the processing sequence, which can improve the data query efficiency, thereby improving the data processing efficiency for large-scale graph data.

[0171] Further, based on the above Figures 1 to 10 The method shown in the specification, one or more embodiments of the specification also provide a computer program product, including a computer program, the computer program in the computer program product can implement the following flow when executed by a processor: In response to a data processing request for target graph data, determine the data processing type corresponding to the data processing request and a plurality of data query conditions, the target graph data contains more than a predetermined number of triples. The processing order of the multiple data query conditions is determined based on the number of triples associated with the term corresponding to each data query condition in the target graph data. Based on the processing order of the multiple data query conditions, each of the data query conditions is sequentially queried and processed in the target graph data to determine the data query results corresponding to the multiple data query conditions; Based on the data processing type and the data query results corresponding to the multiple data query conditions, the data processing result for the data processing request is determined.

[0172] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the above-described embodiment of a computer program product is relatively simple in description because it is fundamentally similar to the method embodiment; relevant parts can be referred to the description of the method embodiment.

[0173] This specification provides a computer program product that, in response to a data processing request for target graph data, determines the data processing type and multiple data query conditions corresponding to the data processing request. The target graph data contains more than a predetermined number of triples. Based on the number of triples associated with each term in the target graph data, the processing order of the multiple data query conditions is determined. Based on this processing order, each data query condition is sequentially queried and processed in the target graph data to determine the data query results corresponding to the multiple data query conditions. Finally, based on the data processing type and the data query results corresponding to the multiple data query conditions, the data processing result for the data processing request is determined. In this way, on the one hand, by determining the processing order of multiple data query conditions, query processing can be performed directly in the target graph data, avoiding the high data storage and maintenance costs and data inconsistency issues that require additional indexing for data querying. On the other hand, since the number of triples associated with different search conditions may be different, when the graph data is large and there are many search conditions, determining the processing order of multiple data query conditions by the number of triples associated with the terms corresponding to the data query conditions in the target graph data, and then querying sequentially according to this processing order, can improve data query efficiency, thereby improving the data processing efficiency for large-scale graph data.

[0174] The above described embodiments of the present description have been described. Other embodiments are within the scope of the following claims. In some cases, the actions or steps recited in the claims can be performed in a different order and still achieve desirable results. Additionally, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.

[0175] In the 1990s, it was possible to distinguish whether an improvement in a technology was a hardware improvement (e.g., an improvement in the circuit structure of a diode, transistor, switch, etc.) or a software improvement (an improvement in a method flow). However, as technology has advanced, many improvements in method flows today can be considered as direct improvements in hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into a hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is an integrated circuit whose logic function is determined by user programming of the device. A digital system is "integrated" on a PLD by the designer programming it themselves, without having to ask a chip manufacturer to design and fabricate a custom integrated circuit chip. Moreover, instead of manually fabricating an integrated circuit chip, this programming is now mostly implemented using "logic compiler" software, which is similar to the software compiler used when developing a program, and the original code before compilation must also be written in a specific programming language, which is called a hardware description language (HDL), and there are many types of HDL, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc., and the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that it is very easy to obtain a hardware circuit that implements a logical method flow by simply logically programming the method flow in one of the above-mentioned hardware description languages and programming it into an integrated circuit.

[0176] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or processor and a computer readable medium storing computer readable program code, such as software or firmware, executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that, in addition to implementing the controller in pure computer readable program code, it is also possible to implement the controller in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc. to perform the same functions by logically programming the method steps. The controller can thus be considered a hardware component, and the means for performing the various functions included therein can also be considered structures within the hardware component. Alternatively, or even in addition, the means for performing the various functions can be considered both software modules that implement the methods and structures within the hardware component.

[0177] For ease of description, the above apparatuses are described in various units by function. Of course, the functions of the units can be implemented in one or more software and / or hardware when implementing one or more embodiments of the present specification.

[0178] Those skilled in the art will appreciate that embodiments of the present specification can be provided as methods, systems, or computer program products. Accordingly, one or more embodiments of the present specification can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of the present specification can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage, etc.) having computer usable program code embodied therein.

[0179] The embodiments of the present specification are described with reference to flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus generate a means for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 one or more flows and / or blocks

[0180] It should also be noted that the terms "comprising", "comprises", "including", "includes" or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article or apparatus. Without further limitation, an element preceded by "comprises a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.

[0181] Those skilled in the art will appreciate that the embodiments of the present specification can be provided as a method, system or computer program product. Therefore, one or more embodiments of the present specification can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of the present specification can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.

[0182] Each of the embodiments in the present specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the difference from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0183] The above only describes the embodiments of the present specification and is not intended to limit the present specification. The present specification can have various changes and modifications for those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of claims of the present specification.

Claims

1. A data processing method, comprising: determining a data processing type and a plurality of data query conditions corresponding to a data processing request in response to the data processing request, the target graph data containing more than a predetermined number of triples; determining a processing order of the plurality of data query conditions according to a number of triples associated with a term corresponding to each of the data query conditions in the target graph data; sequentially querying each of the data query conditions in the target graph data based on the processing order of the plurality of data query conditions to determine data query results corresponding to the plurality of data query conditions; determining a data processing result for the data processing request according to the data processing type and the data query results corresponding to the plurality of data query conditions.

2. The method of claim 1, wherein the determining a data processing type and a plurality of data query conditions corresponding to a data processing request comprises: obtaining a data processing statement corresponding to the data processing request; parsing the data processing statement based on a pre-constructed regular expression to obtain the data processing type and the plurality of data query conditions corresponding to the data processing request.

3. The method of claim 2, wherein the parsing the data processing statement based on a pre-constructed regular expression to obtain the data processing type and the plurality of data query conditions corresponding to the data processing request comprises: determining a first clause and a second clause corresponding to the data processing request based on the pre-constructed regular expression and a preset keyword, wherein the first clause is a select clause or a delete clause, and the second clause is a where clause; determining the data processing type corresponding to the data processing request according to the first clause; determining a plurality of conditional triples included in the second clause based on the pre-constructed regular expression and a preset triple format; determining the plurality of data query conditions corresponding to the data processing request according to the plurality of conditional triples.

4. The method of claim 3, wherein the determining a processing order of the plurality of data query conditions according to a number of triples associated with a term corresponding to each of the data query conditions in the target graph data comprises: determining a plurality of terms corresponding to each of the data query conditions according to a plurality of non-variable values in each of the data query conditions determined based on the pre-constructed regular expression; determining triples in the triples of the target graph data that have a preset association relationship with each term corresponding to the data query conditions; determining a number of triples corresponding to each of the data query conditions according to a number of triples associated with each term corresponding to each of the data query conditions; determining the processing order of the plurality of data query conditions according to the number of triples corresponding to each of the data query conditions.

5. The method of claim 4, wherein the determining triples in the triples of the target graph data that have a preset association relationship with each term corresponding to the data query conditions comprises: determine a target concept in the concept corresponding to the data query condition and related to the object value or the attribute value; determine a sub-database corresponding to the data query condition in an index database according to the target concept in the concept corresponding to the data query condition and the data type of the target concept, wherein the data type of the object value or the attribute value contained in the triple of the target graph data stored in the sub-database of the index database is the same as the data type of the target concept; determine, in each sub-database corresponding to the data query condition, a triple having the preset association relationship with the target concept.

6. The method of claim 1, wherein the determining of the processing order of the multiple data query conditions according to the number of triples associated with the concept corresponding to each data query condition in the target graph data comprises: determining a query type to which the concept corresponding to each data query condition belongs, the query type comprising an object query type and an attribute query type; determining the processing order of the multiple data query conditions according to the query type to which the concept corresponding to each data query condition belongs and the number of triples associated with the concept corresponding to each data query condition in the target graph data.

7. The method of claim 3, wherein the data processing type comprises a data query type corresponding to a select clause, and the determining of the data processing result for the data processing request according to the data processing type and the data query results corresponding to the multiple data query conditions comprises: performing splicing processing on the data query results corresponding to the multiple data query conditions to obtain a target query result; determining a data filtering condition corresponding to the select clause; performing filtering processing on the target query result according to the data filtering condition to obtain the data processing result for the data processing request.

8. The method of claim 3, wherein the data processing type comprises a data deletion type corresponding to a delete clause, and the determining of the data processing result for the data processing request according to the data processing type and the data query results corresponding to the multiple data query conditions comprises: performing splicing processing on the data query results corresponding to the multiple data query conditions to obtain a target query result; deleting a triple corresponding to the target query result in the target graph data.

9. The method of claim 1, wherein the determining of the data query results corresponding to the multiple data query conditions by sequentially performing query processing on each data query condition in the target graph data based on the processing order of the multiple data query conditions comprises: determining a first query condition in the multiple data query conditions according to the processing order of the multiple data query conditions; performing query processing on the target data according to the first query condition to obtain a first data query result; According to the first data query result, a data query condition related to the first query condition among the data query conditions other than the first query condition is updated to obtain an updated data query condition; According to the number of triples associated with the updated data query condition in the target graph data, the processing order of the plurality of data query conditions is updated to obtain an updated processing order of the plurality of data query conditions; Based on the updated processing order of the plurality of data query conditions, the updated data query condition in the target graph data is sequentially queried to obtain a data query result corresponding to the plurality of data query conditions.

10. A data processing device, characterized by A processor and a memory are included, the memory stores programs or instructions executable on the processor, and the programs or instructions are executed by the processor to implement the steps of the data processing method according to any one of claims 1 to 9.

11. A readable storage medium, characterized by, The readable storage medium stores programs or instructions, and the programs or instructions are executed by the processor to implement the steps of the data processing method according to any one of claims 1 to 9.

12. A computer program product, characterised in that, A computer program is included, and the computer program is executed by the processor to implement the steps of the data processing method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Search system based on regular path query

    CN113326284A

  • Data processing method and device, readable storage medium and electronic equipment

    CN115982416A

  • Data processing method and device and terminal equipment

    CN116467370A

  • Graph data query method for graph database and related equipment

    CN117591564A

  • Data processing method and apparatus, readable storage medium, and electronic device

    US20240256613A1