Corpus generation method and device based on graph query language

By building natural language and GQL templates containing placeholders, and using graph data patterns to define the generation type path, the problem of low efficiency in generating natural language and GQL conversion corpus in the existing technology is solved, and a large number of conversion corpus is achieved efficiently.

CN120407866APending Publication Date: 2025-08-01ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510458286.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently generate a large number of conversion corpus between natural languages and graph query languages (GQL), and manual collection and organization are inefficient.

Method used

By building natural language templates and GQL templates containing placeholders, use the schema definition and placeholder path of graph data to generate type paths, and replace placeholders in the template with the definition type of triple elements to generate a large number of conversion corpus.

Benefits of technology

It realizes efficient generation of the conversion corpus between natural language and GQL, improves the generation efficiency, and solves the problem of low manual collection and organization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407866A_ABST
    Figure CN120407866A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a corpus generation method and device based on a graph query language. In the method, a first corpus template used for constructing a conversion corpus of a graph query language is obtained, the first corpus template comprises a natural language template and a corresponding graph query language template, and the two templates comprise a first type of placeholders used for representing definition types of triple elements. Then, extracting a placeholder path formed by a first type of placeholders and a connection relationship thereof from the natural language template or the graph query language template, and generating a plurality of types of paths meeting relationship constraints based on the pattern definition and the placeholder path of the graph data, and using the definition type of the triple element in any type path to replace the corresponding first type placeholder in the natural language template and the graph query language template, and determining a first corpus based on a replacement result. Wherein the mode definition comprises definition types of triple elements and relation constraints among the triple elements. In the processing process, privacy protection needs to be carried out on related privacy data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the technical field of graph database, and in particular, to a method and device for generating corpus based on graph query language. Background Art

[0002] Database query language is a programming language used to interact with a database for query and analysis operations on the database or data warehouse. For example, Structured Query Language (SQL) is a database query and programming language for accessing data and querying, updating, and managing relational database systems. Similarly, Graph Query Language (GQL) is a standard developed by the International Organization for Standardization (ISO) for standardizing graph query languages and is used to execute queries in a graph database. In some complex query scenarios, it is relatively difficult to manually write the corresponding SQL statements, which requires a high level of developers. In recent years, with the development and popularization of large models, using large models to convert query statements described in natural language into corresponding GQL statements has gradually become an effective way to reduce the difficulty of writing and learning graph query languages. Training such large models requires a large amount of conversion corpus between natural language and GQL, and the generation process of this conversion corpus needs to be protected for privacy.

[0003] Currently, there is a need for an improved solution that can more efficiently generate a large amount of conversion corpus between natural language and GQL. Summary of the Invention

[0004] One or more embodiments of this specification describe a method and device for generating corpus based on graph query language to more efficiently generate a large amount of conversion corpus between natural language and GQL. The specific technical solutions are as follows.

[0005] In a first aspect, an embodiment provides a method for generating corpus based on graph query language, including:

[0006] Obtain a first corpus template for constructing a conversion corpus of graph query language, where the natural language template and graph query language template include a natural language template and a corresponding graph query language template, and the first corpus template contains a first type of placeholder for representing the defined type of triple elements;

[0007] Extract a placeholder path formed by the first type of placeholder and its connection relationship from the natural language template or graph query language template;

[0008] Generate several types of paths based on the schema definition of the graph data and the placeholder path; wherein, the schema definition includes the definition types of triple elements and the relationship constraints between them, and the connection relationship of the definition types of the triple elements included in any type of path satisfies the relationship constraints;

[0009] Replace the corresponding first type of placeholder in the natural language template and the graph query language template with the definition types of the triple elements in any one of the type paths, and determine the first corpus based on the replacement result.

[0010] In one implementation, the step of extracting the placeholder path formed by the first type of placeholder and its connection relationship from the natural language template or the graph query language template includes: extracting several first type of placeholders from the natural language template or the graph query language template in the word order; establishing a connection relationship for the several extracted first type of placeholders in the extraction order to obtain the placeholder path.

[0011] In one implementation, the step of establishing a connection relationship for the several extracted first type of placeholders in the extraction order includes: when there is a delimiter between the several first type of placeholders included in the natural language template or the graph query language template, grouping the several extracted first type of placeholders according to the position of the delimiter; for each group, establishing a connection relationship for the several first type of placeholders in the group in the extraction order to obtain the placeholder path corresponding to the group.

[0012] In one implementation, the step of generating several types of paths based on the schema definition of the graph data and the placeholder path includes: sequentially placing the definition types of the triple elements included in the schema definition at the positions corresponding to the first type of placeholders in the placeholder path to form several types of paths; wherein, the definition type of the triple element at the current position is selected from the screening set, and the definition types of the triple elements in the screening set are screened from the schema definition based on the existing definition types of the triple where the current position is located and the relationship constraints.

[0013] In one implementation, at least two placeholder paths are extracted from the natural language template or the graph query language template. The step of generating several types of paths based on the schema definition of the graph data and the placeholder path includes: when the same placeholder is included in the two placeholder paths, using the definition types of the triple elements at the position of the same placeholder in the several type paths corresponding to one of the generated placeholder paths as the definition types of the triple elements at the position of the same placeholder in the other placeholder path.

[0014] In one implementation, the natural language template and the graph query language template further include a second type of placeholder for defining types of non-triple elements, and the second type of placeholder is used to represent attributes of associated triple elements. The step of determining the first corpus based on the replacement result includes: for any second type of placeholder in the replacement result, generating instance data for the second type of placeholder based on the defined type and attributes of the associated triple element, and using the instance data to replace the second type of placeholder in the replacement result to obtain the first corpus.

[0015] In one implementation, the natural language template further includes a third type of placeholder for representing descriptive words; the step of determining the first corpus based on the replacement result further includes:

[0016] For any third type of placeholder in the replacement result, selecting a word from the optional words of different expression forms corresponding to the third type of placeholder stored in advance, and using the selected word to replace the third type of placeholder in the replacement result to obtain the first corpus.

[0017] In one implementation, the first corpus contains natural language instances and corresponding graph query language instances, and the method further includes:

[0018] Inputting the natural language instance into a large model to obtain a generalized natural language instance;

[0019] Adding the first corpus and the generalized natural language instance to the corpus set.

[0020] In one implementation, the corpus set is used to fine-tune a large model so that the large model can convert an input natural language instance into a corresponding graph query language instance, or so that the large model can convert a graph query language instance into a corresponding natural language instance.

[0021] In one implementation, the first corpus template is constructed based on the first syntactic function in the graph query language.

[0022] In a second aspect, an embodiment provides a corpus generation device based on a graph query language, including:

[0023] An acquisition module configured to acquire a first corpus template for constructing a conversion corpus of a graph query language, where the first corpus template includes a natural language template and a corresponding graph query language template, and the natural language template and the graph query language template contain a first type of placeholder for representing the defined type of triple elements;

[0024] An extraction module configured to extract a placeholder path formed by the first type of placeholder and its connection relationship from the natural language template or the graph query language template;

[0025] A generation module, configured to generate a number of type paths based on the schema definition of the graph data and the placeholder path; wherein, the schema definition includes the definition types of triple elements and the relationship constraints between them, and the connection relationship of the definition types of the triple elements included in any type path satisfies the relationship constraints;

[0026] A replacement module, configured to use the definition types of the triple elements in any one of the type paths to replace the corresponding first type of placeholders in the natural language template and the graph query language template, and determine the first corpus based on the replacement result.

[0027] In a third aspect, an embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method according to any one of the first aspects.

[0028] In a fourth aspect, an embodiment provides a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method according to any one of the first aspects is implemented.

[0029] In the method and apparatus provided in the embodiments of the present specification, by pre-constructing a natural language template containing placeholders and a corresponding GQL template, and using the schema definition and placeholder path of the graph data to generate a number of type paths containing the definition types of triple elements, and using the definition types of the triple elements in the type paths to replace the placeholders in the two templates, the required corpus is obtained. The number of type paths constructed in this way is very large, so many conversion corpora can be generated using the templates, and a large number of conversion corpora can be generated using multiple templates. Therefore, this embodiment can more efficiently generate a large number of conversion corpora between natural language and GQL. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0031] Figure 1 It is a schematic diagram of an implementation scenario of an embodiment disclosed in the present application;

[0032] Figure 2 It is a schematic flowchart of a method for generating a corpus based on a graph query language provided by an embodiment;

[0033] Figure 3A schematic diagram of the syntax tree corresponding to the GQL statement;

[0034] Figure 4 A schematic diagram of the syntax tree corresponding to the GQL instance;

[0035] Figure 5 A schematic structural diagram of the overall application architecture of this embodiment;

[0036] Figure 6 A schematic block diagram of a corpus generation device based on a graph query language provided by the embodiment. Detailed implementation manners

[0037] The following describes the solution provided in this specification with reference to the accompanying drawings.

[0038] Figure 1 A schematic diagram of the implementation scenario of an embodiment disclosed in this application. It includes a computing device, a GQL corpus, and a large model. The GQL corpus and the large model can be implemented in devices outside the computing device respectively, or can be implemented in the computing device. The GQL corpus template includes a natural language template and a corresponding GQL template. The computing device can generate a large number of GQL corpora according to the pre-constructed GQL corpus template and the schema definition of the graph data. The GQL corpus includes natural languages for querying subgraphs in the graph database and corresponding GQLs. The computing device can store the generated large number of GQL corpora in the GQL corpus. The corpora in the GQL corpus are used to train the large model so that the large model can convert the input natural language into the corresponding GQL, or convert the GQL into natural language.

[0039] GQL is a domain-specific language in the graph field and is a computer language. A domain-specific language (DSL) is a programming language or standard designed specifically for a certain specific application field. Natural language is a language that evolves naturally with culture and is the main tool for human communication and thinking, such as Chinese, English, etc. After converting natural language into GQL, subgraph queries can be performed on graph data in the graph database.

[0040] A large language model (LLM) is a language model, usually a machine learning model with tens of billions or more parameters and a complex structure, which has powerful semantic understanding capabilities and can understand and generate natural language text. Common large models include GPT, LLaMA, and GLM, etc. In the pre-training stage, self-supervised learning or semi-supervised learning is used to train on a large amount of text data, which can come from various sources such as the Internet, books, news, etc. The large model uses data from various application fields during pre-training, which contains a large amount of knowledge, so it has stronger language processing capabilities and wider applicability. Taking the pre-trained large model as the basic framework and fine-tuning the large model on the labeled dataset designed for specific tasks can enable the large model to quickly adapt to specific task processing, such as performing natural language processing or image recognition tasks.

[0041] Corpus usually refers to the text dataset used for research and analysis. The GQL corpus is a dataset that contains natural language text and corresponding GQL language text. The GQL corpus can be used as a labeled dataset to train the large model. That is, the GQL corpus contains a large number of data pairs formed by natural language - GQL. Currently, in the field of graph databases, such GQL corpora are still scarce and the datasets are relatively lacking, and it is also difficult and inefficient to collect and organize the corpus manually.

[0042] In order to efficiently build a corpus generation system that can automatically generate more natural language - GQL corpora, the embodiments of this specification provide a corpus generation method based on graph query language. The following combines Figure 2 to elaborate on the specific embodiments in detail.

[0043] Figure 2 It is a schematic flowchart of a corpus generation method based on graph query language provided for the embodiment. This method is executed by a computing device, and the computing device can be implemented by any device, equipment, platform, device cluster, etc. with computing and processing capabilities. This method includes the following steps.

[0044] Step S210, obtain the first corpus template P1 for constructing the conversion corpus of the graph query language.

[0045] Among them, the first corpus template P1 is any corpus template in the template library. The following description of the first corpus template P1 also applies to any corpus template in the template library. The first corpus template P1 includes a natural language template and a corresponding GQL template (i.e., the GQL language template), and the function of each template is to query graph data. The first corpus template P1 is used to construct a conversion corpus for graph query languages, that is, a corpus for conversion between natural language and GQL. This conversion corpus can be used for training large models, but is not limited to this.

[0046] The natural language template and the GQL template correspond, or rather, a certain natural language statement corresponds to a certain GQL statement, which means that both represent the same query content, but the difference is that they are expressed in different languages. The first corpus template P1, including the natural language template and the corresponding GQL template, is constructed through placeholders, which contain several placeholders. For example, Table 1 shows the specific content of a certain corpus template. Table 2 is any corpus generated based on this corpus template.

[0047] Table 1

[0048]

[0049] Table 2

[0050]

[0051]

[0052] Among them, the data after # and $ are all placeholders. In Table 1 and Table 2, prompt (hint) is used to represent natural language, and answer (answer) is used to represent GQL. In this case, after generating the corpus, the natural language in it can be constructed into a prompt and input into the large model, and the corresponding GQL obtained through the large model is used as the answer. This is just one implementation method. In actual applications, the GQL can also be constructed into a prompt and input into the large model, and the corresponding natural language obtained through the large model is used as the answer.

[0053] The first corpus template P1 is pre-constructed manually. The following is a detailed description of its construction process.

[0054] The natural language template and the corresponding GQL template contain the first type of placeholder labelVar for representing the defined types of triple elements. The triple elements include any one or more of nodes and edges, specifically including the source point, edge, and target point. Specifically, the natural language template and the corresponding GQL template may only contain the first type of placeholder for representing the defined type of one node, or may contain the first type of placeholder for representing the defined types of several nodes and the first type of placeholder for representing the defined types of several edges. Several includes one or more.

[0055] The defined type of triple elements refers to the defined types of triple elements that can exist in the graph data as defined in the schema definition of the graph data, including node types and edge types. The schema definition of the graph data also defines the attributes that can exist in the graph data and the relationship constraints between nodes. Graph data in different application scenarios has different schema definitions. Attributes include node attributes and edge attributes. Graph data includes nodes and the edges between nodes, and the table reflects the association relationships between nodes.

[0056] Node types define the types of nodes that can exist in the graph data, such as person and software, etc. Edge types define the types of relationships that can exist between nodes, such as knows, creates, and reference (abbreviated as ref), etc. Attributes define the attributes that nodes and edges can have and their data types, such as name (string), age (integer), and weight (floating point), etc.

[0057] Relationship constraints define which types of triples formed by source nodes, edges, and target nodes can be. For example, triples can be person-knows-person, person-creates-software, software-ref-software, etc., but cannot be software-creates-person, etc.

[0058] The schema definition can also include other constraints, such as a certain node type must have a certain attribute, or the value of a certain attribute must satisfy a certain range, etc. Defined types (including node types and edge types) and attributes are defined conceptually and do not have specific data.

[0059] The first corpus template P1 contains several first-class placeholder labelVars, and each first-class placeholder indicates the corresponding triple element. For example, in the prompt template and answer template of Table 1, {LabelVar-V-a}, {LabelVar-V-e}, and {LabelVar-V-b} represent the defined type placeholder of the source node V-a, the defined type placeholder of the edge V-e, and the defined type placeholder of the target node V-b, respectively.

[0060] To make the expression form of the corpus more abundant, the natural language template and the corresponding GQL template may also include a second type of placeholder for the definition type of non-triple elements, that is, the second type of placeholder is not a placeholder for the definition type of triple elements. The second type of placeholder is used to represent the attributes of the associated triple elements, that is, each second type of placeholder has an associated triple element, that is, the second type of placeholder is a placeholder for the attributes of a certain triple element.

[0061] For example, in the prompt template and answer template of Table 1, {NameVar-V-a} is a placeholder for the name attribute of the source node V-a, and {ProjectVar-V-b} is an attribute expression related to the target node V-b.

[0062] For example, the second type of placeholder may include a name placeholder (nameVar), a numerical placeholder (numberVar), a filtering condition placeholder (conditionVar), a variable expression placeholder (ProjectVar), and so on. Among them, the name placeholder is used to indicate the extraction of the name of a node or an edge. The numerical placeholder is used to indicate the generation of a numerical value. For example, limit${NumberVar} can indicate the generation of limit 10, which means limited to the numerical value 10. The filtering condition placeholder is used to indicate the generation of a filtering condition. For example, for match(a where${ConditionVar-a}), filtering conditions such as match(a where a.id=1), match(a where a.id!=1 and a.length>10) can be generated. The variable expression placeholder is used to indicate the generation of a variable expression, which can be used when fetching attributes. For example, return a.id and order by b.high can be generated, which respectively represent returning the identifier id of node a and sorting by the high value of node b.

[0063] Although the GQL template in Table 1 is in the form of a statement, in the graph query language, this statement can correspond to a syntax tree. Figure 3It is a schematic diagram of the syntax tree corresponding to the GQL statement. Among them, the root node is the statement corresponding to the GQL template, and this root node can be split into a match part and a return part. This match part can be further split into child nodes represented by source node a, edge e, and target node b. Source node a can be split into a placeholder {LabelVar-V-a} for the defined type of the source node and a where part, where the where part contains a placeholder {NameVar-V-a} for the name. The child node of edge e contains a placeholder {LabelVar-V-e} for the defined type of the edge, and the child node of target node b contains a placeholder {LabelVar-V-b} for the defined type of the target node. The return part contains a placeholder {ProjectVar-V-b} for the variable expression. The root node of this syntax tree is the GQL statement, and the leaf nodes are various types of placeholders included therein.

[0064] The natural language template can also include a third type of placeholder for representing descriptive words (description). Descriptive words can include, for example, words such as query, all, find, and return, and the Desc identifier can be used in the placeholder. The third type of placeholder is also a placeholder for ordinary descriptions.

[0065] The second type of placeholder and the third type of placeholder are not necessarily present in a corpus template and are absent in some corpus templates. In this embodiment, an example is given where the first corpus template P1 contains several first type of placeholders and does not contain the second type of placeholder and the third type of placeholder. Several includes the case of one or more.

[0066] In order to cover common or most of the grammar points in GQL, when constructing a corpus template, it can be constructed separately based on multiple grammar functions included in GQL. For example, the first corpus template P1 can be constructed based on the first grammar function in the graph query language.

[0067] The first grammar function is any one of the multiple grammar functions in the graph query language. Each grammar function corresponds to a graph query type. For example, the grammar functions include match point, match edge, where, etc. Grammar functions are also called grammar points. The grammar functions in the graph query language are the basic building blocks of this language. By combining these grammar functions, users can construct complex query statements for querying, filtering, aggregating, and modifying graph data, etc.

[0068] Table 3 is a corpus example constructed corresponding to the common grammar functions in the graph query language. For the convenience of understanding, corpus examples are used in Table 3 to reflect the corpus template. By correspondingly replacing the elements in any corpus example with placeholders, the corpus template corresponding to the grammar function can be obtained.

[0069] Table 3

[0070]

[0071]

[0072]

[0073] The left column in Table 3 gives common syntactic functions in the graph query language, and the right column gives corpus examples corresponding to the syntactic functions. The corresponding corpus templates can be obtained according to the corpus examples. One or more corpus templates can be constructed for each syntactic function. Each corpus template is a query operation. Constructing corpus templates corresponding to multiple query operations can basically cover most graph data query situations.

[0074] The process of constructing the corpus template can be executed by a computing device that is the execution subject of this method, or by other devices outside the computing device. The computing device can obtain the constructed corpus template from the other device.

[0075] Step S220, extract a placeholder path formed by the first type of placeholder labelVar and its connection relationship from the first corpus template P1. Specifically, the placeholder path can be extracted from the natural language template or the corresponding GQL template.

[0076] Among them, the number of placeholder paths can be one or more. For each placeholder path, it contains a path connected by several first type of placeholders labelVar.

[0077] For example, the placeholder path extracted from the natural language template or GQL template in Table 1 is: {LabelVar-V-a}—{LabelVar-V-e}—{LabelVar-V-b}. This placeholder path contains the placeholder {LabelVar-V-a} of the source node V-a, the placeholder {LabelVar-V-e} of the edge V-e, and the placeholder {LabelVar-V-b} of the target node V-b.

[0078] The natural language template and the GQL template contain the same first type of placeholder labelVar and its order. The first type of placeholder labelVar generally appears in the order of node, edge, node, edge, and the first first type of placeholder is usually a node. In specific implementation, the placeholder path can be extracted according to the following processes of Step 1 and Step 2.

[0079] Step 1: Extract a number of first-class placeholder labelVars from the first corpus template P1 in the order of word order. Specifically, extract a number of first-class placeholder labelVars from the natural language template or GQL template in the order of word order.

[0080] The number of first-class placeholders extracted in this step can be understood as independent and uncorrelated placeholders. The extracted number of first-class placeholders can be stored in a list.

[0081] Step 2: Establish a connection relationship for the extracted number of first-class placeholder labelVars in the extraction order to obtain a placeholder path.

[0082] Among them, establishing a connection relationship can be understood as associating adjacent first-class placeholder labelVars in the order of node-edge-node, establishing a connection relationship between nodes and edges to form a doubly linked list. That is, for each placeholder, set the prev and next attributes of the placeholder to the previous placeholder and the next placeholder.

[0083] When the first corpus template P1 is a relatively complex template, there may be delimiters between multiple first-class placeholder labelVars, so multiple groups of placeholder paths can be generated.

[0084] Therefore, in Step 2, when there are delimiters between the number of first-class placeholder labelVars included in the natural language template or GQL template, group the extracted number of first-class placeholder labelVars according to the positions of the delimiters. For each group, establish a connection relationship for the number of first-class placeholder labelVars in the group in the extraction order to obtain the placeholder path corresponding to the group, so as to obtain multiple groups of corresponding placeholder paths.

[0085] For example, in Table 3, the GQL template corresponding to the syntax function join is as follows: match(a:${LabelVar-a}wherea.name='${NameVar-a}')-[e:${LabelVar-e}]-(b:${LabelVar-b)},(a:${LabelVar-a)}-[e2:${LabelVar-e2}]->(c:${LabelVar-c})return a,b,c.

[0086] The first type of placeholders included therein are as follows: {LabelVar-a}{LabelVar-e}{LabelVar-b)}, {LabelVar-a)}{LabelVar-e2}{LabelVar-c}. These placeholders are separated by the delimiter ",", so two groups of the first type of placeholders labelVar can be extracted from this corpus template, which are the two parts separated by commas. Connect the first type of placeholders labelVar in these two parts in sequence to obtain the following two paths:

[0087] {LabelVar-a}—{LabelVar-e}—{LabelVar-b)}

[0088] {LabelVar-a)}—{LabelVar-e2}—{LabelVar-c}(1)

[0089] Among them, the above two placeholder paths are both path relationships including node placeholder - edge placeholder - node placeholder, and the source node placeholders in the two placeholder paths are the same. In practical applications, the placeholder path also includes other situations. For example, the placeholder path may only contain one node placeholder, or may include a path relationship such as node placeholder - edge placeholder - node placeholder - edge placeholder - node placeholder (see the corpus example corresponding to the syntax function let (subquery) in Table 3), or may also include more placeholders.

[0090] Step S230, generate several types of paths based on the schema definition of the graph data and the placeholder path.

[0091] As mentioned above, the schema definition includes the definition types of triple elements and the relationship constraints between them. Each first type of placeholder labelVar included in any placeholder path corresponds to a position on this placeholder path. Combine the definition types of nodes and edges in the schema definition according to the position order of the first type of placeholder labelVar in the placeholder path, and the corresponding type path can be obtained. Any type path includes the definition types of several triple elements and their connection relationships, and this connection relationship satisfies the relationship constraints in the schema definition.

[0092] Specifically, when generating the type path, the definition types of the triple elements included in the schema definition can be placed in sequence at the positions corresponding to the first type of placeholder labelVar in the placeholder path to form several type paths.

[0093] Among them, the defined type of the triple element at the current position is selected from the filtering set schemalist. The defined types of the triple elements in the filtering set are filtered from the defined types of the triple elements included in the schema definition based on the existing defined types and relationship constraints of the triple where the current position is located. The triple where the current position is located means that if any one of the source node, edge, and target node in the triple is placed at the current position, then the triple where this element is located is the triple where the current position is located. When determining the defined type of each position in the type path, the filtering set schemalist should correspond to this position, that is, the filtering set schemalist needs to be re-determined at each position so as to select the defined type from the filtering set schemalist that meets the conditions. The first type of placeholder labelVar includes node placeholders and edge placeholders. A node placeholder is a placeholder representing the defined type of a node, and an edge placeholder is a placeholder representing the defined type of an edge.

[0094] When the current position is the first placeholder position, if the first type of placeholder labelVar at this position is a node placeholder, then the node defined types in the schema definition that can be used as the source node are added to the filtering set schemalist.

[0095] When the current position is not the first placeholder position, the defined types are filtered from the schema definition based on the existing defined types of the triple where the current position is located and the above relationship constraints, and are added to the filtering set schemalist.

[0096] When the first type of placeholder labelVar at the previous position is a node placeholder and the first type of placeholder labelVar at the current position is an edge placeholder, the defined types of the edges that meet the relationship constraints in the schema definition can be added to the filtering set schemalist. That is to say, the node defined type at the previous position and the edge defined type in this filtering set schemalist should meet the relationship constraints.

[0097] When the first type of placeholder labelVar at the previous position is an edge placeholder and the first type of placeholder labelVar at the current position is a node placeholder, the defined types that can be used as the target node and meet the relationship constraints in the schema definition can be added to the filtering set schemalist. That is to say, the node defined type before the previous position, the edge defined type at the previous position, and the node defined type in this filtering set should meet the relationship constraints.

[0098] Based on the defined types determined by the previous position and the filtering set schemalist corresponding to the current position, multiple type paths can be generated. The following is an example. Suppose the schema definition includes four node definition types: a1, a2, b1, and b2, and two edge definition types: e1 and e2. The relationship constraints include: a1-e1-b2, a2-e2-b1, b2-e2-a1. The node types that can be source nodes include a1, a2, and b2, and the node types that can be target nodes include a1, b1, and b2.

[0099] Each position in the placeholder path is 1-2-3-4-5, where positions 2 and 4 are edge placeholders, and 1, 3, and 5 are node placeholders.

[0100] When determining the type path, position 1 can be selected from the filtering set composed of the node types a1, a2, and b2 that can be source nodes. When position 1 is the defined type a1, according to the relationship constraints, the defined type of position 2 can be selected from the filtering set composed of e1; when position 1 is the defined type a2, according to the relationship constraints, the defined type of position 2 can be selected from the filtering set composed of e2; when position 1 is the defined type b2, according to the relationship constraints, the defined type of position 2 can be selected from the filtering set composed of e2. Thus, the incomplete type paths obtained are:

[0101] a1-e1…, a2-e2…, b2-e2…(2)

[0102] Among them, the position 2 in formula (2) corresponds to the defined type of the edge. Therefore, when determining the defined type of the edge at position 3, it is necessary to judge according to the triple where position 3 is located, that is, the defined types of position 1 and position 2 and the relationship constraints. When position 1 is the defined type a1 and position 2 is the defined type e1, according to the relationship constraints, the defined type of position 3 can be selected from the filtering set composed of b2; when position 1 is the defined type a2 and position 2 is the defined type e2, according to the relationship constraints, the defined type of position 3 can be selected from the filtering set composed of b1; when position 1 is the defined type b2 and position 2 is the defined type e2, according to the relationship constraints, the defined type of position 3 can be selected from the filtering set composed of a1. Thus, the above type paths can be updated to the following incomplete type paths:

[0103] a1-e1-b2…, a2-e2-b1…, b2-e2-a1…(3)

[0104] Among them, the one corresponding to position 3 in formula (3) is the definition type b2 of the node. Therefore, when determining the definition type of the edge at position 4, it is only necessary to judge according to the definition type and relationship constraints at position 3, without judging according to the definition type of the edge at position 2. When the definition type at position 3 is b2, combined with the relationship constraints, it can be known that the definition type at position 4 can be selected from the screening set composed of e2; when the definition type at position 3 is b1, combined with the relationship constraints, it can be known that there is no selectable edge for the definition type at position 4, that is, the screening set is empty at this time; when the definition type at position 3 is a1, combined with the relationship constraints, it can be known that the definition type at position 4 can be selected from the screening set composed of e1. Thus, the above type path can be updated to the following incomplete type path:

[0105] a1-e1-b2-e2…, a2-e2-b1-×…, b2-e2-a1-e1…(4)

[0106] Among them, in formula (4), the paths with position 4 being × are discarded, and the remaining 2 paths. When the definition type at position 3 in the above type path is b2 and the definition type at position 4 is e2, combined with the relationship constraints, it can be known that the definition type at position 5 can be selected from the screening set composed of a1; when the definition type at position 3 in the above type path is a1 and the definition type at position 4 is e1, combined with the relationship constraints, it can be known that the definition type at position 5 can be selected from the screening set composed of b2. Thus, the above type path can be updated to the following type path:

[0107] a1-e1-b2-e2-a1, b2-e2-a1-e1-b2(5)

[0108] The 2 type paths in the above formula (5) are the type paths obtained in this example.

[0109] In this step, based on the definition type determined by the previous position and the screening set schemalist corresponding to the current position, the backtracking method can be used to generate the above multiple type paths.

[0110] The above is the process description of generating the corresponding type path for a placeholder path. When extracting at least two placeholder paths from the first corpus template P1, if the same placeholder is included in these two placeholder paths, when generating several type paths, the definition types of the triple elements at the position of the same placeholder in the several type paths corresponding to one placeholder path that has been generated can be used as the definition types of the triple elements at the position of the same placeholder in the other placeholder path.

[0111] For example, the two placeholder paths included in formula (1) are extracted from a corpus template, and the two placeholder paths contain the definition type placeholder {LabelVar-a} of the same source node. If multiple type paths k1 to k7 have been generated for the first placeholder path in formula (1), then when generating the type path kx for the second placeholder path in formula (1), the definition type at the position of the placeholder {LabelVar-a} in the multiple type paths k1 to k7 can be used as the definition type at the same position in the type path kx.

[0112] This embodiment can avoid repeatedly determining the definition type, save the operation process, and improve the processing efficiency.

[0113] Step S240: Replace the first type of placeholder corresponding in the first corpus template P1 with the definition type of the triple element in any one of the type paths, and determine the first corpus based on the replacement result. The first corpus includes natural language instances and corresponding GQL instances, and one first corpus is a sample.

[0114] When replacing, the replacement is performed separately for the natural language template and the GQL template included in the first corpus template P1. Specifically, the replacement is performed on both the match matching part and the return return part in the GQL template. When the first corpus template P1 only includes the first type of placeholder, the replacement result can be directly used as the first corpus.

[0115] When replacing the GQL template, the leaf nodes included in the syntax tree corresponding to the GQL statement can be used to replace the leaf nodes with specific definition types, and thus the GQL instance is obtained.

[0116] Taking the GQL template in Table 1 as an example, the GQL syntax tree corresponding to this GQL template is as Figure 3 shown. By replacing the leaf nodes therein with the specific definition types in the type path, the syntax tree corresponding to the GQL instance can be obtained. Figure 4 It is a schematic diagram of the syntax tree corresponding to the GQL instance. The root node of this syntax tree is the GQL instance in Table 2, and the leaf nodes respectively include person, a.name = 'Zhang San', teach, course, and b.name. It can be seen that Figure 3 selecting from any non-leaf node in the syntax tree can generate a corresponding GQL sub-statement.

[0117] When there are n type paths, each type path can correspondingly obtain a first corpus, thereby obtaining k first corpora.

[0118] In this embodiment, the implementation manner in which the placeholders in the first corpus template P1 only include the first type of placeholder labelVar is described. For example, in Table 3, the corpus targets corresponding to the syntax functions match point and match edge only include the case of the first type of placeholder labelVar. Through the method of this embodiment, multiple type paths can be generated based on the schema definition and the placeholder path, so as to obtain multiple corpora. When different corpus templates are constructed, a large number of corpora can be generated. Compared with the number of constructed corpus templates, the number of obtained corpora has changed by an order of magnitude. Therefore, a large number of conversion corpora between natural language and GQL can be generated more efficiently.

[0119] In another embodiment of the present application, as mentioned in step S210, the first corpus template P1 may further include a second type of placeholder for defining the type of non-triple elements, and the second type of placeholder is used to represent the attributes of the associated triple elements. For a detailed description of the second type of placeholder, reference can be made to the relevant description in step S210, which will not be elaborated here.

[0120] When determining the first corpus based on the replacement result in step S240, for any second type of placeholder in the replacement result corresponding to the natural language template, based on the defined type and attributes of the associated triple elements, instance data for this second type of placeholder can be generated, and this instance data is used to replace the second type of placeholder in this replacement result to obtain the natural language instance in the first corpus. For any second type of placeholder in the replacement result corresponding to the GQL template, based on the defined type and attributes of the associated triple elements, instance data for this second type of placeholder can be generated, and this instance data is used to replace the second type of placeholder in this replacement result to obtain the GQL instance in the first corpus. When different instance data is generated, after using the corresponding instance data to replace the second type of placeholder in this replacement result, multiple first corpora can be obtained.

[0121] Among them, the attributes of the defined type can be obtained from the schema definition. When generating the instance data for this second type of placeholder, the instance data can be randomly generated based on the defined type and attributes of the associated triple elements.

[0122] For example, when the first corpus template P1 includes the name placeholder {NameVar-V-a}, the definition type associated with this name placeholder is the definition type of node V-a, and the associated attribute is the name. Then, instance data for the name placeholder {NameVar-V-a} can be generated based on the already determined definition type and its attribute of node V-a. Suppose the definition type of node V-a is person. Then, according to the definition type person and the attribute being the name, a person's name such as Zhang San or Li Si can be randomly generated. For example, based on the natural language template and GQL template in Table 1, the instance data of the definition types of nodes and edges and the relevant attributes can be correspondingly filled into the placeholders to generate the natural language instances and GQL instances in Table 2.

[0123] When the first corpus template P1 includes the filtering condition placeholder {conditionVar-V-a}, if the definition type associated with this filtering condition placeholder is software and the associated attribute is length, then a length condition for the software can be randomly generated, such as length > 50, or length < 100, etc.

[0124] In this embodiment, when the first corpus template P1 contains the second type of placeholder, i.e., the attribute placeholder, a large number of different corpora can be constructed by randomly generating the instance data of this second type of placeholder, improving the richness of the corpora.

[0125] In another embodiment of the present application, when the first corpus template P1 does not include the first type of placeholder but only includes the second type of placeholder, the instance data of the second type of placeholder can be directly generated based on the information associated with the second type of placeholder. For example, the functionCall syntax function in Table 3.

[0126] In another embodiment of the present application, as mentioned in step S210, the natural language template can further include a third type of placeholder for representing descriptive words. Descriptive words can include, for example, words such as query, all, find, and return. The Desc identifier can be used in the placeholder. The third type of placeholder is set to improve the richness of the corpora by using various expression forms of descriptive words.

[0127] When determining the first corpus based on the replacement result in step S240, for any third - type placeholder in the replacement result (including the replacement results corresponding to the natural - language template and the GQL template respectively), a word can be selected from the optional words with different forms of expression corresponding to this third - type placeholder stored in advance, and the selected word is used to replace the third - type placeholder in the replacement result to obtain the first corpus. When different optional words are selected, different first corpora can be obtained, so that multiple first corpora can be generated by changing the form of expression of descriptive words, increasing the richness of the corpus.

[0128] For example, Table 1 contains the third - type placeholder {SelectDesc}, and the optional words corresponding to this placeholder include "select", "find", and "search", etc. Correspondingly, when this placeholder is replaced with different optional words, the natural language in Table 2 can be transformed into the following expressions:

[0129] Select all the courses taught by the nodes named Zhang San;

[0130] Find all the courses taught by the nodes named Zhang San.

[0131] In another embodiment of the present application, after obtaining the first corpus, the natural - language instances in the first corpus can also be generalized to make their forms of expression more diverse and natural. The natural - language instances generated by code may be more templated. Through the conversion of the large - model, the templated language can be converted into a language that more conforms to the natural expression rules. For example, the natural - language instance can be: Find the people who have a friendship relationship with the point named Xiaohong. The generalized expression obtained after being transformed by the large - model is: Find Xiaohong's friends. The meaning of generalization also includes that the large - model converts the natural - language instance into other descriptive forms, thus making the expression more diverse.

[0132] Specifically, the embodiment can also input the natural - language instances in the first corpus into the large - model LLM1 to obtain the generalized natural - language instances, and add the first corpus and the generalized natural - language instances to the corpus set. The first corpus contains the original natural - language instances and the corresponding GQL instances.

[0133] Among them, the above - mentioned corpus set can be used to fine - tune the large - model LLM2 so that the large - model LLM2 can convert the input natural - language instance (i.e., prompt) into the corresponding GQL instance (i.e., answer). Or, this corpus set can be used to fine - tune the large - model LLM3 so that the large - model LLM3 can convert the GQL instance (i.e., prompt) into the corresponding natural - language instance (i.e., answer).

[0134] Each sample in the corpus contains a natural language instance and a corresponding GQL instance. When the corpus is used to train the large model LLM2, the GQL instances in the samples can be used as labels. When the corpus is used to train the large model LLM3, the natural language instances in the samples can be used as labels. For specific details, please refer to Figure 5 。

[0135] Figure 5 FIG. is a schematic structural diagram of an overall application architecture of this embodiment. Among them, the generator generates natural language (NL) instances and GQL instances according to the method of this embodiment. Inputting the NL instances into the large model LLM1, generalized NL can be obtained, and the generalized NL is also added to the corpus. In this way, an NL data set containing NL instances and generalized NL, as well as a GQL instance data set, can be obtained. The large models LLM2 and LLM3 can be fine-tuned using the NL data set and the GQL instance data set respectively. After fine-tuning, the large model LLM2 can be used to interact with users and provide GQL conversion services; the large model LLM3 can provide prediction services, that is, it can predict the corresponding natural language according to the input GQL instance.

[0136] The automated generation method provided in the above embodiment can effectively generate a large amount of natural language-GQL corpus, and can further generalize the description of natural language through the capabilities of large models to obtain a more diverse corpus data set, solving the problems of low efficiency in manual collection and collation and the steps of manually extracting the initial data set. The large models after fine-tuning can provide interactions such as intelligent map queries or graph calculations.

[0137] In this specification, the "first" in terms such as the first type of placeholder and the first corpus, and the corresponding "second" (if any) in the text are only for the convenience of distinction and description, and do not have any restrictive meaning.

[0138] The above content describes specific embodiments of this specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments, and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily have to be performed in the specific order or continuous order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible, or may be advantageous.

[0139] Figure 6 FIG. is a schematic block diagram of a corpus generation device based on a graph query language provided for the embodiment. This device embodiment corresponds to Figure 2 the method embodiment shown. The device 400 is deployed in a computing device and includes:

[0140] An acquisition module 610, configured to acquire a first corpus template for constructing a conversion corpus of a graph query language, where the first corpus template includes a natural language template and a corresponding graph query language template, and the natural language template and the graph query language template contain first-class placeholders for representing defined types of triple elements;

[0141] An extraction module 620, configured to extract a placeholder path formed by the first-class placeholders and their connection relationships from the natural language template or the graph query language template;

[0142] A generation module 630, configured to generate several types of paths based on the schema definition of the graph data and the placeholder path; wherein, the schema definition includes defined types of triple elements and relationship constraints therebetween, and the connection relationships of the defined types of triple elements included in any type of path satisfy the relationship constraints;

[0143] A replacement module 640, configured to use the defined types of triple elements in any one of the type paths to replace the corresponding first-class placeholders in the natural language template and the graph query language template, and determine the first corpus based on the replacement result.

[0144] In one implementation, the extraction module 620 includes:

[0145] An extraction sub-module 21, configured to extract several first-class placeholders from the natural language template or the graph query language template in the word order;

[0146] A connection sub-module 22, configured to establish a connection relationship for the several first-class placeholders extracted in the extraction order to obtain a placeholder path.

[0147] In one implementation, the connection sub-module 22 is specifically configured to:

[0148] When there are delimiters between several first-class placeholders included in the natural language template or the graph query language template, group the several first-class placeholders extracted according to the positions of the delimiters;

[0149] For each group, establish a connection relationship for the several first-class placeholders in the group in the extraction order to obtain a placeholder path corresponding to the group.

[0150] In one implementation, the generation module 630 is specifically configured to:

[0151] Place the definition types of the triple elements included in the pattern definition in sequence at the positions corresponding to the first type of placeholders in the placeholder path to form several type paths; wherein, the definition type of the triple element at the current position is selected from the screening set, and the definition types of the triple elements in the screening set are screened from the pattern definition based on the existing definition types of the triple where the current position is located and the relationship constraint.

[0152] In one implementation, at least two placeholder paths are extracted from a natural language template or a graph query language template. The specific configuration of the generation module 630 includes:

[0153] When the two placeholder paths contain the same placeholder, use the definition type of the triple element at the position of the same placeholder in the several type paths corresponding to one generated placeholder path as the definition type of the triple element at the position of the same placeholder in the other placeholder path.

[0154] In one implementation, the natural language template or the graph query language template further includes a second type of placeholder for the definition type of non-triple elements, and the second type of placeholder is used to represent the attributes of the associated triple elements. When the replacement module 640 determines the first corpus based on the replacement result, it includes:

[0155] For any second type of placeholder in the replacement result, generate instance data for the second type of placeholder based on the definition type and attributes of the associated triple element, and use the instance data to replace the second type of placeholder in the replacement result to obtain the first corpus.

[0156] In one implementation, the natural language template further includes a third type of placeholder for representing descriptive words. When the replacement module 640 determines the first corpus based on the replacement result, it further includes:

[0157] For any third type of placeholder in the replacement result, select a word from the optional words of different expression forms corresponding to the third type of placeholder stored in advance, and use the selected word to replace the third type of placeholder in the replacement result to obtain the first corpus.

[0158] In one implementation, the first corpus contains natural language instances and corresponding graph query language instances, and the apparatus 600 further includes:

[0159] The generalization module 650 is configured to input the natural language instance into a large model to obtain a generalized natural language instance;

[0160] The addition module 660 is configured to add the first corpus and the generalized natural language instance to the corpus set.

[0161] In one implementation, the corpus is used to fine-tune a large model so that the large model can convert an input natural language instance into a corresponding graph query language instance, or so that the large model can convert a graph query language instance into a corresponding natural language instance.

[0162] In one implementation, the first corpus template is constructed based on the first syntactic function in the graph query language.

[0163] The above device embodiments correspond to the method embodiments. For specific descriptions, reference can be made to the descriptions in the method embodiment section, which will not be elaborated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For specific descriptions, reference can be made to the corresponding method embodiments.

[0164] The embodiments of this specification also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute Figures 1 to 5 any of the methods described above.

[0165] The embodiments of this specification also provide a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the following is implemented Figures 1 to 5 any of the methods described above.

[0166] The embodiments in this specification are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the storage medium and the computing device, since they are basically similar to the method embodiments, the descriptions are relatively simple. For the relevant parts, reference can be made to the partial descriptions of the method embodiments.

[0167] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the embodiments of the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0168] The specific embodiments described above further elaborate the objectives, technical solutions, and beneficial effects of the embodiments of the present invention. It should be understood that the above are only the specific embodiments of the embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention should be included in the protection scope of the present invention.

Claims

1. A corpus generation method based on a graph query language, comprising: Obtaining a first corpus template for constructing a conversion corpus of a graph query language, where the first corpus template includes a natural language template and a corresponding graph query language template, and the natural language template and the graph query language template contain first-class placeholders for representing the defined types of triple elements; Extracting a placeholder path formed by the first-class placeholders and their connection relationships from the natural language template or the graph query language template; Generating a number of type paths based on the schema definition of the graph data and the placeholder path; wherein, the schema definition includes the defined types of triple elements and the relationship constraints between them, and the connection relationships of the defined types of triple elements included in any type path satisfy the relationship constraints; Replacing the corresponding first-class placeholders in the natural language template and the graph query language template with the defined types of triple elements in any one of the type paths, and determining the first corpus based on the replacement result.

2. The method according to claim 1, wherein the step of extracting a placeholder path formed by the first-class placeholders and their connection relationships from the natural language template or the graph query language template includes: Extracting a number of first-class placeholders from the natural language template or the graph query language template in the word order; Establishing a connection relationship for the extracted number of first-class placeholders in the extraction order to obtain a placeholder path.

3. The method according to claim 2, wherein the step of establishing a connection relationship for the extracted number of first-class placeholders in the extraction order includes: When there are delimiters between a number of first-class placeholders included in the natural language template or the graph query language template, grouping the extracted number of first-class placeholders according to the positions of the delimiters; For each group, establishing a connection relationship for the number of first-class placeholders in the group in the extraction order to obtain the placeholder path corresponding to the group.

4. The method according to claim 1, wherein the step of generating a number of type paths based on the schema definition of the graph data and the placeholder path includes: Sequentially placing the defined types of triple elements included in the schema definition at positions corresponding to the first-class placeholders in the placeholder path to form a number of type paths; wherein, the defined type of the triple element at the current position is selected from a screening set, and the defined types of triple elements in the screening set are screened from the schema definition based on the existing defined types of the triple where the current position is located and the relationship constraints.

5. The method according to claim 4, wherein, At least two placeholder paths are extracted from the natural language template or the graph query language template; the step of generating a number of type paths based on the schema definition of the graph data and the placeholder path includes: When the same placeholder is included in the two placeholder paths, using the defined types of triple elements at the position of the same placeholder in the number of type paths corresponding to one of the generated placeholder paths as the defined types of triple elements at the position of the same placeholder in the other placeholder path.

6. The method according to claim 1, wherein the natural language template and the graph query language template further include a second type of placeholder for defining types of non-triple elements, and the second type of placeholder is used to represent the attributes of associated triple elements; the step of determining the first corpus based on the replacement result includes: For any one of the second type of placeholders in the replacement result, based on the defined type and attributes of the associated triple element, generate instance data for the second type of placeholder, and use the instance data to replace the second type of placeholder in the replacement result to obtain the first corpus.

7. The method according to claim 6, wherein the natural language template further includes a third type of placeholder for representing descriptive words; the step of determining the first corpus based on the replacement result further includes: For any one of the third type of placeholders in the replacement result, select a word from the optional words with different expression forms corresponding to the third type of placeholder stored in advance, and use the selected word to replace the third type of placeholder in the replacement result to obtain the first corpus.

8. The method according to claim 1, wherein the first corpus contains natural language instances and corresponding graph query language instances, and the method further includes: Input the natural language instance into a large model to obtain a generalized natural language instance; Add the first corpus and the generalized natural language instance to the corpus set.

9. The method according to claim 8, wherein the corpus set is used to fine-tune the large model so that the large model can convert the input natural language instance into a corresponding graph query language instance, or so that the large model can convert the graph query language instance into a corresponding natural language instance.

10. The method according to claim 1, wherein the first corpus template is constructed based on the first syntax function in the graph query language.

11. A corpus generation device based on a graph query language, comprising: An acquisition module configured to acquire a first corpus template for constructing a conversion corpus of a graph query language, the first corpus template including a natural language template and a corresponding graph query language template, and the natural language template and the graph query language template include a first type of placeholder for representing the defined type of triple elements; An extraction module configured to extract a placeholder path formed by the first type of placeholder and its connection relationship from the natural language template or the graph query language template; A generation module configured to generate several types of paths based on the schema definition of the graph data and the placeholder path; wherein, the schema definition includes the defined type of triple elements and the relationship constraints therebetween, and the connection relationship of the defined types of triple elements included in any type of path satisfies the relationship constraints; A replacement module configured to use the defined type of triple elements in any one of the type paths to replace the corresponding first type of placeholder in the natural language template and the graph query language template, and determine the first corpus based on the replacement result.

12. A computer-readable storage medium having a computer program stored thereon, which, when executed on a computer, causes the computer to execute the method according to any one of claims 1-10.

13. A computing device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1-10 is implemented.