A knowledge graph complex query method fusing type constraints

CN122817338APending Publication Date: 2026-09-25UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611002813.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本发明的目的在于提供一种融合类型约束的知识图谱复杂查询方法,来解决现有基于知识图谱的复杂查询方法在电商应用场景中存在的技术问题

Benefits of technology

[0029]1.通过在知识图谱复杂查询的推理过程中引入实体类型约束机制,使关系计算仅在语义合理的实体范围内进行,从而在根本上改变传统方法在全体实体空间中无差别计算的模式。在该约束机制作用下,每一种关系在执行投影计算时均被限定在其对应的头实体类型集合与尾实体类型集合所构成的子空间内,有效避免不符合业务语义的实体参与计算。由此,在电商场景中可以显著减少诸如“用户实体被错误识别为商品”或“类目实体参与商品筛选”等不合理匹配现象,从源头上降低噪声引入的概率,显著提升查询结果的准确性与业务一致性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817338A_ABST
    Figure CN122817338A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of electronic commerce and artificial intelligence, and particularly relates to a knowledge graph complex query method fusing type constraints, comprising: constructing entity nodes based on e-commerce business data; adopting a knowledge graph embedding method to perform vectorization representation on the entity nodes and relationships; obtaining implicit type information of the entities by clustering the entity vectors; constructing a head entity acceptable type set and a tail entity acceptable type set; parsing the complex query into a directed acyclic graph structure; performing calculation only in an entity subspace matching the relationship type constraint according to the type constraint set; performing element-by-element combination on the multi-constraint results to realize fusion processing; and performing reasoning step by step according to the topological structure and outputting the results in descending order of scores. The present application improves the accuracy and business consistency of the query results, reduces the calculation complexity, and meets the needs of e-commerce systems for high concurrency and low latency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of e-commerce and artificial intelligence technology, and in particular to a complex query method for knowledge graphs that incorporates type constraints. Background Technology

[0002] With the rapid development of e-commerce platforms and the continuous growth of data scale, e-commerce platforms have accumulated massive amounts of data resources from different sources and with diverse structures, including product information, merchant information, user behavior data, and platform risk control data. This data can typically be abstracted into entities and their relationships in a knowledge graph, and organized in the form of "entity-relationship-entity" triples. For example, relationships such as (Merchant A, Sales, Product X), (Product X, Category, Electronics), and (User U, Purchase, Product X) can all be uniformly represented as structured data in a knowledge graph.

[0003] In actual business operations, platforms often need to execute complex query tasks with multiple constraints. For example, in product filtering scenarios, it may be necessary to query a set of products that simultaneously meet multiple conditions such as category, brand, rating, and merchant attributes. Or, in risk control scenarios, it may be necessary to identify products sold by specific types of merchants or associated with abnormal behavior. These queries typically involve multi-hop relational reasoning and logical operations such as conjunction, disjunction, and negation, and are typical complex logical query problems. Knowledge graphs, as a structured knowledge representation method, can effectively organize and associate multi-source heterogeneous data in the e-commerce field, providing an important technical foundation for complex queries and reasoning.

[0004] Existing knowledge graph-based complex query methods, when performing relational reasoning in e-commerce applications, typically perform calculations across the entire entity space without explicitly considering the entity type constraints corresponding to the relationship. For example, when processing the "sales" relationship, the model often calculates the matching degree across all entities without restricting the relationship to exist only between "merchant" and "product". This calculation method, lacking type constraints, introduces a large number of semantically irrelevant entities, resulting in query results containing content that clearly does not conform to business logic. Furthermore, in multi-hop reasoning and complex query structures, if irrelevant entities are introduced in the initial stage, this noise will accumulate and amplify continuously in subsequent reasoning processes, ultimately causing the query results to deviate from the true target. Summary of the Invention

[0005] The purpose of this invention is to provide a complex query method for knowledge graphs that incorporates type constraints, in order to solve the technical problems existing in the complex query methods based on knowledge graphs in e-commerce application scenarios.

[0006] This invention provides a method for complex querying of knowledge graphs with fused type constraints, comprising:

[0007] A knowledge graph is constructed based on e-commerce business data, in which products, merchants, users and categories are regarded as different types of entity nodes, and sales relationships, category affiliation relationships, purchase behavior relationships and evaluation relationships are regarded as edges connecting entity nodes.

[0008] The knowledge graph embedding method is used to vectorize the entities and relations in the knowledge graph, resulting in entity embedding matrix and relation embedding matrix;

[0009] Clustering entity vectors to obtain implicit type information of entities, classifying entities into at least one of product type, merchant type, user type and category type;

[0010] For each relation in the knowledge graph, we statistically analyze the distribution of head entity types and tail entity types connected in the training data, and construct the set of acceptable head entity types and the set of acceptable tail entity types corresponding to the relation.

[0011] The complex query input by the user is parsed into a directed acyclic graph structure, where nodes represent entities or variables to be determined, and edges represent relational constraints.

[0012] When performing relation projection operations during reasoning, computation is performed only within the entity subspace that matches the relation type constraints, based on a pre-built set of type constraints.

[0013] When a node is simultaneously subject to multiple relations or conditions, the results of each constraint are combined element by element to achieve fusion processing.

[0014] The relation projection and result update are performed step by step according to the topology of the directed acyclic graph until the result representation of the target node is obtained; the candidate entities are scored and sorted according to the result vector of the target node, and the result set that meets the query conditions is output.

[0015] In some embodiments, the e-commerce business data includes at least basic product information, merchant registration and operation information, user behavior data, and platform risk control data. The basic product information includes product identification, category, brand, price range, and rating. The merchant registration and operation information includes store identification, registered location, business qualifications, and historical violation records. The user behavior data includes browsing, clicking, adding to cart, and purchasing behaviors. The risk control data includes complaint records, violation tags, and risk level markings.

[0016] In some embodiments, the original business data is cleaned and standardized, including removing duplicate data, filling in missing values, and standardizing field formats. The structured data is then converted into triples to construct an e-commerce knowledge graph.

[0017] In some embodiments, when using a knowledge graph embedding method for vectorization, each entity is mapped to a vector representation in a low-dimensional vector space. Training is performed by minimizing the distance difference between real triples and negative sample triples, making semantically related entities closer in the vector space.

[0018] In some embodiments, the trained entity embedding matrix is ​​E∈R^(d×|E|), and the relation embedding matrix is ​​R∈R^(d×|R|), where d is the embedding dimension. The training is performed by designing an interval-based ranking loss or cross-entropy loss as the objective function.

[0019] In some embodiments, the method for clustering entity vectors includes the K-means clustering algorithm or the hierarchical clustering algorithm. Let the set of entity types obtained after clustering be T={T1, T2, ..., Tk}, where k is the number of types, and each entity is assigned to one or more types.

[0020] In some embodiments, the set of acceptable types for the head entity, Ah(r), is defined as: Ah(r) = {t | there exists a triple (h, r, t') such that h is of type t}; the set of acceptable types for the tail entity, At(r), is defined as: At(r) = {t | there exists a triple (h', r, t) such that t is of type t}.

[0021] In some embodiments, when performing a projection operation on a relation r;

[0022] The candidate head entity set Sh is determined based on the set of acceptable head entity types Ah(r);

[0023] The candidate tail entity set St is determined based on the set of acceptable tail entity types At(r);

[0024] Projection calculations are performed based on relation embedding vectors within the entity subspace consisting of the candidate head entity set Sh and the candidate tail entity set St.

[0025] In some embodiments, projection calculation is represented as:

[0026] h' = σ((E_Sh)ᵀ · r_vec · (E_St) · h), where E_Sh and E_St are submatrices extracted from the entity embedding matrix E, corresponding to the candidate head entity set and the candidate tail entity set, respectively, r_vec is the embedding vector corresponding to relation r, and σ is the activation function.

[0027] In some embodiments, combining the constraint results element by element includes performing at least one of conjunction, disjunction and negation operations. For queries that include negation conditions, a complement operation is performed on the relation projection result set, and then the category filtering results and brand exclusion results are combined element by element to obtain a product set that satisfies all conditions.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] 1. By introducing an entity type constraint mechanism into the reasoning process of complex queries in knowledge graphs, relation calculations are performed only within semantically reasonable entities, fundamentally changing the traditional method's indiscriminate calculation across the entire entity space. Under this constraint mechanism, each relation is confined to the subspace formed by its corresponding head entity type set and tail entity type set during projection calculation, effectively preventing entities that do not conform to business semantics from participating in the calculation. Therefore, in e-commerce scenarios, this can significantly reduce unreasonable matching phenomena such as "user entities being incorrectly identified as products" or "category entities participating in product filtering," reducing the probability of noise introduction at the source and significantly improving the accuracy and business consistency of query results.

[0030] 2. Because the computational scope during the reasoning process is reduced from a high-dimensional space covering all entities to a low-dimensional subspace limited by type constraints, the computational scale involved in relation projection and intermediate result updates is effectively compressed. In actual e-commerce knowledge graphs, the number of entities typically reaches hundreds of thousands or even higher, while the subset of entities corresponding to a single type often accounts for only a small portion of the whole. Therefore, type constraints can reduce the number of participating entities in each step of the computation by more than an order of magnitude, thereby significantly reducing time and space complexity. This characteristic enables the method of this invention to meet the actual needs of e-commerce systems for high concurrency and low latency while ensuring query accuracy, and it has good real-time response capabilities.

[0031] 3. In multi-hop reasoning and complex query structures, this invention effectively blocks the propagation path of irrelevant entities in the graph structure by continuously applying type constraints at each level of reasoning. Unlike existing methods that introduce noise in the initial stage and amplify it in subsequent processes, this invention can semantically filter intermediate results at each computation step, thereby significantly reducing the cumulative effect of errors. This feature is particularly important in query tasks involving multi-condition filtering, multi-path constraints, and complex dependencies, and can significantly improve the stability and robustness of reasoning results.

[0032] 4. At the engineering implementation level, the type constraint mechanism adopted in this invention does not require the introduction of complex nonlinear logic operators or additional high-overhead computation modules. Instead, it can be achieved by dynamically pruning the computation space during the inference stage by utilizing existing knowledge graph structures and statistical information. Therefore, this method is characterized by its simplicity, low computational overhead, and ease of integration. It can be seamlessly embedded into the existing e-commerce platform's search system, recommendation system, or risk control system, achieving performance improvements without requiring large-scale modifications to the original architecture.

[0033] 5. The method of this invention has good adaptability to query structures. It is not only suitable for simple single-path queries, but also supports directed acyclic graph structure queries that include multiple relational constraints and combinations of multiple conditions. In actual e-commerce business, user filtering conditions often exhibit multi-dimensional combination characteristics, such as simultaneously involving category, brand, price range, and merchant attributes. The method of this invention can uniformly handle the above complex structures and output results that meet multiple business constraints while ensuring efficiency, thereby improving the overall service capability of the system. Attached Figure Description

[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0035] Figure 1 This is a flowchart of the complex query method for knowledge graphs with fusion type constraints for e-commerce platforms according to the present invention.

[0036] Figure 2 This is a schematic diagram of the complex query method for knowledge graphs with fusion type constraints for e-commerce platforms according to the present invention.

[0037] Figure 3 This is a flowchart of complex query reasoning for knowledge graphs with integrated type constraints for e-commerce platforms, based on the present invention. Detailed Implementation

[0038] The following will be based on embodiments of the present invention. Figures 1-3 The technical solutions in the embodiments of the present invention will be clearly and completely described together. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0039] Example 1

[0040] refer to Figure 2The complex query method for knowledge graphs with fusion type constraints for e-commerce platforms of the present invention can be understood as including the steps of importing a new knowledge graph, uploading a knowledge graph file, writing to a database, determining whether it has been trained, executing the training model, and completing the training of the new graph.

[0041] This method constructs a knowledge graph based on e-commerce business data, and uses knowledge graph embedding methods to vectorize entities and relations. It obtains implicit type information of entities by clustering entity vectors, and then constructs a set of type constraints corresponding to the relations. When executing complex queries, the query is parsed into a directed acyclic graph structure. During inference, relation projection calculations are performed only within semantically reasonable entity subspaces based on the type constraint set, achieving result fusion processing of multiple constraints. Finally, the inference is executed step by step according to the topological structure of the directed acyclic graph, and results that satisfy the query conditions are output.

[0042] The first step in importing the new knowledge graph involves receiving multi-source business data from the e-commerce platform. This data includes basic product information, merchant registration and operational information, user behavior data, and platform risk control data. Basic product information includes attributes such as product identifiers, categories, brands, price ranges, and ratings, describing the basic characteristics of products sold on the platform. Merchant registration and operational information includes store identifiers, registered locations, business qualifications, and historical violation records, used to characterize the merchant's basic attributes and operational status. User behavior data includes browsing, clicking, adding to cart, and purchasing behaviors, recording various user actions on the platform. Risk control data includes complaint records, violation tags, and risk level markings, used to identify merchants or products that may pose a risk.

[0043] In the step of uploading the knowledge graph file, the received raw business data is preprocessed. First, the raw data is cleaned to remove duplicate records, fill in missing field values, and standardize inconsistent field formats. For example, date formats from different data sources are unified to YYYY-MM-DD, and price fields are standardized to integers in cents. The cleaned structured data is then converted into triples for storage, using a standard (head entity, relation, tail entity) representation.

[0044] Specifically, the following triples are used to represent "a merchant sells a product" (merchant A, sells, product X), "the product belongs to a category" (product X, belongs to category, electronic products), "the product belongs to a brand" (product X, brand, brand B), "a user buys a product" (user U, buys, product X), "the product has a rating" (product X, rating, 4.5 points), and "the merchant has qualifications" (merchant A, has qualifications, business license), etc.

[0045] In the database writing step, the constructed triplet data is persistently stored in a graph database. The graph database employs a storage engine suitable for storing large-scale graph structures, supporting efficient graph traversal and relation query operations. During the writing process, a unique identifier is automatically assigned to each entity and relation, and a corresponding index structure is established to accelerate subsequent query retrieval. The database also stores the original attribute information of entities and the metadata of relations, providing a statistical basis for subsequent type constraint construction.

[0046] In the step of determining whether training has been completed, it checks whether the current knowledge graph has completed the training of the embedding model. If the knowledge graph has been trained and the embedding model parameters have been saved in the model library, it directly proceeds to the new graph training completion step, skipping the model training process. If the knowledge graph has not been trained or needs to be retrained, it proceeds to the step of executing the training model to train the knowledge graph embedding model.

[0047] In the model training step, a knowledge graph embedding method is used to vectorize the entities and relations in the knowledge graph. Assuming the entity set is represented by E and the relation set by R, the trained entity embedding matrix is ​​represented as E∈R^(d×|E|), and the relation embedding matrix is ​​represented as R∈R^(d×|R|), where d represents the dimension of the embedding vector. During training, each entity is mapped to a vector representation in a d-dimensional vector space, and optimization is performed by minimizing the distance difference between the true triples and the negative triples. It should be noted that the negative triples are obtained by randomly replacing the head or tail entities in the true triples.

[0048] The specific training objective function adopts a margin-based ranking loss function, which is formally defined as follows:

[0049] L = Σ_{(h,r,t)∈Ω} Σ_{(h',r,t')∈Ω'} [γ + sim(f(h,r,t)) - sim(f(h',r,t'))]₊,

[0050] Where Ω represents the set of positive sample triples, Ω' represents the set of negative sample triples, γ represents the interval hyperparameter, sim(·) represents the triple score function, [x]₊ represents max(x,0), and f(h,r,t) represents the score calculation of triple (h,r,t).

[0051] During training, the parameters of the entity embedding matrix and relation embedding matrix are updated using stochastic gradient descent, making entity pairs (h,t) that satisfy the triplet (h,r,t) closer in the vector space, while entity pairs that do not satisfy the triplet are farther apart. After training, semantically similar entities have higher similarity in the vector space.

[0052] After completing entity representation learning, the implicit type information of the entities is automatically obtained by clustering the entity vectors. The clustering algorithm uses K-means clustering or hierarchical clustering to divide the entities into several type sets.

[0053] Let the set of entity types obtained after clustering be T = {T1, T2, ..., T}. k} where k is the number of types. Each entity e∈E is assigned to one or more types, and the same entity can be divided into multiple implicit types based on its semantic features. In e-commerce knowledge graph scenarios, the entity type set typically includes product type Tproduct, merchant type Tmerchant, user type Tuser, and category type Tcategory, etc.

[0054] Based on the entity type information obtained from training, a set of type constraints corresponding to the relations is further constructed. For each relation in the knowledge graph, the distribution of head entity types and tail entity types connected to it in the training data are statistically analyzed.

[0055] Specifically, for a relation *r*, the set of acceptable types for the head entity, *A_h(r)*, is defined as *A_h(r)* = {t | there exists a triple (h, r, t') such that the type of *h* is *t*}; the set of acceptable types for the tail entity, *A_t(r)*, is defined as *A_t(r)* = {t | there exists a triple (h', r, t) such that the type of *t* is *t*}. For example, by counting all instances of the "sales" relation, we can determine that its head entity mainly belongs to the merchant type and its tail entity mainly belongs to the product type, thus constructing the corresponding type constraint sets *A_h(sales)* = {merchant type} and *A_t(sales)* = {product type}. For the "belongs to category" relation, its head entity belongs to the product type and its tail entity belongs to the category type; therefore, *A_h(belongs to category)* = {product type} and *A_t(belongs to category)* = {category type}.

[0056] For the "purchase" relationship, the head entity belongs to the user type, and the tail entity belongs to the product type. Therefore, A_h(purchase) = {user type}, and A_t(purchase) = {product type}. These type constraint sets are stored in the form of a mapping table for limiting the computational scope in subsequent reasoning processes.

[0057] In the complex query reasoning stage, the first step is to select a knowledge graph, choosing the target graph for the query from the imported knowledge graphs.

[0058] In the query input step, complex query conditions are received from the user. For example, if the user inputs "filter for products that belong to the electronics category but do not belong to a specified brand", this query includes category constraints and brand exclusion conditions, which is a typical multi-condition combination query.

[0059] In the step of determining whether a problem can be represented in first-order logic form, the input natural language query is analyzed to see if it can be converted into a standard first-order logic form. If the query can be represented as a combination of conjunctions, disjunctions, or negations of atomic formulas, the process proceeds to the step of converting the problem into a first-order logic form. If the query cannot be directly represented in first-order logic form, a query parser based on natural language processing is invoked to convert the natural language into a logical form.

[0060] In the step of converting the problem into a first-order logical form, the input query conditions are parsed into a first-order logical form. Taking filtering products that belong to the electronics category but do not belong to a specified brand as an example, the query is parsed into the first-order logical form: ∃x (Product(x) ∧ BelongsToCategory(x, electronics) ∧ ¬BelongsToBrand(x, specified brand)), where x represents the product variable to be queried, Product represents the product type predicate, BelongsToCategory represents the category relationship predicate, BelongsToBrand represents the brand relationship predicate, and ¬ represents the logical negation operator.

[0061] In the step of identifying query types and generating query graphs, the query types are identified according to first-order logic forms, and corresponding directed acyclic graph structures are generated. The query graph structure is formally represented as G=(V,E), where V is the set of nodes and E is the set of edges. In this embodiment, the query graph includes a target node v_target representing the set of products to be filtered, a category constraint node v_category representing the "electronic products" entity, a brand constraint node v_category representing the "specified brand" entity, edge e1 representing the "belongs to category" relationship from product to category, and edge e2 representing the "brand" relationship from product to brand.

[0062] Before inference begins, the nodes in the query graph are initialized. For known entity nodes, such as "electronic product category" and "specified brand", they are initialized as indicator vectors with the corresponding position set to 1. For variable nodes to be determined, i.e., product set nodes, they are initialized as candidate vectors for the full space of entities of the corresponding type.

[0063] The specific initialization process is as follows: h_category = one-hot (electronic products), h_brand = one-hot (specified brand), h_target = all-1 vector (representing all product entities).

[0064] In the step of calling the LPT-DAG model, the system performs type constraint relationship projection calculation according to the topological order of the query graph. The core of the LPT-DAG model is to continuously apply type constraints during each level of inference, limiting the calculation scope to the entity subspace that satisfies the set of type constraints.

[0065] When performing the projection operation for the "belongs to category" relationship, the candidate head entity set S_h is first determined based on the head entity's acceptable type set A_h (belongs to category) = {product type}, and the candidate tail entity set S_t is determined based on the tail entity's acceptable type set A_t (belongs to category) = {category type}. The candidate head entity set S_h includes all entities in the knowledge graph that belong to the product type, and the candidate tail entity set S_t includes all entities in the knowledge graph that belong to the category type.

[0066] Subsequently, within the entity subspace comprised of the candidate head entity set S_h and the candidate tail entity set S_t, projection calculation is performed based on the relation embedding vector. Let the embedding vector corresponding to the "belongs to category" relation be r_category_vec. The candidate head entity submatrix extracted from the entity embedding matrix E is represented as E_{S_h}, and the candidate tail entity submatrix is ​​represented as E_{S_t}. The projection calculation is expressed as: h'target = σ((E{S_h})^T · r_category_vec · (E_{S_t}) · h_category), where σ is the activation function. This calculation process propagates the constraint information of the electronic product node along the "belongs to category" relation edge to the product node, obtaining the vector representation h'_target of the candidate product set belonging to the electronic product category.

[0067] When performing the projection operation on brand relationships, the candidate entity set is determined based on the set of acceptable types for the head entity, A_h(brand) = {product type}, and the set of acceptable types for the tail entity, A_t(brand) = {brand type}. Projection calculations are performed within the candidate entity subspace to obtain the vector representation h_brand_result of the product set belonging to the specified brand. Since the query includes a negative condition, i.e., requiring "not belonging to the specified brand", a complement operation is performed on the brand relationship projection result set.

[0068] Let h_brand_initial be an all-one vector representing all product entities. The negation operation is represented as: h_brand_negated = h_brand_initial - h_brand_result. This operation removes all products belonging to the specified brand from the candidate product set.

[0069] When multiple constraints apply to the same target node, the results of each condition are merged. Specifically, the category filtering results and brand exclusion results are combined element by element to obtain a set of products that simultaneously meet all conditions: h_final = h'_target ⊙ h_brand_negated, where ⊙ represents element-wise multiplication.

[0070] Throughout the reasoning process, because each step is subject to type constraints, irrelevant entities are filtered out at the initial stage. Specifically, when performing a projection of a category relationship, calculations are performed only within the subspace of the product type entity, completely excluding the possibility of non-product type entities such as users and merchants participating in the calculation. When performing a projection of a brand relationship, calculations are performed only within the subspaces of the product type entity and the brand type entity.

[0071] Following the topology of the directed acyclic graph, relation projection and result update operations are performed step by step. In each step, nodes that have been computed are removed until the result representation of the target node is finally obtained. Type constraints are continuously applied at each level of reasoning to effectively block the propagation path of irrelevant entities in the graph structure.

[0072] In the result return step, candidate entities are scored and sorted according to the result vector of the target node, and a set of products that meet the query conditions is output. Specifically, this includes: normalizing the result vector and sorting it in descending order of the scores of each entity. In this embodiment, the normalization process uses the softmax function to convert the result vector into a probability distribution form, representing the probability that each product meets the query conditions. Finally, the system outputs the N products with the highest scores as the query results.

[0073] For each output result, its reasoning path information is also retained. The reasoning path records all the reasoning steps that the product went through from the initial candidate set through various layers of constraint filtering to finally become the output result, including the types of relations, constraints, and intermediate results.

[0074] In practical applications, the method of this invention can be integrated into the product search and recommendation module of an e-commerce platform. When a user inputs filtering conditions or the system executes an automatic recommendation task, the above-mentioned reasoning process can be invoked in real time to quickly filter out a set of products that meet the conditions from the knowledge graph. Since the reasoning computation is performed only in a subspace limited by type constraints, compared with the traditional method of performing computation in the entire entity space, the computational complexity is significantly reduced, enabling a high response speed to be maintained even in large-scale data environments.

[0075] Specifically, when a knowledge graph includes hundreds of thousands of entities, the subset of entities corresponding to a single type usually accounts for only a small part of the whole. By using type constraints, the number of participating entities in each step of the calculation can be reduced by more than an order of magnitude, thereby significantly reducing time and space complexity and meeting the actual needs of e-commerce systems for high concurrency and low latency.

[0076] Example 2

[0077] This invention is applied to risk control scenarios on e-commerce platforms, where it is necessary to identify products sold by specific types of merchants and associated with abnormal behavior.

[0078] Suppose the query condition is "find merchants registered in a specific region who have sold the complained-about product". This query involves multi-hop relation reasoning. First, the projection operation of the "sales" relation is performed. Based on the entity subspace determined by A_h(sales) = {merchant type} and A_t(sales) = {product type}, the set of merchants who have sold the complained-about product is calculated.

[0079] Then, the projection operation of the "registration location" relationship is performed. Based on the entity subspace determined by A_h(registration location) = {merchant type} and A_t(registration location) = {region type}, the set of merchants whose registration location is in a specific region is calculated.

[0080] Finally, the results of the two constraints are conjuncted to obtain the set of merchants that simultaneously satisfy both conditions. In this process, the type constraint ensures that the calculation of the "sales" relationship is performed only between merchants and products, and the calculation of the "registration location" relationship is performed only between merchants and regions, avoiding unreasonable matching across types.

[0081] Example 3

[0082] This invention is applied to personalized recommendation scenarios on e-commerce platforms, where products need to be recommended based on users' historical behavior and preferences. Assuming the query condition is "find products belonging to the same category as products already purchased by the user and whose ratings are higher than the average rating of products the user has historically purchased," this query involves projecting purchase behavior relationships and rating comparison constraints.

[0083] First, the set of products that the user has purchased is determined based on the "purchase" relationship and the set of type constraints. Then, the set of other products under the same category is determined based on the "belongs to category" relationship. Next, products with ratings higher than the threshold are selected based on the "rating" relationship. Finally, the results of each constraint are fused to obtain the recommended candidate set.

[0084] Throughout the reasoning process, type constraints ensure that the calculation of the "purchase" relationship is performed only between the user and the product, and the calculation of the "rating" relationship is performed only between the product and the rating value, thereby guaranteeing the semantic rationality of the reasoning results.

[0085] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the scope of the invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

[0086] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment includes only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A complex query method for knowledge graphs that incorporates type constraints, characterized in that, include: A knowledge graph is constructed based on e-commerce business data, in which products, merchants, users and categories are regarded as different types of entity nodes, and sales relationships, category affiliation relationships, purchase behavior relationships and evaluation relationships are regarded as edges connecting entity nodes. The knowledge graph embedding method is used to vectorize the entities and relations in the knowledge graph, resulting in entity embedding matrix and relation embedding matrix; Clustering entity vectors to obtain implicit type information of entities, classifying entities into at least one of product type, merchant type, user type and category type; For each relation in the knowledge graph, we statistically analyze the distribution of head entity types and tail entity types connected in the training data, and construct the set of acceptable head entity types and the set of acceptable tail entity types corresponding to the relation. The complex query input by the user is parsed into a directed acyclic graph structure, where nodes represent entities or variables to be determined, and edges represent relational constraints. When performing relation projection operations during reasoning, computation is performed only within the entity subspace that matches the relation type constraints, based on a pre-built set of type constraints. When a node is simultaneously subject to multiple relations or conditions, the results of each constraint are combined element by element to achieve fusion processing. The relation projection and result update are performed step by step according to the topology of the directed acyclic graph until the result representation of the target node is obtained; the candidate entities are scored and sorted according to the result vector of the target node, and the result set that meets the query conditions is output.

2. The method according to claim 1, characterized in that, The e-commerce business data includes at least basic product information, merchant registration and operation information, user behavior data, and platform risk control data. Among them, basic product information includes product identification, category, brand, price range, and rating; merchant registration and operation information includes store identification, registered location, business qualifications, and historical violation records; user behavior data includes browsing, clicking, adding to cart, and purchasing behaviors; and risk control data includes complaint records, violation tags, and risk level markings.

3. The method according to claim 2, characterized in that, The original business data is cleaned and standardized, including removing duplicate data, filling in missing values, and standardizing field formats. The structured data is then converted into triples to build an e-commerce knowledge graph.

4. The method according to claim 1, characterized in that, When using knowledge graph embedding methods for vectorization, each entity is mapped to a vector representation in a low-dimensional vector space. Training is performed by minimizing the distance difference between real triples and negative sample triples, making semantically related entities closer in the vector space.

5. The method according to claim 4, characterized in that, The trained entity embedding matrix is ​​E∈R^(d×|E|), and the relation embedding matrix is ​​R∈R^(d×|R|), where d is the embedding dimension. The training is performed by designing a margin-based ranking loss or cross-entropy loss as the objective function.

6. The method according to claim 1, characterized in that, Methods for clustering entity vectors include K-means clustering algorithm or hierarchical clustering algorithm. Let the set of entity types obtained after clustering be T={T1, T2, ..., Tk}, where k is the number of types, and each entity is assigned to one or more types.

7. The method according to claim 1, characterized in that, The set of acceptable types for the head entity, Ah(r), is defined as: Ah(r) = {t | there exists a triple (h, r, t') such that h is of type t}; the set of acceptable types for the tail entity, At(r), is defined as: At(r) = {t | there exists a triple (h', r, t) such that t is of type t}.

8. The method according to claim 7, characterized in that, When performing a projection operation on a relation r; The candidate head entity set Sh is determined based on the set of acceptable head entity types Ah(r); The candidate tail entity set St is determined based on the set of acceptable tail entity types At(r); Projection calculations are performed based on relation embedding vectors within the entity subspace consisting of the candidate head entity set Sh and the candidate tail entity set St.

9. The method according to claim 7, characterized in that, Projection calculation is expressed as: h' = σ((E_Sh)ᵀ · r_vec · (E_St) · h), where E_Sh and E_St are submatrices extracted from the entity embedding matrix E, corresponding to the candidate head entity set and the candidate tail entity set, respectively, r_vec is the embedding vector corresponding to relation r, and σ is the activation function.

10. The method according to claim 1, characterized in that, The constraints are combined element by element, including performing at least one of conjunction, disjunction and negation operations. For queries that include negation conditions, the set of relation projection results is complemented. Then, the category filtering results and the brand exclusion results are combined element by element to obtain a set of products that simultaneously meet all conditions.