Knowledge graph determination method and device applied to tobacco production scheduling

By extracting data and aligning multiple data sources, combining relationship fusion processing, an accurate knowledge graph in the field of tobacco production scheduling is built, solving the problems of low construction efficiency and inaccurate relationship information in the existing technology.

CN120218216APending Publication Date: 2025-06-27CHINA TOBACCO ZHEJIANG IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510375658.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is inefficient in building a knowledge graph in the field of tobacco production scheduling, and it is difficult to accurately determine the relationship information between entities, especially when dealing with heterogeneous data sources.

Method used

By performing data extraction and processing on multiple data sources associated with tobacco production scheduling, an initial tobacco knowledge graph is constructed, and the target tobacco knowledge graph is determined through entity alignment and relationship fusion processing.

Benefits of technology

The accurate construction of the knowledge graph in the field of tobacco production scheduling is achieved, the construction efficiency is improved, and the heterogeneity and ambiguity problems in heterogeneous data sources are effectively dealt with.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218216A_ABST
    Figure CN120218216A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge graph determination method and device applied to tobacco production scheduling. The method comprises the following steps: respectively performing data extraction on at least two data sources associated with tobacco production scheduling to obtain tobacco production scheduling information and first relation information; based on the tobacco production scheduling information and the first relation information of each data source, determining an initial tobacco knowledge graph of each data source; performing entity alignment on the initial entity information of the at least two initial tobacco knowledge maps, and determining an entity information group and target entity information; determining target relation information according to the first relation information, second relation information of the target entity information in the at least two data sources and third relation information of the target entity information in the initial tobacco knowledge graph; and fusing the at least two initial tobacco knowledge maps based on the target entity information and the target relationship information to determine the target tobacco knowledge map, thereby realizing accurate construction of the target tobacco knowledge map, and providing technical guidance for tobacco production scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a method and device for determining a knowledge graph applied to tobacco production scheduling. Background Art

[0002] A knowledge graph is a graph data structure used to represent and store a large number of entities and their corresponding relationships. The knowledge graph is usually presented in the form of a graph, where the nodes in the graph represent entities, and the edges represent the relationships between these entities.

[0003] Currently, a knowledge graph corresponding to the tobacco production scheduling process is usually constructed manually, that is, through predefined ontology information, the corresponding staff extracts entity information and relationship information from the data source, and constructs the entity information and relationship information into a knowledge graph according to the ontology information. However, when constructing a knowledge graph based on a large-scale heterogeneous data source, the above method not only has low construction efficiency, but also cannot accurately determine the relationship information between entities. In addition, the above method is also difficult to handle the heterogeneity and ambiguity problems corresponding to the heterogeneous data source, resulting in an inaccurate knowledge graph constructed. Summary of the Invention

[0004] The present invention provides a method and device for determining a knowledge graph applied to tobacco production scheduling, realizing the accurate construction of the tobacco knowledge graph in the field of tobacco production scheduling, and providing technical guidance for subsequent tobacco production scheduling.

[0005] According to one aspect of the present invention, there is provided a method for determining a knowledge graph applied to tobacco production scheduling, the method comprising:

[0006] For at least two data sources associated with tobacco production scheduling, perform data extraction processing on each data source to obtain tobacco production scheduling information corresponding to each data source and first relationship information between the tobacco production scheduling information;

[0007] Based on the tobacco production scheduling information and the first relationship information corresponding to each data source, determine an initial tobacco knowledge graph corresponding to each data source, wherein the tobacco production scheduling information is used as initial entity information in the initial tobacco knowledge graph;

[0008] Perform entity alignment processing on the initial entity information in at least two initial tobacco knowledge graphs to determine an entity information group corresponding to the at least two initial tobacco knowledge graphs and target entity information corresponding to the entity information group, wherein the entity information group includes initial entity information with a matching relationship;

[0009] Determine the second relationship information of the target entity information in at least two data sources, and the third relationship information in the initial tobacco knowledge graph. Based on the first relationship information, the second relationship information, and the third relationship information, determine the target relationship information corresponding to the target entity information;

[0010] Based on the target entity information and the target relationship information, fuse at least two initial tobacco knowledge graphs to determine the target tobacco knowledge graph.

[0011] According to another aspect of the present invention, there is provided a knowledge graph determination device applied to tobacco production scheduling. The device includes:

[0012] A relationship information determination module, configured to perform data extraction processing on each of at least two data sources associated with tobacco production scheduling, and obtain the first relationship information between the tobacco production scheduling information corresponding to each data source respectively;

[0013] An initial tobacco knowledge graph determination module, configured to determine the initial tobacco knowledge graph corresponding to each data source respectively based on the tobacco production scheduling information and the first relationship information corresponding to each data source respectively, where the tobacco production scheduling information is used as the initial entity information in the initial tobacco knowledge graph;

[0014] An entity alignment module, configured to perform entity alignment processing on the initial entity information in at least two initial tobacco knowledge graphs, determine the entity information group corresponding to at least two initial tobacco knowledge graphs, and the target entity information corresponding to the entity information group, where the entity information group includes initial entity information with a matching relationship;

[0015] A target relationship information determination module, configured to determine the second relationship information of the target entity information in at least two data sources, and the third relationship information in the initial tobacco knowledge graph. Based on the first relationship information, the second relationship information, and the third relationship information, determine the target relationship information corresponding to the target entity information;

[0016] A target tobacco knowledge graph determination module, configured to fuse at least two initial tobacco knowledge graphs based on the target entity information and the target relationship information to determine the target tobacco knowledge graph.

[0017] According to another aspect of the present invention, there is provided an electronic device, which includes:

[0018] At least one processor; and

[0019] A memory communicatively connected to at least one processor; wherein,

[0020] The memory stores a computer program that can be executed by at least one processor. The computer program is executed by at least one processor to enable the at least one processor to execute the knowledge graph determination method applied to tobacco production scheduling according to any embodiment of the present invention.

[0021] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the knowledge graph determination method applied to tobacco production scheduling according to any embodiment of the present invention when executed.

[0022] According to another aspect of the present invention, there is provided a computer program product including a computer program, characterized in that the computer program implements the knowledge graph determination method applied to tobacco production scheduling according to any embodiment of the present invention when executed by a processor.

[0023] The technical solution of the embodiment of the present invention is directed to at least two data sources associated with tobacco production scheduling. Data extraction processing is performed on each data source to obtain the tobacco production scheduling information corresponding to each data source and the first relationship information between the tobacco production scheduling information. According to the tobacco production scheduling information and the first relationship information corresponding to each data source, an initial tobacco knowledge graph corresponding to each data source is determined. Based on this, the tobacco knowledge graph corresponding to each data source associated with tobacco production scheduling is determined, providing data support for the construction of the subsequent target tobacco knowledge graph. Entity alignment processing is performed on the initial entity information in at least two initial tobacco knowledge graphs to obtain an entity information group corresponding to the at least two initial tobacco knowledge graphs and the target entity information corresponding to the entity information group. Based on this, the entity information required for constructing the target tobacco knowledge graph is accurately determined. The second relationship information of the target entity information in at least two data sources and the third relationship information in the initial tobacco knowledge graph are determined to perform relationship fusion processing according to the first relationship information, the second relationship information, and the third relationship information to obtain the target relationship information. Based on the target entity information and the target relationship information, at least two initial tobacco knowledge graphs are fused to obtain the target tobacco knowledge graph. The present invention solves the problems in the prior art of low efficiency in constructing a knowledge graph and inability to accurately determine the relationship information between entities, and by performing entity and relationship fusion processing on the initial tobacco knowledge graphs of at least two data sources, solves the problem in the prior art that the constructed knowledge graph is inaccurate due to the heterogeneity and ambiguity corresponding to heterogeneous data sources, realizing the accurate construction of the tobacco knowledge graph in the field of tobacco production scheduling and providing technical guidance for subsequent tobacco production scheduling.

[0024] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. Description of the Drawings

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0026] Figure 1 is a flowchart of a method for determining a knowledge graph applied to tobacco production scheduling provided by an embodiment of the present invention;

[0027] Figure 2 is a flowchart of a method for determining a knowledge graph applied to tobacco production scheduling provided by an embodiment of the present invention;

[0028] Figure 3 is a schematic structural diagram of a device for determining a knowledge graph applied to tobacco production scheduling provided by an embodiment of the present invention;

[0029] Figure 4 is a schematic structural diagram of an electronic device for implementing the method for determining a knowledge graph applied to tobacco production scheduling according to an embodiment of the present invention. Detailed Embodiments

[0030] In order to enable those skilled in the art to better understand the solution of the present invention, the following clearly and completely describes the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0031] It should be noted that the terms "first", "second", etc. in the description, claims and the above drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0032] Embodiment 1

[0033] Figure 1 It is a flowchart of a method for determining a knowledge graph applied to tobacco production scheduling provided by Embodiment 1 of the present invention. This embodiment is applicable to the situation of constructing a target tobacco knowledge graph in the field of tobacco production scheduling. This method can be executed by a device for determining a knowledge graph applied to tobacco production scheduling. The device for determining a knowledge graph applied to tobacco production scheduling can be implemented in the form of hardware and / or software, and the device for determining a knowledge graph applied to tobacco production scheduling can be configured in an electronic device such as a mobile phone, a computer or a server. As Figure 1 shown, the method includes:

[0034] S110. For at least two data sources associated with tobacco production scheduling, perform data extraction processing on each data source to obtain the tobacco production scheduling information corresponding to each data source and the first relationship information between the tobacco production scheduling information.

[0035] Among them, tobacco production scheduling can be understood as the production and scheduling process of corresponding tobacco products. At least two data sources associated with tobacco production scheduling can be databases or data tables associated with tobacco production scheduling. Optionally, the at least two data sources may include a database storing business data of tobacco production scheduling, a processing technology document of tobacco products, and other databases or data tables associated with tobacco production scheduling. This embodiment does not limit the specific data sources.

[0036] Tobacco production scheduling information may include various information associated with tobacco production scheduling. For example, tobacco production scheduling information may include: production order information of tobacco products, production resource information, production process information, production plan information, and quality index information, etc. Among them, the production order information of tobacco products is used to describe the order information of tobacco products, and the production order information includes: order number, product type, order quantity, and delivery date, etc. The production resource information is used to describe various resource information required for the production of tobacco products, and the production resource information includes: production raw materials, production equipment, and production personnel, etc. The production process information is used to describe each process in the tobacco production process. Each process may include: leaf blending, cutting, tipping, and packaging, etc. The production plan information can be understood as the production task arrangement information specified according to the production order and production resources. The production plan information may include: production batch, production start time, equipment used in production, etc. The quality index information can be used to represent various indicators for measuring the quality of tobacco products. The quality index information may include: cigarette qualification rate, tobacco filling value, and cigarette weight, etc. Among them, the tobacco filling value can be understood as the volume that unit weight of tobacco can occupy under standard conditions, and is used to reflect the looseness and filling performance of tobacco.

[0037] Optionally, corresponding attribute information can be extracted from the data source. For example, for the production order information of tobacco products, it may include attribute information such as order number, product type, order quantity, and delivery date. Among them, the order number is the code used to uniquely represent the production order and belongs to the data type attribute. The product type can be understood as the type of tobacco products required to be produced by the production order and belongs to the object type attribute. The order quantity can be understood as the quantity of tobacco products required to be produced by the production order and belongs to the data type attribute. The delivery date can be understood as the delivery date required by the production order and belongs to the data type attribute. The attribute information can also be used as tobacco production scheduling information.

[0038] The first relationship information can be understood as the relationship information between each tobacco production scheduling information. For example, there is an inclusion relationship between the production order information and the product type in the tobacco production scheduling information, that is, the production order information includes the product type. For example, the first relationship information between the production plan information and the production order information in the tobacco production scheduling information can be: the production order information is formulated based on the production plan information. For example, the first relationship information between the production plan information and the production resource information can be: the production resource information is allocated or determined according to the production plan information.

[0039] Specifically, for at least two data sources associated with tobacco production scheduling, natural language processing technology is used to perform data extraction processing on each data source respectively, obtaining the tobacco production scheduling information corresponding to each data source and the first relationship information between every two pieces of tobacco production scheduling information. Optionally, before performing data extraction on at least two data sources, ontology information in the field of tobacco production scheduling can be constructed to perform data extraction processing on each data source based on the ontology information of tobacco production scheduling. It should be noted that the ontology information can provide a structured framework for defining entities, attributes, relationships, and their constraints in the field of tobacco production scheduling, and can provide support for the subsequent construction of the knowledge graph.

[0040] Exemplarily, after determining that the current domain scope is the tobacco production scheduling domain, formal representations can be made for all aspects of information corresponding to tobacco production scheduling according to various aspects of information related to tobacco production scheduling, such as production plans, resource allocation, process optimization, and quality control, etc., to construct a structured ontology in the field of tobacco production scheduling, and evaluate the ontology to obtain an evaluation result, so as to adjust the ontology based on the evaluation result. Tobacco production scheduling information and first relationship information are extracted from at least two data sources according to the ontology information of the ontology.

[0041] In the embodiment of the present invention, the specific manner of determining the tobacco production scheduling information and the first relationship information of each data source may be: performing data extraction processing on the unstructured data in each data source to determine the first data corresponding to each data source, and performing classification processing on the structured data of each data source to obtain a classification result, and determining the second data of each data source according to the classification result; based on the first data and the second data, determining the tobacco production scheduling information of each data source and the first relationship information between the tobacco production scheduling information.

[0042] Among them, unstructured data can be understood as data without a predefined data model or data structure. Unstructured data can be data in formats such as text, images, videos, and audios. Structured data can be data stored in a corresponding format and rules, usually stored in a relational database. Structured data is usually data in two formats: numbers and numerical values. The first data can be data obtained after performing natural language processing (NLP) on the unstructured data in the data source. The classification result can be obtained after classifying the structured data according to predefined types. The second data can be data extracted from the classification result.

[0043] Specifically, for at least two data sources, determine the structured data and unstructured data associated with tobacco production scheduling in each data source. Use natural language processing techniques to perform text analysis and data extraction on the unstructured data of each data source to obtain the first data corresponding to each data source. Use machine learning techniques or clustering techniques to classify the structured data of each data source to obtain a classification result, and extract the second data from the classification result. Perform data preprocessing and data integration on the first data and the second data to obtain the first relationship information between the tobacco production scheduling information of each data source and the tobacco production scheduling information.

[0044] Exemplarily, from data sources such as a database storing business data of tobacco production scheduling, processing technology documents of tobacco products, etc., determine at least two heterogeneous data sources associated with tobacco production scheduling. For each data source, determine the structured data and unstructured data in each data source. Use natural language processing techniques to perform entity recognition and relationship extraction on the unstructured data to obtain knowledge elements such as entities, relationships, and attributes, that is, the first data. For the structured data, use machine learning techniques or clustering techniques to classify the structured data to obtain a classification result, and determine the rules or data related to tobacco production scheduling from the classification result, that is, the second data. Integrate the first data and the second data to determine the tobacco production scheduling information and the first relationship information.

[0045] S120. Based on the tobacco production scheduling information and the first relationship information respectively corresponding to each data source, determine the initial tobacco knowledge graph respectively corresponding to each data source, where the tobacco production scheduling information is used as the initial entity information in the initial tobacco knowledge graph.

[0046] Among them, the initial tobacco knowledge graph can be constructed based on the tobacco production scheduling information and the first relationship information and is used to represent the association relationship between tobacco production scheduling information. The tobacco production scheduling information can be used as the initial entity information for constructing the initial tobacco knowledge graph. The first relationship information can be used as the relationship information for constructing the initial tobacco knowledge graph.

[0047] Specifically, for at least two data sources, according to the predefined ontology information associated with tobacco production scheduling, use the tobacco production scheduling information corresponding to the current data source as the initial entity information, perform mapping processing on the initial entity information and the first relationship information to obtain the standardized triple corresponding to the current data source. Perform semantic analysis on the standardized triple to determine the entity type and relationship type corresponding to the current data source. Construct the initial tobacco knowledge graph corresponding to the current data source according to the entity type, relationship type, and the standardized triple.

[0048] In an embodiment of the present invention, the method for determining the initial tobacco knowledge graph may be as follows: for at least two data sources, map the tobacco production scheduling information and the first relationship information of the current data source into triples to be processed; perform semantic analysis processing on the triples to be processed based on a pre-trained type annotation model to determine the initial entity type corresponding to the tobacco production scheduling information and the relationship type corresponding to the first relationship information; determine the initial tobacco knowledge graph of the current data source based on the initial entity type, the relationship type, and the triples to be processed.

[0049] Among them, the triples to be processed are a preliminary graph structure constructed by the tobacco production scheduling information and the first relationship information, and are used to represent the semantic relationship between the tobacco production scheduling information. Optionally, the triples to be processed may be standardized Resource Description Framework (RDF) triples. The pre-trained type annotation model may be an entity and relationship type annotation model based on a graph neural network. The type annotation model is used to determine the entity type of the tobacco production scheduling information, that is, the initial entity type. The type annotation model may be used to determine the relationship type corresponding to the first relationship information.

[0050] Specifically, for at least two data sources, map the tobacco production scheduling information and the first relationship information of the current data source into standardized RDF triples, that is, the triples to be processed. Perform entity and relationship type annotation processing on the triples to be processed according to the pre-trained type annotation model to determine the initial entity type corresponding to the tobacco production scheduling information and the relationship type corresponding to the first relationship information. According to the initial entity type, the relationship type, and the triples to be processed corresponding to the current data source, the initial tobacco knowledge graph corresponding to the current data source can be constructed.

[0051] Optionally, the determination method of the initial entity type and relationship type can be as follows: Process the triple to be processed through at least one graph convolutional layer of a pre-trained type annotation model to determine the entity encoding information corresponding to the tobacco production scheduling information and the relationship encoding information corresponding to the first relationship information; Process the entity encoding information and the relationship encoding information through the attention layer of the type annotation model to determine the entity attention parameters corresponding to the entity encoding information and the relationship attention parameters corresponding to the relationship encoding information; Perform splicing processing on the entity encoding information and the entity attention parameters to obtain the entity feature to be processed corresponding to each entity encoding information, and perform splicing processing on the relationship encoding information and the relationship attention parameters to obtain the relationship feature to be processed corresponding to the relationship encoding information; Determine the entity similarity information between each entity feature to be processed and the preset entity semantic features of at least one preset entity type, and based on the entity similarity information, determine the initial entity type of the entity feature to be processed; And, determine the relationship type corresponding to the relationship feature to be processed according to the relationship similarity information between the relationship feature to be processed and the preset relationship semantic features of at least one preset relationship type.

[0052] Among them, the pre-trained type annotation model can at least include: at least one graph convolutional layer, an attention layer, a feature splicing layer, a fully connected layer, a feature extraction layer, etc. At least one graph convolutional layer is used to perform convolutional processing on the triple to be processed to obtain entity encoding information and relationship encoding information. The entity encoding information can be obtained by encoding the tobacco production scheduling information in the triple to be processed using the graph convolutional layer. Correspondingly, the relationship encoding information is obtained by encoding the first relationship information in the triple to be processed using the graph convolutional layer.

[0053] The attention layer can be used to aggregate the neighbor information of the current entity encoding information and the current relationship encoding information through an attention mechanism and an attention mask. The neighbor information can be other tobacco production scheduling information or first relationship information directly associated with the tobacco production scheduling information corresponding to the current entity encoding information. The entity attention parameter and the relationship attention parameter can be the attention aggregation parameters output by the attention layer. The entity attention parameter is the attention aggregation parameter corresponding to the entity encoding information. The relationship attention parameter is the attention aggregation parameter corresponding to the relationship encoding information.

[0054] The feature splicing layer is used to perform feature splicing processing on the entity encoding information and the entity attention parameters, and perform feature splicing processing on the relationship encoding information and the relationship attention parameters. The fully connected layer is used to process the entity splicing features output by the feature splicing layer to obtain the entity features to be processed corresponding to the entity splicing features, and process the relationship splicing features to obtain the relationship features to be processed.

[0055] The feature extraction layer is used to determine the entity similarity information between the entity features to be processed and the preset entity semantic features of the preset entity types, and to determine the relationship similarity information between the relationship features to be processed and the preset relationship semantic features of the preset relationship types. Among them, the preset entity types can be entity types preset according to actual needs. The preset entity semantic features can be features obtained by extracting features from the preset entity types. Optionally, the preset entity semantic features can be used to represent the semantic central features of the preset entity types. The preset relationship types can be relationship types set according to actual needs. The preset relationship semantic features can be features obtained by extracting features from the preset relationship types. Optionally, the preset relationship semantic features can be used to represent the semantic central features of the preset relationship types. The entity similarity information can be used to represent the similarity degree between the entity features to be processed and each preset entity semantic feature. The relationship similarity information can be used to represent the similarity degree between the relationship features to be processed and each preset relationship semantic feature. The initial entity type can be the preset entity type corresponding to the maximum entity similarity information determined from at least one entity similarity information. Correspondingly, the relationship type can be the preset relationship type corresponding to the maximum relationship similarity information determined from at least one relationship similarity information.

[0056] Specifically, at least one graph convolutional layer in the pre-trained type annotation model is used to perform convolutional processing on the input triple to be processed, to determine the entity encoding information corresponding to the tobacco production scheduling information and the relationship encoding information corresponding to the first relationship information. According to the attention mechanism and attention weights in the attention layer of the type annotation model, attention aggregation processing is performed on the entity encoding information and the relationship encoding information to determine the entity attention parameters corresponding to the entity encoding information and the relationship attention parameters corresponding to the relationship encoding information. Through the feature splicing layer of the type annotation model, feature splicing processing is performed on the entity encoding information and the entity attention parameters to obtain the entity features to be processed corresponding to each entity encoding information. And, feature splicing processing is performed on the relationship encoding information and the relationship attention parameters to obtain the relationship features to be processed corresponding to each relationship encoding information.

[0057] Calculate the entity similarity information between each entity feature to be processed and the preset entity semantic features corresponding to at least one preset entity type. From at least one entity similarity information, determine the maximum entity similarity information, and use the preset entity type corresponding to the maximum entity similarity information as the initial entity type of the current entity feature to be processed. Correspondingly, calculate the relationship similarity information between each relationship feature to be processed and the preset relationship semantic features corresponding to at least one preset relationship type. From at least one relationship similarity information, determine the maximum relationship similarity information, and use the preset relationship type corresponding to the maximum relationship similarity information as the initial relationship type of the current relationship feature to be processed.

[0058] Exemplarily, taking the pre-trained type annotation model as an entity and relationship type annotation model based on a graph convolutional neural network for illustration, taking the tobacco production scheduling information as entity v, the first relationship information as relationship e, all the tobacco production scheduling information corresponding to the current data source as entity set V, and all the first relationship information as relationship set E for illustration. Correspondingly, the triple to be processed can be represented by G=(V, E).

[0059] Let the triple G=(V, E) to be processed corresponding to the current data source, where V is the entity set, E is the relationship set, and each entity v∈V and each relationship e∈E respectively correspond to a feature vector.

[0060] For the Graph Convolutional Networks (GCN), the processing process of the l-th layer of graph convolutional layer is as follows:

[0061]

[0062] where, H (l) represents the output feature of the l-th layer of graph convolutional layer. If the l-th layer is the last graph convolutional layer, then H (l) represents the entity encoding information. σ is the activation function. Optionally, the activation function can be the ReLU activation function. represents the sum of the adjacency matrix A and the self-connection matrix I of the triple to be processed. The adjacency matrix A is determined according to the connection relationship between the entities corresponding to the triple to be processed. If there is a connection between entity i and entity j, the value of the corresponding element in the matrix is 1, otherwise it is 0. The values of the elements on the diagonal of the self-connection matrix I are 1, and the elements in other positions are 0. is the degree matrix of, where the degree matrix is a diagonal matrix, the value of the i-th element in the matrix is the degree of entity i in the triple to be processed. Degree represents the number of edges connected to entity i. W represents the preset edge weight matrix. ⊙ is the Hadamard product operation, represents the matrix multiplied element by element with the matrix W. T represents the entity type embedding matrix, and the T matrix is determined according to the entity embedding vectors corresponding to at least one preset entity type. is the broadcast addition operation, W (l) represents the preset weight matrix of the l-th layer of graph convolutional layer, H (l-1) represents the output feature of the (l - 1)-th layer of graph convolutional layer. It should be noted that in the graph convolutional neural network, due to the introduction of the edge weight matrix W and the entity type embedding matrix T, the relationship strength and type information between the entities corresponding to the triple to be processed can be better captured.

[0063] After obtaining the entity encoding information and relationship encoding information through the above at least one graph convolutional layer, the neighbor information can be aggregated through the attention layer. Among them, the neighbor information can be used to represent other entities and relationships connected to the current entity.

[0064] For entity v, the set of neighbor entities in the neighbor information is N v , the neighbor entity is u, u ∈ N v , the set of neighbor relationships is R v , the neighbor relationship is r, r ∈ R v , then the attention parameter of entity v can be expressed as follows:

[0065]

[0066] Among them, represents the attention parameter of entity v, h u represents the entity encoding information of neighbor entity u, h r represents the relationship encoding information of neighbor relationship r, α vu represents the attention weight of neighbor entity u, β vr represents the attention weight of neighbor relationship r, M vu represents the attention mask of neighbor entity u, M vr represents the attention mask corresponding to neighbor relationship r. Among them, the attention mask can be set according to the preset tobacco scheduling rules. For example, for the device parameters related to the quality index information, the corresponding neighbor information can be set with a higher attention mask value.

[0067] Among them, the attention weights α vu and β vr can be calculated in the following way.

[0068]

[0069] Among them, f() and g() are preset attention scoring functions, h v represents the entity encoding information of entity v, h u represents the entity encoding information of neighbor entity u, P vu represents the preset prior embedding vector of neighbor entity u, h u′ represents belonging to the set of neighbor entities N v of other neighbor entities u', h vu′ represents the prior embedding vector of other neighbor entities u'. h r represents the relationship encoding information of neighbor relationship r, P vr represents the preset prior embedding vector of neighbor relationship r, h r′ represents the relationship encoding information of other neighbor relationships r', P vr′Indicates belonging to the neighbor relationship set R v The prior embedding vector of other neighbor relationships r'.

[0070] The entity encoding information h of entity v v and the attention parameter are concatenated and processed through a fully connected layer to obtain the entity feature z to be processed for entity v v , and the specific process is as follows:

[0071]

[0072] It should be noted that the entity encoding information h v and the attention parameter are vector information. When concatenating, the above two vector informations are connected end to end to form a new feature vector z v . Correspondingly, the determination method of the relationship feature to be processed is similar to the determination method of the entity feature to be processed above, which will not be elaborated here. Below, the preset entity semantic feature of the preset entity type is p c , and the preset relationship semantic feature of the preset relationship type is p r is used for illustration.

[0073] For entity v, calculate the entity similarity information between the entity feature z to be processed for entity v v and the preset entity semantic feature p corresponding to the preset entity type c c .

[0074]

[0075] Among them, sim(z v , p c ) represents the entity similarity information, c' represents other preset entity types except the current preset entity type, C represents the set of preset entity types, that is, at least one preset entity type mentioned in the above embodiments, and S cc′ represents the correlation weight between the preset entity type c and other preset entity types c', and p c′ represents the preset entity semantic feature of other preset entity types c'.

[0076] Determine the preset entity type c with the largest entity similarity information * as the initial entity type of the entity feature to be processed, that is,

[0077]

[0078] Among them, c * represents the initial entity type of the entity feature to be processed, z v represents the entity feature to be processed of entity v, and p cRepresents the preset entity semantic features. It should be noted that the method of the relationship type is similar to the method for determining the above-mentioned initial entity type, which will not be elaborated here.

[0079] S130. Perform entity alignment processing on the initial entity information in at least two initial tobacco knowledge graphs, determine the entity information groups corresponding to the at least two initial tobacco knowledge graphs, and the target entity information corresponding to the entity information groups.

[0080] Among them, the entity information group includes initial entity information with a matching relationship. The target entity information can be determined from the entity information group and is used to represent the entity information of the entity information group.

[0081] Specifically, for at least two initial tobacco knowledge graphs, perform entity alignment processing on the initial entity information in every two initial tobacco knowledge graphs, determine whether there is a matching relationship between each initial entity information in every two initial tobacco knowledge graphs and other initial entity information, so as to obtain at least one entity information group corresponding to each initial entity information. According to the at least one entity information group corresponding to the current initial entity information in the at least two initial tobacco knowledge graphs, determine the target entity information corresponding to the current initial entity information degree. Based on this, obtain at least one target entity information from the at least two initial tobacco knowledge graphs.

[0082] S140. Determine the second relationship information of the target entity information in at least two data sources and the third relationship information in the initial tobacco knowledge graph. Based on the first relationship information, the second relationship information, and the third relationship information, determine the target relationship information corresponding to the target entity information.

[0083] Among them, the second relationship information can be the relationship information associated with the target entity information determined from at least two data sources. The third relationship information can be the relationship information corresponding to the target entity information determined from at least two initial tobacco knowledge graphs.

[0084] Specifically, after determining the target entity information, determine the second relationship information corresponding to the target entity information from at least two data sources. And determine the third relationship information corresponding to the target entity information from at least two initial tobacco knowledge graphs. According to the first relationship information, the second relationship information, and the third relationship information, determine the total relationship information. Perform clustering processing on the total relationship information to obtain at least two groups of relationship information combinations. Perform abstraction processing on each group of relationship information combinations to obtain the target relationship information corresponding to each group of relationship information combinations. Based on this, obtain at least two target relationship information corresponding to the target entity information.

[0085] S150. Based on the target entity information and the target relationship information, fuse at least two initial tobacco knowledge graphs to determine the target tobacco knowledge graph.

[0086] Among them, the target tobacco knowledge graph can be a knowledge graph obtained by fusing at least two initial tobacco knowledge graphs.

[0087] Specifically, according to the target entity information and target relationship information, at least two initial tobacco knowledge graphs corresponding to the tobacco production scheduling field are fused to obtain a target tobacco knowledge graph corresponding to the tobacco production scheduling field.

[0088] Optionally, after obtaining the target tobacco knowledge graph, the method further includes: when obtaining incremental production scheduling information, determining at least one candidate linked entity matching the incremental production scheduling information from the target tobacco knowledge graph; for at least one candidate linked entity, determining the entity similarity result and relationship similarity result between the incremental production scheduling information and each candidate linked entity; based on a preset link evaluation function, the entity similarity result corresponding to each candidate linked entity, and the relationship similarity result, determining the link evaluation attribute corresponding to each candidate linked entity; based on at least one link evaluation attribute, determining the target linked entity, and linking the incremental production scheduling information to the target tobacco knowledge graph based on the target linked entity to obtain an updated target tobacco knowledge graph.

[0089] Among them, the incremental production scheduling information can be newly obtained production scheduling information different from the tobacco production scheduling information. The candidate linked entity can be an entity in the target tobacco knowledge graph that is relevant to the incremental production scheduling information. Optionally, the candidate linked entity can be determined according to the semantic similarity between each entity in the target tobacco knowledge graph and the incremental production scheduling information or preset ontology information, or can be determined by the matching degree of keywords between each entity in the target tobacco knowledge graph and the incremental production scheduling information. The specific determination method of the candidate linked entity in this embodiment is not limited.

[0090] The entity similarity result can be the concept level similarity between the incremental production scheduling information and the candidate linked entity determined according to the preset entity concept hierarchy corresponding to the predefined ontology information. The relationship similarity result can be the relationship similarity between the incremental production scheduling information and the candidate linked entity determined according to the preset relationship constraint information corresponding to the predefined ontology information. Among them, the preset entity concept hierarchy can be understood as the upper and lower position relationships between various entities in the tobacco production scheduling field. For example, tobacco leaves are the lower-level entities of raw materials. The preset relationship constraint information is the possible relationship types and relationship restrictions between different entities. For example, the planting relationship only exists between farmers and tobacco leaves. It should be noted that the preset entity concept hierarchy and preset relationship constraint information are both predefined according to actual needs.

[0091] The preset link evaluation function can be used to determine the link possibility between the incremental production scheduling information and the candidate link entity. The link evaluation attribute can be used to characterize the link possibility between the incremental production scheduling information and the candidate link entity. The target link entity can be the candidate link entity with the largest link evaluation attribute.

[0092] Specifically, when the incremental production scheduling information is obtained, at least one candidate link entity that matches the incremental production scheduling information is determined from the target tobacco knowledge graph. According to the preset entity concept hierarchy corresponding to the predefined ontology information, the entity similarity result between the incremental production scheduling information and each candidate link entity is determined. According to the preset relationship constraint information corresponding to the predefined ontology information, the relationship similarity result between the incremental production scheduling information and each candidate link entity is determined. For each candidate link entity, based on the preset link evaluation function, the entity similarity result and the relationship similarity result of each candidate link entity are processed to obtain the link evaluation attribute corresponding to each candidate link entity. The maximum link evaluation attribute is determined from at least one link evaluation attribute, and the candidate link entity corresponding to the maximum link evaluation attribute is used as the target link entity. The entity corresponding to the incremental production scheduling information is determined, and the entity corresponding to the incremental production scheduling information is linked to the target tobacco knowledge graph through the target link entity to update the target tobacco knowledge graph and obtain the updated target tobacco knowledge graph. Optionally, a knowledge link model based on knowledge graph embedding and attention mechanism can be used to implement the link processing of the incremental production scheduling information to update the target tobacco knowledge graph. Based on this, the incremental update of the target tobacco knowledge graph is realized, and the problem of reconstructing the target tobacco knowledge graph every time new incremental production scheduling information is introduced is avoided.

[0093] Optionally, when the incremental relationship information is obtained, the relationship similarity result between the incremental relationship information and each candidate link relationship can also be determined in the above manner to update and adjust the candidate link relationship according to the relationship similarity result, and the updated target tobacco knowledge graph is obtained.

[0094] Exemplarily, taking the incremental production scheduling information as x, the candidate link entity as y, and at least one candidate link entity as the candidate link entity set C x , y ∈ C x .

[0095] The attention unit in the knowledge link model based on knowledge graph embedding and attention mechanism determines the first similarity information α between the incremental production scheduling information x and each candidate link entity y x,y .

[0096] Among them, the first similarity information α x,y can be determined by the following function.

[0097]

[0098] Among them, e x represents the embedding vector of the incremental production scheduling information x, and e y represents the embedding vector of the candidate link entity y, and y' represents the other candidate link entities in the candidate link entity set C x except the candidate link entity y. e y' represents the embedding vector of the other candidate link entity y'.

[0099] It should be noted that the above embedding vectors can be determined by the knowledge graph embedding method. For example, for an entity or a relationship, its corresponding embedding vector can be determined by the knowledge graph embedding method. For an entity v ∈ V, where V is the set of entities, the embedding vector corresponding to the entity v can be represented as e v . For a relationship e ∈ E, where E is the set of relationships, the embedding vector corresponding to the relationship e can be represented as r e .

[0100] According to the gated unit in the knowledge link model, the preset entity concept hierarchy and the preset relationship constraint information corresponding to the preset ontology information are integrated into the incremental linking process. That is, for each candidate link entity y, calculate its concept similarity cs(x, y) with the incremental production scheduling information x at the ontology level, that is, the entity similarity result mentioned in the above embodiment. And, the relationship similarity result rs(x, y) between each candidate link entity y and the incremental production scheduling information x at the ontology level. It should be noted that the entity similarity result cs(x, y) is used to characterize whether the relationship between the candidate link entity y and the incremental production scheduling information x conforms to the preset entity concept hierarchy defined by the ontology information. The relationship similarity result rs(x, y) is used to characterize whether the relationship between the candidate link entity y and the incremental production scheduling information x conforms to the preset relationship constraint information defined by the ontology information.

[0101] Use the gated unit to determine the second similarity information g between the incremental production scheduling information x and each candidate link entity y.

[0102] g = σ(W g · [e x ; e y ; cs(x, y); rs(x, y)] + b g )

[0103] Among them, σ is the activation function, and W g and b g represent the preset parameters of the gated unit.

[0104] The first similarity information α x,y, substitute the second similarity information g, the entity similarity result cs(x,y), and the relationship similarity result rs(x,y) into the preset link evaluation function to obtain the adjusted weight parameter

[0105]

[0106] Among them, represents the adjusted weight parameter, g represents the second similarity information, and α x,y represents the first similarity information, cs(x,y) represents the entity similarity result, and rs(x,y) represents the relationship similarity result.

[0107] Based on the adjusted weight parameter determine the link evaluation attribute between the current candidate link entity and the incremental production scheduling information. Take the candidate link entity corresponding to the maximum link evaluation attribute as the target link entity, and link the entity corresponding to the incremental production scheduling information to the target tobacco knowledge graph through the target link entity.

[0108] The technical solution of this embodiment is directed to at least two data sources associated with tobacco production scheduling. Data extraction processing is performed on each data source to obtain the first relationship information between the tobacco production scheduling information corresponding to each data source and the tobacco production scheduling information. Based on the tobacco production scheduling information and the first relationship information corresponding to each data source, the initial tobacco knowledge graph corresponding to each data source is determined. Based on this, the tobacco knowledge graph corresponding to each data source associated with tobacco production scheduling is determined, providing data support for the construction of the subsequent target tobacco knowledge graph. Entity alignment processing is performed on the initial entity information in at least two initial tobacco knowledge graphs to obtain an entity information group corresponding to at least two initial tobacco knowledge graphs and the target entity information corresponding to the entity information group. Based on this, the entity information required for constructing the target tobacco knowledge graph is accurately determined. Determine the second relationship information of the target entity information in at least two data sources and the third relationship information in the initial tobacco knowledge graph, so as to perform relationship fusion processing according to the first relationship information, the second relationship information, and the third relationship information to obtain the target relationship information. Based on the target entity information and the target relationship information, at least two initial tobacco knowledge graphs are fused to obtain the target tobacco knowledge graph. The present invention solves the problems in the prior art of low efficiency in constructing a knowledge graph and inability to accurately determine the relationship information between entities, and by performing entity and relationship fusion processing on the initial tobacco knowledge graphs of at least two data sources, solves the problem in the prior art that the constructed knowledge graph is inaccurate due to the heterogeneity and ambiguity corresponding to heterogeneous data sources, realizes the accurate construction of the tobacco knowledge graph in the field of tobacco production scheduling, and provides technical guidance for subsequent tobacco production scheduling.

[0109] Embodiment 2

[0110] Figure 2 FIG. is a flowchart of a method for determining a knowledge graph applied to tobacco production scheduling provided in Embodiment 2 of the present invention. This embodiment of the present invention is a preferred embodiment of the above-mentioned invention embodiment. The specific implementation manner can refer to the technical solution of this embodiment. Among them, the same or corresponding technical terms as those in the above embodiment will not be described in detail here.

[0111] As Figure 2 shown, the method includes:

[0112] S210. For at least two data sources associated with tobacco production scheduling, perform data extraction processing on each data source to obtain the tobacco production scheduling information corresponding to each data source and the first relationship information between the tobacco production scheduling information.

[0113] S220. Based on the tobacco production scheduling information and the first relationship information corresponding to each data source, determine the initial tobacco knowledge graph corresponding to each data source, where the tobacco production scheduling information is used as the initial entity information in the initial tobacco knowledge graph.

[0114] S230. Perform entity alignment processing on the initial entity information in at least two initial tobacco knowledge graphs to determine the entity information group corresponding to at least two initial tobacco knowledge graphs and the target entity information corresponding to the entity information group.

[0115] Among them, the entity information group includes initial entity information with a matching relationship.

[0116] In the embodiment of the present invention, the manner of determining the entity information group and the target entity information may be: based on at least two initial tobacco knowledge graphs, determine a plurality of initial entity information; process the plurality of initial entity information through the embedding layer of a pre-trained entity alignment model to determine the entity embedding vector corresponding to each initial entity information; process the plurality of entity embedding vectors through the attention layer of the entity alignment model to determine the weighted embedding vector corresponding to each initial entity information; based on the similarity evaluation function in at least one dimension, determine the similarity result between every two weighted embedding vectors, where at least one dimension includes at least one of the embedding similarity dimension, the structure similarity dimension, and the semantic similarity dimension; based on the similarity result, determine the entity information group and the target entity information corresponding to the entity information group.

[0117] Among them, the initial entity information can be the entity information corresponding to at least two initial tobacco knowledge graphs. The entity alignment model is used to perform entity alignment processing on the initial entity information. Optionally, the entity alignment model can be an entity alignment model based on the knowledge graph embedding method and the attention mechanism. The embedding layer in the entity alignment model is used to determine the entity embedding vector corresponding to each initial entity information. The entity embedding vector can be understood as the feature vector corresponding to the initial entity information. Optionally, the embedding layer can use the knowledge graph embedding method to determine the entity embedding vector. That is, at least one of the algorithms such as TransE, TransR, TransH, TransD, RotatE, etc. is used to process the initial entity information to determine the low-dimensional vector representation of the initial entity information, that is, the entity embedding vector.

[0118] The weighted embedding vector can be a feature vector obtained by weighting the entity embedding vector with the attention weight parameter of the attention layer. The similarity evaluation function in at least one dimension can be used to determine the similarity result between two weighted embedding vectors. At least one dimension includes at least one of the embedding similarity dimension, the structure similarity dimension, and the semantic similarity dimension. Among them, the similarity evaluation function in the embedding similarity dimension can be used to determine the cosine similarity between the weighted embedding vectors. The similarity evaluation function in the structure similarity dimension can be used to determine the neighbor entity set similarity between the weighted embedding vectors. The similarity evaluation function in the semantic similarity dimension can be used to determine the semantic similarity between the weighted embedding vectors. The similarity result can be a comprehensive similarity result determined based on at least one similarity determined by the similarity evaluation function in at least one dimension.

[0119] Specifically, multiple initial entity information between at least two initial tobacco knowledge graphs is determined. The multiple initial entity information is processed according to the embedding layer of the pre-trained entity alignment model to determine the entity embedding vector corresponding to each initial entity information. The multiple entity embedding vectors are processed according to the attention layer in the entity alignment model to determine the weighted embedding vector corresponding to each entity embedding vector. According to the similarity evaluation function in at least one dimension, the to-be-fused similarity result of each two weighted embedding vectors in at least one dimension among the embedding similarity dimension, the structural similarity dimension, and the semantic similarity dimension is determined. According to the preset weight coefficient and the to-be-fused similarity result in at least one dimension, the similarity result of each two weighted embedding vectors is determined. According to at least one similarity result between the weighted embedding vector corresponding to the current initial entity information and the weighted embedding vectors corresponding to other initial entity information, the initial entity information corresponding to the maximum similarity result is determined as the aligned entity information corresponding to the current initial entity information, that is, the aligned entity information and the current initial entity information are determined as the entity information group corresponding to the current initial entity information. The entity information group is abstracted to determine the target entity information corresponding to the entity information group.

[0120] Optionally, for at least two initial tobacco knowledge graphs, the entity alignment model can be used to align the initial entity information in each two initial tobacco knowledge graphs to determine the entity alignment relationship between each two initial tobacco knowledge graphs, so as to determine the entity alignment relationship between at least two initial tobacco knowledge graphs, and thus determine the entity information group and the target entity information corresponding to each initial entity information in at least two initial tobacco knowledge graphs.

[0121] The specific method can be: for at least two initial tobacco knowledge graphs, determine the multiple initial entity information corresponding to each two initial tobacco knowledge graphs. The initial entity information corresponding to each two initial tobacco knowledge graphs is processed according to the embedding layer of the pre-trained entity alignment model to determine the entity embedding vector corresponding to each initial entity information. The entity embedding vectors corresponding to each two initial tobacco knowledge graphs are processed according to the attention layer of the entity alignment model to determine the weighted embedding vector corresponding to each entity embedding vector. According to the similarity evaluation function in at least one dimension, the similarity result between each two weighted embedding vectors is determined. According to the similarity result, determine the multiple entity information groups corresponding to each two initial tobacco knowledge graphs, and the target entity information corresponding to the multiple entity information groups. Based on this, the entity information group and the target entity information corresponding to at least two initial tobacco knowledge graphs are obtained.

[0122] Exemplarily, taking two initial tobacco knowledge graphs as examples for illustration, the two initial tobacco knowledge graphs corresponding to different data sources are denoted as G1 and G2. The set of initial entity information corresponding to the initial tobacco knowledge graph G1 is V1, and the initial entity information is v1, where v1 ∈ V1. The set of initial entity information corresponding to the initial tobacco knowledge graph G2 is V2, and the initial entity information is v2, where v2 ∈ V2.

[0123] For each initial tobacco knowledge graph, the embedding layer in the entity alignment model is used to determine the low-dimensional vector representation of each initial entity information, that is, the entity embedding vector. The entity embedding vector of the initial entity information v1 can be denoted as e v1 , and the entity embedding vector of the initial entity information v2 can be denoted as e v2 .

[0124] The attention layer based on the entity alignment model measures the importance of the initial entity information v1 and v2 in different feature dimensions.

[0125] The attention weight parameter of the initial entity information v1 and v2 in the k-th feature dimension can be determined by the following function.

[0126]

[0127] Among them, α k represents the attention weight parameter of the initial entity information v1 and v2 in the k-th feature dimension, [e v1 ; e v2 represents the feature concatenation processing of the entity embedding vectors e v1 and e v2 , k' represents the k'-th feature dimension. For example, if there are a total of k feature dimensions, then k' ∈ [0, k], represents the attention weight vector in the k'-th feature dimension. represents the attention weight vector in the k-th feature dimension. Among them, the attention weight vector is a model parameter, and each dimension has a corresponding weight vector to enable the entity alignment model to recognize the importance of each dimension for entity alignment.

[0128] According to the attention weight parameter α k , the weighted embedding vectors of the initial entity information v1 and v2 are determined.

[0129]

[0130] Among them, α k represents the attention weight parameter of the initial entity information v1 and v2 in the k-th feature dimension, e v1 represents the entity embedding vector of the initial entity information v1, Represents the weighted embedding vector of the initial entity information v1.

[0131]

[0132] Among them, α k Represents the attention weight parameter of the initial entity information v1 and v2 on the k-th feature dimension, and e v2 Represents the entity embedding vector of the initial entity information v2, Represents the weighted embedding vector of the initial entity information v2.

[0133] Based on the weighted embedding vectors of the initial entity information v1 and v2, calculate the similarity results of the initial entity information v1 and v2. That is, calculate the embedding similarity result, the structure similarity result, and the semantic similarity result respectively through the following formula.

[0134] Based on the similarity evaluation function in the embedding similarity dimension, determine the embedding similarity result between the weighted embedding vectors and The embedding similarity result between them.

[0135]

[0136] Among them, s emb (v1, v2) represents the embedding similarity result between the weighted embedding vectors and The embedding similarity result between them.

[0137] Based on the similarity evaluation function in the structure similarity dimension, determine the structure similarity result between the weighted embedding vectors and The structure similarity result between them.

[0138]

[0139] Among them, s struct (v1, v2) represents the structure similarity result between the weighted embedding vectors and The structure similarity result between them. N(v1) represents the set of neighbor entity information of the initial entity information v1. N(v2) represents the set of neighbor entity information of the initial entity information v2.

[0140] Based on the similarity evaluation function in the semantic similarity dimension, determine the semantic similarity result between the weighted embedding vectors and The semantic similarity result between them.

[0141]

[0142] Among them, s sem (v1, v2) represents the semantic similarity result between the initial entity information v1 and v2, The word embedding vector representing the initial entity information v1 The word embedding vector representing the initial entity information v2. Among them, the word embedding vector can be obtained by performing word embedding processing on the entity name of the initial entity information, that is, using a word embedding model to convert the entity name of the initial entity information into a dense vector representation, namely the word embedding vector.

[0143] Perform weighted summation processing on the embedding similarity result, the structure similarity result, and the semantic similarity result to obtain the similarity result.

[0144] s(v1, v2) = λ1s emb (v1, v2) + λ2s struct (v1, v2) + λ3s sem (v1, v2)

[0145] Among them, s(v1, v2) represents the similarity result of the initial entity information v1 and v2, and s emb (v1, v2) represents the weighted embedding vector and The embedding similarity result between them, λ1 is the weight coefficient of the embedding similarity result, and s struct (v1, v2) represents the weighted embedding vector and The structure similarity result between them, λ2 is the weight coefficient of the structure similarity result, and s sem (v1, v2) represents the semantic similarity result between the initial entity information v1 and v2, and λ3 represents the weight coefficient of the semantic similarity result.

[0146] For each initial entity information v1 ∈ V1, select the initial entity information v2 ∈ V2 with the largest similarity result as the aligned entity information, that is, determine the entity information group corresponding to the current initial entity information by combining the aligned entity information with the current initial entity information. Perform abstraction processing on the entity information group to determine the target entity information corresponding to the entity information group.

[0147] Optionally, the method for determining the entity information group and the target entity information according to the similarity result can be: for multiple initial entity information, determine at least one entity information group corresponding to the current initial entity information according to at least one similarity result between the current initial entity information and other initial entity information except the current initial entity information; in the case of multiple entity information groups corresponding to the current initial entity information, perform entity disambiguation processing on the multiple entity information groups to obtain the target entity information corresponding to the multiple entity information groups.

[0148] Among them, the situation of multiple entity information groups corresponding to the current initial entity information may be due to multiple maximum similarity results between the current initial entity information and other initial entity information except the current initial entity information.

[0149] Specifically, for multiple initial entity information, at least one similarity result between the current initial entity information and other initial entity information except the current initial entity information is used to determine the maximum similarity result, and the other initial entity information corresponding to the maximum similarity result is used as the alignment entity information of the current initial entity information. According to the current initial entity information and the alignment entity information, an entity information group corresponding to the current initial entity information is obtained. If there are multiple maximum similarity results, the current initial entity information corresponds to multiple entity information groups. Entity disambiguation processing can be performed on multiple entity information groups, and the target alignment entity information corresponding to the current initial entity information is determined according to the entity ambiguity degree and information entropy between each entity information group, so as to obtain the target entity information based on the current initial entity information and the target alignment entity information.

[0150] Exemplarily, when performing entity disambiguation processing on multiple entity information groups corresponding to the current initial entity information, the ambiguity degree between the current initial entity information and other initial entity information in the entity information group can be calculated first, that is, the reciprocal of the difference between the maximum similarity result and the second maximum similarity result of the current initial entity information.

[0151]

[0152] Among them, amb(v1) is used to represent the ambiguity degree of the initial entity information v1, and s(v1, v2) represents the similarity result between the initial entity information v1 and v2. v2′ represents other initial entity information except the initial entity information v2 in the initial entity information set V2, and s(v1, v2′) represents the similarity result between the initial entity information v1 and other initial entity information v2′.

[0153] Determine the information entropy of the initial entity information v1, and the calculation formula is as follows:

[0154]

[0155] Among them, H(v1) is used to represent the information entropy of the initial entity information v1, and s(v1, v2) represents the similarity result between the initial entity information v1 and v2. v2′ represents other initial entity information except the initial entity information v2 in the initial entity information set V2, and s(v1, v2′) represents the similarity result between the initial entity information v1 and other initial entity information v2′.

[0156] Select the initial entity information v2 with a high degree of ambiguity and a large information entropy as the target alignment entity information corresponding to the initial entity information v1, so as to obtain the target entity information based on the current initial entity information and the target alignment entity information.

[0157] Optionally, the initial entity information v2 with a high degree of ambiguity and a large information entropy can also be labeled and used as a training sample of the entity alignment model to retrain the entity alignment model, so as to update the model parameters with the newly added labeled data and improve the alignment performance. Repeat the training until the number of entities with a relatively large degree of ambiguity is lower than a preset threshold or a predetermined number of iteration rounds is reached.

[0158] S240. Determine the second relationship information of the target entity information in at least two data sources and the third relationship information in the initial tobacco knowledge graph, and determine the target relationship information corresponding to the target entity information based on the first relationship information, the second relationship information, and the third relationship information.

[0159] In the embodiment of the present invention, the method for determining the second relationship information may be: taking any two target entity information as a target entity pair; for the target entity information in any target entity pair, performing text extraction processing on at least two data sources to determine the text data corresponding to the target entity pair; processing the text data based on a pre-trained word embedding model to determine the word embedding vector of the text data, and determining the first feature based on the word embedding vector and the entity embedding vector of the target entity pair; determining the relationship embedding vector corresponding to the target entity pair based on at least two initial tobacco knowledge graphs, and using the relationship embedding vector as the second feature; splicing the first feature and the second feature to obtain the target feature of the target entity pair; inputting the target features of at least one target entity pair into a relationship classification model to obtain the second relationship information.

[0160] Among them, the target entity pair can be an entity pair determined according to any two target entity information. That is, the target entity pair contains two target entity information. The text data can be text data extracted from at least two data sources and associated with the target entity pair. The word embedding model is used to determine the word embedding vector corresponding to the text data. The first feature can be a feature vector obtained by vector splicing of the word embedding vector and the entity embedding vector. The relationship embedding vector can be an embedding vector determined by the relationship corresponding to the target entity pair in at least two initial tobacco knowledge graphs. The second feature can be the relationship embedding vector. The target feature can be understood as the feature after feature splicing of the first feature and the second feature. The relationship classification model can be used to determine the relationship information corresponding to the target entity pair in at least two data sources. The second relationship information can be the relationship information corresponding to the target entity pair in at least two data sources output by the relationship classification model.

[0161] Specifically, any two pieces of target entity information are taken as a target entity pair. For the target entity information in each target entity pair, text extraction processing is performed on at least two data sources according to the target entity information to obtain text data corresponding to the target entity information. The text data is input into a pre-trained word embedding model to convert the text data into a word embedding sequence. The word embedding sequence is encoded based on the sequence encoder in the word embedding model, and after combining the context information of the word embedding sequence, a word embedding vector corresponding to the word embedding sequence is output. An entity embedding vector corresponding to each piece of target entity information in the target entity pair is determined. The entity embedding vector is concatenated with the corresponding word embedding vector to obtain a first feature. For at least two initial tobacco knowledge graphs, a relationship embedding vector corresponding to the target entity pair is determined. The relationship embedding vector is used as a second feature and concatenated with the second feature to obtain a target feature. The target features corresponding to at least one target entity pair are input into a relationship classification model to obtain second relationship information.

[0162] Exemplarily, taking the two pieces of target entity information in the target entity pair as v1 and v2, and the text data as s for illustration. The text data s is converted into a word embedding sequence through a pre-trained word embedding model, and the word embedding sequence is encoded using the sequence encoder. After combining the context information of the word embedding sequence, a word embedding vector corresponding to the word embedding sequence is output. Among them, the word embedding vector corresponding to the target entity information v1 is denoted as The word embedding vector corresponding to the target entity information v2 is denoted as Determine the entity embedding vector corresponding to the target entity information v1 And the entity embedding vector corresponding to the target entity information v2 And perform a concatenation process with the corresponding word embedding vector. The specific calculation can be shown as follows.

[0163]

[0164] Among them, v1' represents the vector obtained by concatenating the word embedding vector And the entity embedding vector Corresponding to the first feature mentioned in the above embodiment, i1 represents the position index of the target entity information v1 in the text data s.

[0165]

[0166] Among them, v2' represents the vector obtained by concatenating the word embedding vector And the entity embedding vector Corresponding to the first feature mentioned in the above embodiment, i2 represents the position index of the target entity information v2 in the text data s.

[0167] Determine the relationship paths corresponding to the known target entity pairs from at least two initial tobacco knowledge graphs, that is, at least one relationship path p with a preset path length between the target entity information v1 and v2. Determine the relationship sequence corresponding to each relationship path p, and perform encoding processing on the relationship sequence to obtain a comprehensive path representation, that is, a relationship embedding vector. Concatenate v1', v2' and the relationship embedding vector to obtain the target feature. Use a multi-layer perceptron, that is, the relationship classification model mentioned in the above embodiment, to perform relationship classification on the target feature to obtain the relationship type between the target entity information v1 and v2, that is, the second relationship information. Based on this, the accuracy and generalization ability of relationship extraction are improved.

[0168] In the embodiment of the present invention, the manner of determining the target relationship information according to the first relationship information, the second relationship information, and the third relationship information may be: based on the first relationship information, the second relationship information, and the third relationship information, determine multiple relationship types to be fused; determine the semantic similarity information between the multiple relationship types to be fused, and based on the semantic similarity information, perform clustering processing on the multiple relationship types to be fused to obtain at least one relationship type combination; perform confidence evaluation processing on each relationship type combination based on a preset attribute evaluation function to determine the confidence evaluation result corresponding to each relationship type combination; abstract the relationship type combination whose confidence evaluation result meets the preset condition as the target relationship type, and based on the target relationship type, determine the target relationship information.

[0169] Among them, the relationship types to be fused may be the relationship types determined from the first relationship information, the second relationship information, and the third relationship. For example, based on the first relationship information, determine the relationship type A to be fused, based on the second relationship information, determine the relationship type B to be fused, and based on the third relationship information, determine the relationship type C to be fused, so as to obtain multiple relationship types to be fused A + B + C. The semantic similarity information can be used to represent the semantic similarity between the relationship types to be fused. The relationship type combination may be obtained by performing clustering processing on the relationship types to be fused. Each relationship type combination contains at least one relationship type to be fused. The preset attribute evaluation function may be preset and used to evaluate the confidence of the relationship type combination. The confidence evaluation result is used to represent the confidence of the relationship type combination. Optionally, the confidence evaluation result may be a confidence score, the preset condition may be preset, and whether the confidence score is higher than the preset confidence score. The target relationship type may be a relationship type combination whose confidence evaluation result meets the preset condition, that is, a relationship combination with a confidence score higher than the preset confidence score.

[0170] Specifically, based on the first relationship information, the second relationship information, and the third relationship information, determine multiple relationship types to be fused. Determine the semantic similarity information between each pair of relationship types to be fused. According to the semantic similarity information between each pair of relationship types to be fused, perform clustering processing on the multiple relationship types to be fused, and determine relationship types to be fused with similar semantics as a relationship type combination. Use a preset attribute evaluation function to perform a confidence evaluation process on at least one relationship type combination to determine the confidence evaluation result corresponding to each relationship type combination. When the confidence evaluation result meets the preset conditions, abstract this relationship type combination into a target relationship type. Obtain the target relationship information according to the target relationship type.

[0171] Exemplarily, based on the first relationship information, the second relationship information, and the third relationship information, determine multiple relationship types to be fused {r1, r2,..., r k}, and determine the semantic similarity information between different relationship types to be fused.

[0172]

[0173] Among them, r i represents the i-th relationship type to be fused, r j represents the j-th relationship type to be fused, and sim(r i , r j ) represents the semantic similarity information between the i-th relationship type to be fused and the j-th relationship type to be fused.

[0174] Based on the semantic similarity information between every two relationship types to be fused, use the hierarchical clustering algorithm to cluster the relationship types to be fused, and merge the relationship types to be fused with similar semantics into a relationship type combination. The clustering result can be represented by a relationship type tree.

[0175] For each relationship type combination r, calculate its confidence score, that is, the confidence evaluation result mentioned in the above embodiment.

[0176]

[0177] Among them, r i represents the relationship type to be fused in the relationship type combination r, w i is preset, and is the weight coefficient corresponding to the data source to which the relationship type to be fused belongs. conf(r i ) represents the confidence score of the relationship type to be fused r i .

[0178] When the confidence score of the relationship type combination r is higher than the preset confidence score, abstract this relationship type combination r into a target relationship type to obtain the target relationship information.

[0179] Optionally, a global consistency optimization method can also be used to correct the target relationship type to obtain an updated target relationship type, and then determine the target relationship information based on the updated target relationship type. Among them, the global consistency optimization method can be at least one of a Markov logic network, integer linear programming, or rule reasoning.

[0180] Based on this, the effective integration of the first relationship information, the second relationship information, and the third relationship information is realized, redundant relationship information is reduced, and the quality and consistency of the relationship information are ensured.

[0181] S250. Based on the target entity information and the target relationship information, fuse at least two initial tobacco knowledge graphs to determine a target tobacco knowledge graph.

[0182] S260. Input at least one target entity information and / or target relationship information in the target tobacco knowledge graph, as well as a random noise vector, into a pre-trained triple generation network model for processing to obtain triples to be used corresponding to at least one target entity information and / or target relationship information.

[0183] Among them, the random noise vector can be a vector determined based on random Gaussian noise. The pre-trained triple generation network model can be obtained by training a pre-constructed generative adversarial network model with triples corresponding to the target entity information and target relationship information in the target tobacco knowledge graph and a random noise vector. The triples to be used can be different from the triples corresponding to the target entity information and target relationship information in the target tobacco knowledge graph. It can be understood that the triples to be used contain missing entity information or relationship information in the target tobacco knowledge graph. Optionally, the triples to be used can include entity information - relationship information - entity information.

[0184] Specifically, input at least one target entity information and / or target relationship information in the target tobacco knowledge graph, as well as a random noise vector, into a pre-trained triple generation network model to obtain triples to be used corresponding to at least one target entity information and / or target relationship information based on the triple generation network model.

[0185] S270. Based on the triples to be used, determine the missing entity information and / or missing relationship information in the target tobacco knowledge graph.

[0186] Among them, the missing entity information can be entity information not included in the target tobacco knowledge graph. The missing relationship information can be relationship information not included in the target tobacco knowledge graph.

[0187] Specifically, based on the entity information and relationship information corresponding to the triple to be used, determine the consistency between the entity information and relationship information in the triple to be used and the target entity information and target relationship information corresponding to the target tobacco knowledge graph, so as to obtain missing entity information different from the target entity information and / or missing relationship information different from the target relationship information.

[0188] S280. Based on the missing entity information and / or missing relationship information, perform an adjustment process on the target tobacco knowledge graph to obtain an adjusted target tobacco knowledge graph.

[0189] Specifically, based on the missing entity information and / or missing relationship information in the triple to be used and the target entity information and / or target relationship information in the triple to be used, determine the association relationship between the missing entity information and / or missing relationship information and the target entity information and / or target relationship information, so as to link the missing entity information and / or missing relationship information to the target tobacco knowledge graph based on this association relationship to obtain an adjusted target tobacco knowledge graph.

[0190] Exemplarily, to ensure the integrity of the target tobacco knowledge graph, the missing entity information and / or missing relationship information can be determined according to the triple generation network model to complete the target tobacco knowledge graph. Taking the triple generation network model as a generative adversarial network model for illustration. Among them, the generative adversarial network model includes a generator G and a discriminator D.

[0191] Input a random noise vector into the generator G to generate a virtual entity embedding vector e v' and a relationship embedding vector r e' . The discriminator D receives the real entity embedding vector e v 、the real relationship embedding vector r e , as well as the entity embedding vector e v' and the relationship embedding vector r e' generated by the generator, and discriminates whether the relationship embedding vector and the entity embedding vector are real or generated. Adjust the generator G according to the discrimination result output by the discriminator D. During this process, the generator G gradually learns to generate embedding vectors closer to the real distribution, and learns to capture the potential patterns and structures in the knowledge graph, establishing a mapping from the random noise space to the knowledge graph embedding space, so as to reflect the potential structure of the target tobacco knowledge graph. Based on the discrimination result output by the discriminator D, the generator G adjusts its parameters through backpropagation so that the generator G can generate real entity embedding vectors and relationship embedding vectors.

[0192] Among them, the generator G and the discriminator D are trained through a minimax game, as follows.

[0193]

[0194] Among them, G represents the generator, D represents the discriminator, and e v represents the real entity embedding vector, and r e represents the real relationship embedding vector. (e v , r e ) represents the real triple in the target knowledge graph, and p data represents the existing triple distribution (real data distribution) in the target tobacco knowledge graph. represents the expectation of the real data distribution. logD(e v , r e ) represents the logarithm of the discriminator's discrimination result for the real triple. z represents the random noise vector, and p z represents the noise distribution. represents the expectation of the noise distribution. G(z) represents the triple generated by the generator G based on the random noise vector. D(G(z)) represents the discriminator's discrimination result for the triple generated by the generator. log(1 - D(G(z))) represents the logarithmic probability that the discriminator judges the generated triple as false.

[0195] Through the above training, the generator G generates entity embedding vectors and relationship embedding vectors similar to the distribution of the target tobacco knowledge graph, while the discriminator D distinguishes between real and generated embedding vectors. After the training is completed, the generator G is used as a triple generation network model to determine the missing entity information or missing relationship information based on the input random noise vector and at least one target entity information and / or target relationship information, so as to link the missing entity information and / or missing relationship information to the target tobacco knowledge graph.

[0196] For the technical solution of this embodiment, for at least two data sources associated with tobacco production scheduling, data extraction processing is performed on each data source to obtain the tobacco production scheduling information corresponding to each data source and the first relationship information between the tobacco production scheduling information. According to the tobacco production scheduling information and the first relationship information corresponding to each data source, the initial tobacco knowledge graph corresponding to each data source is determined. Based on this, the tobacco knowledge graph corresponding to each data source associated with tobacco production scheduling is determined, providing data support for the construction of the subsequent target tobacco knowledge graph. Entity alignment processing is performed on the initial entity information in at least two initial tobacco knowledge graphs to obtain an entity information group corresponding to the at least two initial tobacco knowledge graphs and the target entity information corresponding to the entity information group. Based on this, the entity information required for constructing the target tobacco knowledge graph is accurately determined. The second relationship information of the target entity information in at least two data sources and the third relationship information in the initial tobacco knowledge graph are determined, so as to perform relationship fusion processing according to the first relationship information, the second relationship information, and the third relationship information to obtain the target relationship information. Based on the target entity information and the target relationship information, at least two initial tobacco knowledge graphs are fused to obtain the target tobacco knowledge graph. At least one target entity information and / or target relationship information in the target tobacco knowledge graph, and a random noise vector are input into a pre-trained triple generation network model for processing to obtain the triples to be used corresponding to the at least one target entity information and / or target relationship information. Based on the triples to be used, the missing entity information and / or missing relationship information in the target tobacco knowledge graph are determined. Based on the missing entity information and / or missing relationship information, the target tobacco knowledge graph is adjusted to obtain the adjusted target tobacco knowledge graph, realizing the completion processing of the target tobacco knowledge graph and ensuring the integrity of the target tobacco knowledge graph. The present invention solves the problems of low efficiency in constructing a knowledge graph and inability to accurately determine the relationship information between entities in the prior art, and solves the problem that the knowledge graph constructed in the prior art is inaccurate due to the heterogeneity and ambiguity corresponding to heterogeneous data sources by performing entity and relationship fusion processing on the initial tobacco knowledge graphs of at least two data sources, realizing the accurate construction of the tobacco knowledge graph in the field of tobacco production scheduling and providing technical guidance for subsequent tobacco production scheduling.

[0197] Embodiment III

[0198] Figure 3 It is a schematic structural diagram of a knowledge graph determination device applied to tobacco production scheduling provided by Embodiment III of the present invention. As Figure 3 shown, the device includes: a relationship information determination module 310, an initial tobacco knowledge graph determination module 320, an entity alignment module 330, a target relationship information determination module 340, and a target tobacco knowledge graph determination module 350.

[0199] A relationship information determination module 310, configured to perform data extraction processing on each of at least two data sources associated with tobacco production scheduling, so as to obtain first relationship information between the tobacco production scheduling information and the tobacco production scheduling information respectively corresponding to each data source; An initial tobacco knowledge graph determination module 320, configured to determine an initial tobacco knowledge graph respectively corresponding to each data source based on the tobacco production scheduling information and the first relationship information respectively corresponding to each data source, wherein the tobacco production scheduling information is used as initial entity information in the initial tobacco knowledge graph; An entity alignment module 330, configured to perform entity alignment processing on the initial entity information in at least two initial tobacco knowledge graphs, determine an entity information group corresponding to the at least two initial tobacco knowledge graphs, and target entity information corresponding to the entity information group, wherein the entity information group includes initial entity information with a matching relationship; A target relationship information determination module 340, configured to determine second relationship information of the target entity information in at least two data sources and third relationship information in the initial tobacco knowledge graph, and determine target relationship information corresponding to the target entity information based on the first relationship information, the second relationship information, and the third relationship information; A target tobacco knowledge graph determination module 350, configured to fuse at least two initial tobacco knowledge graphs based on the target entity information and the target relationship information to determine a target tobacco knowledge graph.

[0200] The technical solution of this embodiment is directed to at least two data sources associated with tobacco production scheduling. Data extraction processing is performed on each data source to obtain the tobacco production scheduling information corresponding to each data source and the first relationship information between the tobacco production scheduling information. According to the tobacco production scheduling information and the first relationship information corresponding to each data source, the initial tobacco knowledge graph corresponding to each data source is determined. Based on this, the tobacco knowledge graph corresponding to each data source associated with tobacco production scheduling is determined, providing data support for the subsequent construction of the target tobacco knowledge graph. Entity alignment processing is performed on the initial entity information in at least two initial tobacco knowledge graphs to obtain an entity information group corresponding to the at least two initial tobacco knowledge graphs and the target entity information corresponding to the entity information group. Based on this, the entity information required for constructing the target tobacco knowledge graph is accurately determined. The second relationship information of the target entity information in at least two data sources and the third relationship information in the initial tobacco knowledge graph are determined, so as to perform relationship fusion processing according to the first relationship information, the second relationship information, and the third relationship information to obtain the target relationship information. Based on the target entity information and the target relationship information, at least two initial tobacco knowledge graphs are fused to obtain the target tobacco knowledge graph. The present invention solves the problems in the prior art of low efficiency in constructing knowledge graphs and inability to accurately determine the relationship information between entities, and by performing entity and relationship fusion processing on the initial tobacco knowledge graphs of at least two data sources, solves the problem in the prior art that the constructed knowledge graph is inaccurate due to the heterogeneity and ambiguity corresponding to heterogeneous data sources, realizing the accurate construction of the tobacco knowledge graph in the field of tobacco production scheduling and providing technical guidance for subsequent tobacco production scheduling.

[0201] Based on the above embodiment, optionally, the relationship information determination module is configured to perform data extraction processing on the unstructured data in each data source to determine the first data corresponding to each data source, and perform classification processing on the structured data of each data source to obtain a classification result, and determine the second data of each data source according to the classification result; based on the first data and the second data, determine the tobacco production scheduling information corresponding to each data source and the first relationship information between the tobacco production scheduling information.

[0202] Optionally, the initial tobacco knowledge graph determination module includes: a relationship information mapping unit configured to map the tobacco production scheduling information and the first relationship information of the current data source into a triple to be processed for at least two data sources; a triple semantic analysis unit configured to perform semantic analysis processing on the triple to be processed based on a pre-trained type annotation model to determine the initial entity type corresponding to the tobacco production scheduling information and the relationship type corresponding to the first relationship information; an initial tobacco knowledge graph construction unit configured to determine the initial tobacco knowledge graph of the current data source based on the initial entity type, the relationship type, and the triple to be processed.

[0203] Optionally, a triple semantic analysis unit is configured to process a triple to be processed based on at least one graph convolutional layer of a pre-trained type annotation model, and determine entity encoding information corresponding to tobacco production scheduling information and relationship encoding information corresponding to first relationship information; process the entity encoding information and the relationship encoding information based on the attention layer of the type annotation model, and determine an entity attention parameter corresponding to the entity encoding information and a relationship attention parameter corresponding to the relationship encoding information; perform splicing processing on the entity encoding information and the entity attention parameter to obtain a to-be-processed entity feature corresponding to each entity encoding information, and perform splicing processing on the relationship encoding information and the relationship attention parameter to obtain a to-be-processed relationship feature corresponding to the relationship encoding information; determine entity similarity information between each to-be-processed entity feature and preset entity semantic features of at least one preset entity type, and determine an initial entity type of the to-be-processed entity feature based on the entity similarity information; and determine a relationship type corresponding to the to-be-processed relationship feature according to relationship similarity information between the to-be-processed relationship feature and preset relationship semantic features of at least one preset relationship type.

[0204] Optionally, an entity alignment module includes an initial entity information determination unit configured to determine a plurality of initial entity information based on at least two initial tobacco knowledge graphs; an entity embedding vector determination unit configured to process the plurality of initial entity information based on an embedding layer of a pre-trained entity alignment model, and determine an entity embedding vector corresponding to each initial entity information; a weighted embedding vector determination unit configured to process the plurality of entity embedding vectors based on the attention layer of the entity alignment model, and determine a weighted embedding vector corresponding to each initial entity information; a similarity result determination unit configured to determine a similarity result between every two weighted embedding vectors based on a similarity evaluation function in at least one dimension, where the at least one dimension includes at least one of an embedding similarity dimension, a structural similarity dimension, and a semantic similarity dimension; and a target entity information determination unit configured to determine an entity information group and target entity information corresponding to the entity information group based on the similarity result.

[0205] Optionally, the target entity information determination unit is configured to, for a plurality of initial entity information, determine at least one entity information group corresponding to the current initial entity information according to at least one similarity result between the current initial entity information and other initial entity information except the current initial entity information; and perform entity disambiguation processing on the plurality of entity information groups in the case of the plurality of entity information groups corresponding to the current initial entity information, to obtain target entity information corresponding to the plurality of entity information groups.

[0206] Optionally, the target relationship information determination module includes: a second relationship information determination unit, configured to use any two pieces of target entity information as a target entity pair; for the target entity information in any target entity pair, perform text extraction processing on at least two data sources to determine text data corresponding to the target entity pair; process the text data based on a pre-trained word embedding model to determine the word embedding vector of the text data, and determine a first feature based on the word embedding vector and the entity embedding vector of the target entity pair; determine a relationship embedding vector corresponding to the target entity pair based on at least two initial tobacco knowledge graphs, and use the relationship embedding vector as a second feature; splice the first feature and the second feature to obtain the target feature of the target entity pair; input the target features of at least one target entity pair into a relationship classification model to obtain second relationship information.

[0207] Optionally, the target relationship information determination module includes: a relationship fusion module, configured to determine multiple relationship types to be fused based on the first relationship information, the second relationship information, and the third relationship information; determine the semantic similarity information between the multiple relationship types to be fused, and perform clustering processing on the multiple relationship types to be fused based on the semantic similarity information to obtain at least one relationship type combination; perform confidence evaluation processing on each relationship type combination based on a preset attribute evaluation function to determine the confidence evaluation result corresponding to each relationship type combination; abstract the relationship type combination whose confidence evaluation result meets the preset condition into a target relationship type, and determine target relationship information based on the target relationship type.

[0208] Optionally, the apparatus further includes: a target tobacco knowledge graph update module, configured to, when obtaining incremental production scheduling information, determine at least one candidate link entity in the target tobacco knowledge graph that matches the incremental production scheduling information; for at least one candidate link entity, determine the entity similarity result and the relationship similarity result between the incremental production scheduling information and each candidate link entity; determine the link evaluation attribute corresponding to each candidate link entity based on a preset link evaluation function, the entity similarity result corresponding to each candidate link entity, and the relationship similarity result; determine a target link entity based on at least one link evaluation attribute, and link the incremental production scheduling information to the target tobacco knowledge graph based on the target link entity to obtain an updated target tobacco knowledge graph.

[0209] Optionally, the device further includes: a target tobacco knowledge graph adjustment module, configured to input at least one target entity information and / or target relationship information in the target tobacco knowledge graph, and a random noise vector, into a pre-trained triple generation network model for processing, to obtain triples to be used corresponding to the at least one target entity information and / or target relationship information; determine missing entity information and / or missing relationship information in the target tobacco knowledge graph based on the triples to be used; and adjust the target tobacco knowledge graph based on the missing entity information and / or missing relationship information, to obtain an adjusted target tobacco knowledge graph.

[0210] The knowledge graph determination device for tobacco production scheduling provided in the embodiments of the present invention can execute the knowledge graph determination method for tobacco production scheduling provided in any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0211] Embodiment 4

[0212] Figure 4 FIG. 10 is a schematic structural diagram of an electronic device provided in Embodiment 4 of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0213] As Figure 4 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0214] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0215] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the knowledge graph determination method applied to tobacco production scheduling.

[0216] In some embodiments, the knowledge graph determination method applied to tobacco production scheduling can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the knowledge graph determination method applied to tobacco production scheduling described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the knowledge graph determination method applied to tobacco production scheduling in any other suitable manner (e.g., by means of firmware).

[0217] The various embodiments of the systems and technologies described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0218] The computer program for implementing the method for determining a knowledge graph applied to tobacco production scheduling according to the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0219] Embodiment 5

[0220] Embodiment 5 of the present invention further provides a computer-readable storage medium storing computer instructions for causing a processor to execute a method for determining a knowledge graph applied to tobacco production scheduling, the method including:

[0221] For at least two data sources associated with tobacco production scheduling, performing data extraction processing on each data source to obtain the tobacco production scheduling information corresponding to each data source and the first relationship information between the tobacco production scheduling information; based on the tobacco production scheduling information and the first relationship information corresponding to each data source, determining the initial tobacco knowledge graph corresponding to each data source, wherein the tobacco production scheduling information is used as the initial entity information in the initial tobacco knowledge graph; performing entity alignment processing on the initial entity information in at least two initial tobacco knowledge graphs to determine the entity information group corresponding to at least two initial tobacco knowledge graphs and the target entity information corresponding to the entity information group, wherein the entity information group includes initial entity information having a matching relationship; determining the second relationship information of the target entity information in at least two data sources and the third relationship information in the initial tobacco knowledge graph, and based on the first relationship information, the second relationship information, and the third relationship information, determining the target relationship information corresponding to the target entity information; based on the target entity information and the target relationship information, fusing at least two initial tobacco knowledge graphs to determine the target tobacco knowledge graph.

[0222] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0223] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0224] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0225] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0226] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitations are imposed herein.

[0227] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A knowledge graph determination method for tobacco production scheduling, characterized in that: include: For at least two data sources associated with tobacco production scheduling, data extraction processing is performed on each data source to obtain tobacco production scheduling information corresponding to each data source and first relationship information between the tobacco production scheduling information; Based on the tobacco production scheduling information and the first relationship information respectively corresponding to each data source, determining the initial tobacco knowledge graph respectively corresponding to each data source, wherein the tobacco production scheduling information serves as the initial entity information in the initial tobacco knowledge graph; Performing entity alignment processing on the initial entity information in at least two of the initial tobacco knowledge graphs to determine entity information groups corresponding to at least two of the initial tobacco knowledge graphs, and target entity information corresponding to the entity information groups, wherein the entity information groups include initial entity information having a matching relationship; Determine the second relationship information of the target entity information in the at least two data sources and the third relationship information in the initial tobacco knowledge graph, and determine the target relationship information corresponding to the target entity information based on the first relationship information, the second relationship information and the third relationship information; Based on the target entity information and the target relationship information, at least two initial tobacco knowledge graphs are fused to determine a target tobacco knowledge graph.

2. The method according to claim 1, characterized in that The data extraction process is performed on each data source to obtain the tobacco production scheduling information corresponding to each data source and the first relationship information between the tobacco production scheduling information, including: Performing data extraction processing on the unstructured data in each data source to determine the first data corresponding to each data source, and performing classification processing on the structured data of each data source to obtain a classification result, and determining the second data of each data source according to the classification result; Based on the first data and the second data, first relationship information between the tobacco production scheduling information of each data source and the tobacco production scheduling information is determined.

3. The method according to claim 1, characterized in that The determining, based on the tobacco production scheduling information and the first relationship information respectively corresponding to each data source, the initial tobacco knowledge graph respectively corresponding to each data source comprises: For the at least two data sources, mapping the tobacco production scheduling information of the current data source and the first relationship information into a triplet to be processed; Performing semantic analysis on the triples to be processed based on a pre-trained type annotation model to determine the initial entity type corresponding to the tobacco production scheduling information and the relationship type corresponding to the first relationship information; Based on the initial entity type, the relationship type, and the triples to be processed, an initial tobacco knowledge graph of the current data source is determined.

4. The method according to claim 3, characterized in that The performing semantic analysis on the to-be-processed triples based on the pre-trained type annotation model to determine the initial entity type corresponding to the tobacco production scheduling information and the relationship type corresponding to the first relationship information includes: Processing the to-be-processed triples based on at least one graph convolutional layer of a pre-trained type annotation model to determine entity encoding information corresponding to the tobacco production scheduling information and relationship encoding information corresponding to the first relationship information; Processing the entity encoding information and the relationship encoding information based on the attention layer of the type annotation model to determine an entity attention parameter corresponding to the entity encoding information and a relationship attention parameter corresponding to the relationship encoding information; The entity encoding information and the entity attention parameter are concatenated to obtain the entity features to be processed corresponding to each of the entity encoding information, and the relationship encoding information and the relationship attention parameter are concatenated to obtain the relationship features to be processed corresponding to the relationship encoding information; Determine entity similarity information between each entity feature to be processed and a preset entity semantic feature of at least one preset entity type, and determine the initial entity type of the entity feature to be processed based on the entity similarity information; and determine the relationship type corresponding to the relationship feature to be processed based on the relationship similarity information between the relationship feature to be processed and a preset relationship semantic feature of at least one preset relationship type.

5. The method according to claim 1, characterized in that The performing entity alignment processing on the initial entity information in at least two of the initial tobacco knowledge graphs to determine entity information groups corresponding to at least two of the initial tobacco knowledge graphs and target entity information corresponding to the entity information groups includes: Determining a plurality of initial entity information based on at least two of the initial tobacco knowledge graphs; Processing the plurality of initial entity information based on an embedding layer of a pre-trained entity alignment model to determine an entity embedding vector corresponding to each of the initial entity information; Processing the plurality of entity embedding vectors based on the attention layer of the entity alignment model to determine a weighted embedding vector corresponding to each initial entity information; Determine a similarity result between each two weighted embedding vectors based on a similarity evaluation function in at least one dimension, wherein the at least one dimension includes at least one of an embedding similarity dimension, a structural similarity dimension, and a semantic similarity dimension; Based on the similarity result, an entity information group and target entity information corresponding to the entity information group are determined.

6. The method according to claim 5, characterized in that The determining, based on the similarity result, an entity information group and target entity information corresponding to the entity information group includes: For the plurality of initial entity information, determining at least one entity information group corresponding to the current initial entity information according to at least one similarity result between the current initial entity information and other initial entity information except the current initial entity information; In the case where the current initial entity information corresponds to multiple entity information groups, entity disambiguation processing is performed on the multiple entity information groups to obtain target entity information corresponding to the multiple entity information groups.

7. The method according to claim 1, characterized in that The determining the second relationship information of the target entity information in the at least two data sources includes: Taking any two target entity information as a target entity pair; For target entity information in any of the target entity pairs, performing text extraction processing on the at least two data sources to determine text data corresponding to the target entity pair; Processing the text data based on a pre-trained word embedding model to determine a word embedding vector of the text data, and determining a first feature based on the word embedding vector and an entity embedding vector of the target entity pair; Based on at least two of the initial tobacco knowledge graphs, determining a relation embedding vector corresponding to the target entity pair, and using the relation embedding vector as a second feature; Concatenate the first feature and the second feature to obtain a target feature of the target entity pair; The target feature of at least one target entity pair is input into the relation classification model to obtain second relation information.

8. The method according to claim 1, characterized in that The determining, based on the first relationship information, the second relationship information, and the third relationship information, target relationship information corresponding to the target entity information includes: Determine a plurality of relationship types to be merged based on the first relationship information, the second relationship information, and the third relationship information; Determining semantic similarity information between the multiple relationship types to be fused, and based on the semantic similarity information, performing clustering processing on the multiple relationship types to be fused to obtain at least one relationship type combination; Performing confidence evaluation processing on each of the relationship type combinations based on a preset attribute evaluation function to determine a confidence evaluation result corresponding to each of the relationship type combinations; The relationship type combination whose confidence evaluation result meets the preset conditions is abstracted into a target relationship type, and the target relationship information is determined based on the target relationship type.

9. The method according to claim 1, characterized in that: After determining the target tobacco knowledge graph, the method further includes: When the incremental production scheduling information is obtained, determining at least one candidate link entity matching the incremental production scheduling information from the target tobacco knowledge graph; For at least one of the candidate link entities, determining an entity similarity result and a relationship similarity result between the incremental production scheduling information and each of the candidate link entities; Determine a link evaluation attribute corresponding to each of the candidate link entities based on a preset link evaluation function, an entity similarity result corresponding to each of the candidate link entities, and a relationship similarity result; Based on at least one of the link evaluation attributes, a target link entity is determined, and based on the target link entity, the incremental production scheduling information is linked to the target tobacco knowledge graph to obtain an updated target tobacco knowledge graph.

10. The method according to claim 1, characterized in that After determining the target tobacco knowledge graph, the method further includes: Inputting at least one target entity information and / or target relationship information in the target tobacco knowledge graph, and a random noise vector, into a pre-trained triple generation network model for processing to obtain a triple to be used corresponding to the at least one target entity information and / or target relationship information; Based on the triples to be used, determining missing entity information and / or missing relationship information in the target tobacco knowledge graph; Based on the missing entity information and / or the missing relationship information, the target tobacco knowledge graph is adjusted to obtain an adjusted target tobacco knowledge graph.

11. A knowledge graph determination device applied to tobacco production scheduling, characterized in that: include: A relationship information determination module is used to extract data from at least two data sources associated with tobacco production scheduling, and obtain tobacco production scheduling information corresponding to each data source and first relationship information between the tobacco production scheduling information; An initial tobacco knowledge graph determination module is used to determine the initial tobacco knowledge graph corresponding to each data source based on the tobacco production scheduling information and the first relationship information corresponding to each data source, wherein the tobacco production scheduling information serves as the initial entity information in the initial tobacco knowledge graph; An entity alignment module, used to perform entity alignment processing on the initial entity information in at least two of the initial tobacco knowledge graphs, determine entity information groups corresponding to at least two of the initial tobacco knowledge graphs, and target entity information corresponding to the entity information groups, wherein the entity information groups include initial entity information with matching relationships; A target relationship information determination module, used to determine the second relationship information of the target entity information in the at least two data sources and the third relationship information in the initial tobacco knowledge graph, and determine the target relationship information corresponding to the target entity information based on the first relationship information, the second relationship information and the third relationship information; The target tobacco knowledge graph determination module is used to fuse at least two initial tobacco knowledge graphs based on the target entity information and the target relationship information to determine the target tobacco knowledge graph.