Cross-border industry knowledge graph construction method and system based on prompt project
By constructing a cross-border industry knowledge graph based on prompting engineering, the problems of insufficient domain knowledge adaptability and language differences are solved, and cross-language entity alignment and high-quality knowledge graph construction are achieved, supporting cross-border industrial chain analysis.
Patent Information
- Application Number
- CN202510811787.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-06-18
AI Technical Summary
Existing methods for constructing cross-border domain industry knowledge graphs face challenges such as insufficient domain knowledge adaptability, inconsistencies between output results and facts, and language differences, resulting in insufficient accuracy of models in cross-language entity alignment.
By adopting a prompting engineering approach, a progressive architecture is used to construct a cross-border industry knowledge graph, including entity relationship graph construction, entity and relationship type definition, entity and relationship extraction, self-consistent triple purification, and multilingual embedding coding model.
It effectively constructed a cross-border industry knowledge graph, improved the completeness and accuracy of knowledge association, and provided a reliable data infrastructure to support cross-border industrial chain analysis.
Smart Images

Figure CN120893533A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application provides a cross-border industry knowledge graph construction method and system based on prompt engineering, and belongs to the technical field of knowledge graph construction. BACKGROUND
[0002] As a knowledge representation method based on graph structure, knowledge graph constructs a semantic network of triples of "entity-relation-attribute", constructs a cross-domain association network, and provides a multi-dimensional cognitive framework for industrial economic research. This structured knowledge system not only supports basic functions such as semantic retrieval and intelligent question answering, but also can deeply analyze the spatial layout characteristics of the industrial chain, the cross-country investment flow rules and the market supply and demand matching mechanism through graph computing technologies such as path reasoning and community discovery, providing data-driven decision support for government departments to formulate industrial policies and enterprises to plan cross-border investment.
[0003] A domain knowledge graph is a structured representation of knowledge in a specific domain. It extracts, integrates, and organizes entities and their semantic relationships within a domain to provide strong support for information retrieval, semantic reasoning, and decision-making in that domain. Compared to general knowledge graphs (Wiki Data, Google Knowledge Graph), domain knowledge graphs focus on the in-depth integration of professional knowledge in vertical industries or disciplines. In the medical field, researchers capture disease and symptom-related entities and relationships from electronic medical records and combine them with the Google Health Knowledge Graph (GHKG) to create a knowledge graph containing diseases, symptoms, and their relationships. In the field of network security, Jia et al. developed a domain ontology and proposed a technique for building a network security knowledge graph, ultimately constructing a network security graph based on a five-tuple model. In the financial field, knowledge graphs are applied to market analysis and stock price prediction. Liu et al. combined the semantic information of knowledge graphs with the feature extraction capabilities of deep learning to build a feature combination model for predicting the stock price changes of well-known companies. The above work shows that domain knowledge graphs are driving the cognitive intelligence upgrade in industries such as healthcare, network security, education, and finance. However, existing methods face more complex cross-language knowledge integration challenges in cross-border industry environments and need to achieve real-time entity disambiguation and semantic mapping in dynamic multilingual data streams. With the breakthrough of large language models (LLMs) in the trillion-parameter era, they have shown significant advantages in cognitive reasoning tasks in complex contexts. In open information extraction tasks, researchers use prompt engineering strategies to guide large models to complete entity recognition and relationship extraction tasks by designing domain-specific instruction templates and in-context learning. This effectively alleviates the problem of a lack of professional field annotation data. This parameterized knowledge transfer learning paradigm significantly reduces the dependence on manual annotation in traditional methods.
[0004] However, existing federated distillation-based frameworks still face many challenges:
[0005] Lack of domain knowledge adaptability. Large models are trained on general pre-training, but there are many specific domain knowledge in professional fields. Large models have problems such as lack of domain knowledge when building knowledge graphs in professional fields (e.g. specific relationships in the industry field, litigation, shareholders, etc.);
[0006] Output results are inconsistent with facts: the biggest problem of large models is the illusion problem, which will over-predict knowledge unrelated to the task. The triples generated by the model often have factual errors such as format errors and logical contradictions (such as the triple <Ningde Times, executive, NIO> which is a relationship between the executive and the enterprise). This phenomenon will cause semantic differences with the ontology definition, which will directly affect the recommendation system based on knowledge graph and other related downstream tasks.
[0007] Language differences: different language systems have significant differences in vocabulary representation (such as the polysemy of terms, the cultural specificity of named entities), syntactic structure (such as the logical bias of the predicate caused by the difference in word order), and pragmatic rules (such as the context dependence of industry terms). The deep semantic gap between languages causes problems such as insufficient accuracy in cross-language entity alignment. SUMMARY
[0008] To solve the above problems, the present application proposes a cross-border industry knowledge graph construction method and system based on prompt engineering. The method effectively constructs a cross-border industry knowledge graph through a progressive architecture of knowledge extraction, knowledge purification, and knowledge fusion, and the constructed cross-border industry knowledge graph structure is more reasonable and the knowledge association is more complete.
[0009] The technical scheme of the present application is: a cross-border industry knowledge graph construction method based on prompt engineering, the method comprising:
[0010] Step 1, through the ontology definition module, first construct an entity relationship graph as a reference framework, and then define the entity and entity relationship type;
[0011] Step 2, through the knowledge extraction module, perform entity extraction and relationship extraction; specifically, extract entities and triples from text according to the defined entity and entity relationship type through the prompt large model;
[0012] Step 3, through the knowledge purification module, use the self-consistency of the large model to purify the preliminarily extracted triples to obtain triples with higher quality;
[0013] Step 4, through the knowledge fusion module, use a multilingual embedding coding model to vectorize the entities, and then combine the large model to realize knowledge fusion; calculate the cosine similarity between each entity, select the top several entities with the highest similarity as the candidate entity pool, input them into the large model and construct prompts to realize entity alignment, and finally obtain the knowledge graph through the entities and triples after the knowledge purification module and entity alignment.
[0014] Further, the Step 1 comprises:
[0015] The definition of the entity comprises:
[0016] In the ontology construction of the industrial knowledge graph, the entities include subjects, decision makers, spatial elements, and business carriers covering economic activities;
[0017] The defined entity relationship types include:
[0018] The relationship design is based on the industry chain logic and the actual business scenario: the relationship between enterprises reflects the market structure and competition dynamics; the relationship between enterprises and persons reveals the governance structure and the allocation of rights and responsibilities; the relationship between enterprises and places maps the regional economic layout; and the relationship between enterprises and projects describes the resource allocation and business development.
[0019] Further, the Step2 includes:
[0020] Step2.1, for the entity extraction task, extracting entities e_j from text X, and the entity type E is predefined by analyzing the data set; a prompt is constructed to let the large model perform entity extraction, which includes task instructions, related examples, and entity types and test texts;
[0021] Step2.2, after extracting the entities in the first step, the relationship between the entities is extracted as a triple; according to the given entity relationship type, the entities are extracted from the text in the first step, and the large model is prompted to extract the relationship between the entities according to the entity relationship type;
[0022] Step2.3, introduce dynamic selection examples to construct prompts, select the k examples with the highest similarity to the input text;
[0023] An example library is constructed to select examples, which contains example texts and entities and triples contained in the example texts;
[0024] First, in the preprocessing stage, the effectiveness of the example library is filtered, and the invalid data with missing core fields is removed, and the text features are extracted based on the TF-IDF vector space model: the term frequency-inverse document frequency TF-IDF weighting strategy is used to strengthen the representation of domain keywords, N-gram features are integrated to capture the semantics of Thai compound words, and a stop word filtering mechanism is integrated to eliminate redundant interference, and finally a K-nearest neighbor index structure supporting cosine similarity retrieval is constructed;
[0025] In the real-time query stage, after the input text is mapped by the homologous vector, K-nearest neighbor search is performed through the pre-constructed index.
[0026] Further, the Step3 includes:
[0027] Step3.1, relationship verification, through the function Validate(e i ,r,e j) returns different results according to whether the relation r satisfies a certain condition:
[0028] If the relation r belongs to a predefined set R, i.e. r∈R, the function returns the complete triple (e i ,r,e j );
[0029] If the purification condition φ(r|R) of the relation r is non-empty, i.e. the function returns an empty set indicating an invalid relation;
[0030]
[0031] Step 3.2, triple verification, count the frequency of triple t in N times of sampling;
[0032] If the frequency exceeds the threshold τ, then F(t) = 1, and it is determined as a valid triple;
[0033] Otherwise, F(t) = 0, and it is determined as a noise triple;
[0034] The statistical consistency of multiple samplings improves the robustness of the results;
[0035]
[0036] where t is the triple (e i ,r,e j ) to be verified, T k is a set of triples generated by independent sampling, N is the number of independent sampling times, and I represents the indicator function;
[0037] Return 1 when the internal condition is met, otherwise return 0; t∈T k in the inner formula is used to determine whether the triple t appears in the k-th sampling, and τ is a set consistency threshold.
[0038] Further, the Step 4 comprises:
[0039] Step 4.1, construct a cross-language semantic space through a multi-language embedding coding model, and realize the deep fusion of the triple-language knowledge graph through a hierarchical alignment strategy; the core is to construct a hierarchical alignment mechanism combining cross-language semantic space and multi-modal verification; based on the deep semantic coding ability of the multi-language embedding coding model, a number of language entities are mapped to a unified 1024-dimensional semantic space; through vector normalization processing, the length difference of different language entity vectors is eliminated, so that the cross-language similarity calculation has spatial consistency;
[0040]
[0041] wherein, v is a standardized vector, i ||v is the original input, i ||2 represents its L2 norm, here, the introduction of epsilon avoids the problem of zero vector, and the max function ensures that the denominator is not zero;
[0042] Step4.2, for the semantic category of the target entity, the type compatible candidate set is screened in the target language, and the cosine similarity is calculated, and the Top-k candidate entities with similar semantics are retained to construct a preliminary alignment pool;
[0043] Step4.3, the multilingual name, entity type and context description of the candidate entity are integrated into a structured prompt template to drive the large language model to perform semantic equivalence discrimination; secondly, by comparing the semantic coverage of the translated text, the morphological differences specific to the language are analyzed;
[0044] Step4.4, verification and disambiguation step, finally execute attribute constraint verification, ensure the logical consistency of the core attribute; set dynamic confidence threshold to process the verification result in stages, when the model output confidence reaches the preset standard, the alignment relationship is automatically generated; by constructing a prompt to stimulate the potential of the large model, the repeated and incorrect entities are disambiguated from two aspects of translation verification and attribute constraint verification, and finally the aligned entities are obtained.
[0045] The application also provides a cross-border industry knowledge graph construction system based on prompt engineering, the system comprises: a method for executing the cross-border industry knowledge graph construction method based on prompt engineering.
[0046] The application also provides an electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to implement the cross-border industry knowledge graph construction method based on prompt engineering.
[0047] The application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the cross-border industry knowledge graph construction method based on prompt engineering.
[0048] The application also provides a computer program product comprising a computer program, wherein the computer program is executed by a processor to implement the cross-border industry knowledge graph construction method based on prompt engineering.
[0049] The application has the following advantages:
[0050] 1、The application collects industry field texts in Chinese, Vietnamese and Thai, and constructs industry field ontology according to industry characteristics and application requirements.
[0051] 2. This invention proposes a method for constructing a China-Vietnam-Thailand knowledge graph based on prompting engineering, which is divided into knowledge extraction, knowledge purification, and knowledge fusion, and effectively constructs a China-Vietnam-Thailand cross-border industrial knowledge graph.
[0052] 3. This invention conducted a series of knowledge graph representation learning experiments on the China-Vietnam-Thailand cross-border industrial knowledge graph. The experimental results proved the structural rationality and knowledge association completeness of the knowledge graph. Attached Figure Description
[0053] Figure 1 This is an overall framework diagram of the method for constructing a cross-border industry knowledge graph based on prompting engineering, provided in an embodiment of the present invention. Detailed Implementation
[0054] The embodiments of the present invention will now be described with reference to the accompanying drawings. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0055] like Figure 1 The diagram shows the overall framework of the cross-border industry knowledge graph construction method based on prompting engineering provided in this invention. This invention first collects multilingual industry texts covering Chinese, Vietnamese, and Thai, and designs a domain ontology containing several entities and various types of relationships, focusing on representing relationships such as investment and financing, and executive appointments between enterprises and between enterprises and individuals. To address the insufficient adaptability of large-scale models to domain knowledge, during the knowledge extraction stage, prompting engineering guides large language models to complete structured information extraction, standardizing the output format through predefined entity and relationship type constraint prompt templates. Innovatively, a dynamic example selection mechanism is introduced in prompt construction, adaptively selecting Top-k relevant examples from the industry example library based on text similarity to enhance the model's reasoning ability. To address the semantic illusion problem in the model output, a self-consistent filtering strategy is adopted during the knowledge purification stage, generating candidate results through multiple rounds of sampling and setting voting thresholds to effectively eliminate low-confidence triples. To address the core challenges of multilingual knowledge fusion, this invention constructs an alignment mechanism based on dual verification: entities are mapped to a unified semantic space using a multilingual embedding model, and a Top-k candidate entity pool is built through normalization and similarity calculation; then, a large-model-driven bilingual prompting strategy is employed to achieve entity alignment from two dimensions: cross-language translation verification and semantic consistency verification, ultimately completing the fusion and disambiguation of cross-language triples. The China-Vietnam-Thailand cross-border industrial knowledge graph constructed in this invention provides reliable data infrastructure support for cross-border industrial chain analysis. To verify the topological rationality of the knowledge graph, this invention further implements link prediction experiments, providing quantitative evidence for graph quality assessment.
[0056] The cross-border industry knowledge graph construction method based on the prompt engineering comprises the following steps:
[0057] Step 1, through the ontology definition module, first taking the business entity relationship graph constructed by the Qichacha platform as a benchmark framework, and then defining the entity and the entity relationship type;
[0058] Further, the Step 1 comprises:
[0059] Defining the entity comprises:
[0060] In the ontology construction of the industry knowledge graph, the entity includes the subject (enterprise), decision maker (person), spatial element (place) and business carrier (project) covering economic activities;
[0061] The defined entity relationship type comprises:
[0062] The relationship design is based on the industry chain logic and the actual business scene: the relationship between enterprises (such as investment, litigation) reflects the market structure and competition dynamics; the relationship between enterprises and persons (such as shareholders, senior managers) reveals the governance structure and the distribution of rights and responsibilities; the relationship between enterprises and places (such as registered address) maps the regional economic layout; and the relationship between enterprises and projects (such as investment, ownership) depicts the resource allocation and business development.
[0063] Step 2, through the knowledge extraction module, entity extraction and relationship extraction are performed; specifically, the defined entity and the entity relationship type are extracted from the text by the large model according to the prompt;
[0064] Further, the Step 2 comprises:
[0065] Step 2.1, for the entity extraction task, the entity e_j is extracted from the text X, the entity is the most basic element in the knowledge graph, also known as instance or individual, each e_j is a specific thing in the real world, such as enterprise, person, etc.; the entity type E is defined in advance by analyzing the data set; the large model is prompted to extract the entity by constructing the prompt, which includes task instructions, related examples, entity type and test text;
[0066] Step 2.2, after the entity is extracted by the first step, the relationship between the entities is extracted as a triple; according to the given entity relationship type, the entity is extracted from the text by the first step, and the large model extracts the relationship between the entities according to the entity relationship type;
[0067] Step 2.3, dynamic selection examples are introduced to construct the prompt, and the k examples with the highest similarity to the input text are selected;
[0068] Build a sample library to select samples, which contains sample text as well as entities and triples contained in the sample text;
[0069] First, in the preprocessing stage, the example library is filtered for validity to remove invalid data with missing core fields. Then, text features are extracted based on the TF-IDF vector space model: the domain keyword representation is strengthened by the term frequency-inverse document frequency TF-IDF weighting strategy, N-gram features are integrated to capture the semantics of Thai compound words, and a stop word filtering mechanism is integrated to eliminate redundant interference. Finally, a K-nearest neighbor (KNN) index structure that supports cosine similarity retrieval is constructed.
[0070] During the real-time query phase, the input text is mapped using homologous vectorization, and a K-nearest neighbor search is performed using a pre-built index. Based on cosine similarity, the system dynamically matches the top-k most relevant examples. Its core advantage lies in achieving context-sensitive retrieval through feature space alignment, enhancing domain relevance discrimination through TF-IDF's rare word reinforcement mechanism, and being compatible with multilingual mixed text processing. This provides accurate contextual reference examples for language models, effectively solving the problem of the disconnect between static examples and input semantics.
[0071] Step 3: Through the knowledge purification module, the self-consistency of the large model is used to purify the initially extracted triples to obtain higher quality triples.
[0072] Furthermore, Step 3 includes:
[0073] Step 3.1: Relationship validation, using the function Validate(e i ,r,e j Different results are returned based on whether relation r satisfies certain conditions:
[0074] If the relation r belongs to a predefined set R, i.e. r∈R, the function returns a complete triple (e i ,r,e j );
[0075] If the purification condition φ(r|R) of relation r is non-empty, i.e. The function returns an empty set. Indicates an invalid relationship;
[0076]
[0077] Step 3.2, triplet verification: count the frequency of triplet t in N samplings;
[0078] If the frequency exceeds the threshold τ, then F(t) = 1, and it is determined to be a valid triplet;
[0079] Otherwise, F(t) = 0, and it is determined to be a noise triple;
[0080] Robustness of the result is improved by statistical consistency of multiple sampling;
[0081]
[0082] Wherein, t is the triplets (e i ,r,e j ) to be verified, T k is a set of triplets generated by independent sampling, N is the number of independent sampling, I represents the indicator function;
[0083] Return 1 when the internal condition is met, otherwise return 0; t in the inner formula T k is used to judge whether the triplet t appears in the k-th sampling, τ is a set consistency threshold.
[0084] Step4, through the knowledge fusion module, the entity is vectorized by using the multilingual embedding coding model, and then the knowledge fusion is realized by combining the large model; The cosine similarity between each entity is calculated, and the top several highest similarity is selected as the candidate entity pool, which is input to the large model and the prompt is constructed to realize entity alignment, and finally the knowledge graph is obtained after the entity and triplets of the knowledge purification module and entity alignment.
[0085] Further, the Step4 includes:
[0086] Step4.1, build a cross-language semantic space through a multilingual embedding coding model, and realize the deep fusion of the three-language knowledge graph through a hierarchical alignment strategy; The core is to build a hierarchical alignment mechanism combined with cross-language semantic space and multi-modal verification; Based on the deep semantic coding ability of the multilingual embedding coding model, the entities of several languages are mapped to a unified 1024-dimensional semantic space; Through vector normalization processing, the length difference of different language entity vectors is eliminated, so that the cross-language similarity calculation has spatial consistency;
[0087]
[0088] Wherein, is the standardized vector, v i is the original input, and ||v i ||2 represents its L2 norm, and here the introduction of ∈ avoids the zero vector problem, and the max function ensures that the denominator is not zero;
[0089] Step4.2, for the semantic category of the target entity, filter the type compatible candidate set in the target language and calculate the cosine similarity, and retain the Top-k candidate entities with similar semantics to construct a preliminary alignment pool;
[0090] Step4.3, integrate the multilingual name of the candidate entity, entity type and context description into a structured prompt template, drive the large language model to perform semantic equivalence discrimination; secondly, by comparing the semantic coverage of the translated text, analyze the morphological differences specific to the language;
[0091] Step4.4, verification and disambiguation step, finally perform attribute constraint verification to ensure the logical consistency of the core attributes; set a dynamic confidence threshold to process the verification results in stages, and automatically generate aligned relationships when the model output confidence reaches the preset standard; by constructing prompts to stimulate the potential of large models, disambiguate repeated and incorrect entities from two aspects of translation verification and attribute constraint verification, and finally obtain the aligned entities.
[0092] The application also provides a cross-border industry knowledge graph construction system based on prompt engineering, which comprises the cross-border industry knowledge graph construction method based on prompt engineering.
[0093] The application also provides an electronic device comprising a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to realize the cross-border industry knowledge graph construction method based on prompt engineering.
[0094] The application also provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to realize the cross-border industry knowledge graph construction method based on prompt engineering.
[0095] The application also provides a computer program product comprising a computer program, wherein the computer program is executed by a processor to realize the cross-border industry knowledge graph construction method based on prompt engineering.
[0096] The application constructs the first multilingual cross-border industry dataset, with a data size of about 50,000 news texts, 15,000 Chinese texts and 15,000 Vietnamese texts, and about 20,000 Thai texts. The data is mainly collected from authoritative news portals of various countries, including Qichacha, News Network and other media platforms in China, VnExpress in Vietnam and Thai news website thairth in Thailand, covering manufacturing, agriculture, e-commerce, education and other industry fields. The data covers market fluctuation trends, enterprise equity changes, cross-border supply chain relationships, senior management job changes and other industry elements. The cross-border industry knowledge graph constructed in Chinese, Vietnamese and Thai provides bottom support, and focuses on serving cross-border investment decision-making, industry chain risk early warning and policy coordination analysis fields.
[0097] The application carries out a link prediction experiment on a China-Vietnam-Thailand cross-border industrial knowledge graph. The data format of the knowledge graph constructed by the application is presented in the form of triples (head entity, tail entity, relationship). The extracted triples are divided into a training set, a test set and a validation set. The triple data is split according to a ratio of (8:1:1), 80% for the training set, and 10% for the test set and the validation set. Then the application preprocesses the split data set, and according to the input format of the model, the application maps each entity and relationship to a unique number, and the number corresponds to the entity relationship.
[0098] In knowledge graph representation learning, the application takes MRR and Hits@10 as evaluation indicators. MRR (Mean Reciprocal Rank) calculates the reciprocal of the correct entity (head entity or tail entity) in the model prediction ranking for each test triple, and then takes the average of all samples. The reciprocal of the correct entity is 1 / k, and the ranking is k. Hits@n (hit rate@n, such as Hits@10) indicates the proportion of samples in which the correct entity appears in the top n of the prediction results.
[0099] In order to verify the effectiveness of the application, the application displays the knowledge graph entities and triples constructed by the application, as follows:
[0100] Table 1 shows the number of entities and triples in Chinese, Vietnamese, Thai, and China-Vietnam and China-Vietnam-Thailand. It is divided into two stages before and after purification, including the number of each type of entity and triple. The number of triples after purification is much less than that before purification, because although the application has set the output format when constructing the prompt, there are still a lot of noise data in the output of the model, such as the extraction of relationships outside the ontology definition in the reasoning process of the model and the output format error of the triples. The number of entities and relationships in Chinese is 50% higher than that before purification, and the number of entities and relationships in Vietnamese is 60% higher than that before purification. Due to the insufficient adaptability of deepseek to Vietnamese, the extraction performance is not good, and the purification rate is only 20%. From the data in the table, it can be seen that after knowledge purification, high-quality triples can be obtained.
[0101] Table 1 is the number of entities and triples of the knowledge graph
[0102]
[0103] In the link prediction task, the application performs comparative experiments on the knowledge graph before and after knowledge purification. The detailed experimental results are described in Table 2. The application evaluates the knowledge graph constructed by the application on the TransE and ComplEx models, and evaluates different languages respectively, including Chinese, Vietnamese, Thai, Sino-Vietnamese, Sino-Vietnamese-Thai. As can be seen from the table, after purification, the indicators on Chinese data on TransE and other models are improved by about 20%, and the performance on Vietnamese data is improved by 1%-3%. However, in the deepseek-v3 large model, compared with Chinese and Vietnamese, the adaptability to Thai is insufficient, resulting in relatively low quality of extracted entities and triple relationships, so there is a problem of poor effect on Thai data on TransE and other models, resulting in that Sino-Vietnamese-Thai only reaches 10%-17% in indicators.
[0104] Table 2 is the experimental result of knowledge graph representation learning
[0105]
[0106] In the description of the present specification, the description referring to the terms "one embodiment", "an example", "a specific example" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0107] The preferred embodiments of the application disclosed above are only used to help explain the application. The preferred embodiments do not describe all the details and limit the application to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the present specification. The present specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the application, so that those skilled in the art can well understand and utilize the application. The application is limited only by the claims and their full scope and equivalents.
Claims
1. A method for constructing a cross-border industry knowledge graph based on prompting engineering, characterized in that, include: Step 1: Using the ontology definition module, first construct the entity relationship graph as the baseline framework, and then define the entity and entity relationship types; Step 2: Entity extraction and relation extraction are performed through the knowledge extraction module; specifically, it is used to extract entities and triples from text based on predefined entities and entity relation types using a prompting model. Step 3: Through the knowledge purification module, the self-consistency of the large model is used to purify the initially extracted triples to obtain higher quality triples. Step 4: Through the knowledge fusion module, entities are vectorized using a multilingual embedding coding model, and then knowledge fusion is achieved by combining with the large model. The cosine similarity between each entity is calculated, and the top few with the highest similarity are selected as the candidate entity pool. This pool is then input into the large model and construction prompts to achieve entity alignment. Finally, the knowledge graph is obtained through the knowledge purification module and the entities and triples after entity alignment.
2. The method for constructing a cross-border industry knowledge graph based on prompting engineering according to claim 1, characterized in that: Step 1 includes: Defining entities includes: In the ontology construction of the industry knowledge graph, entities include subjects of economic activities, decision-makers, spatial elements, and business carriers. The defined entity relationship types include: Relationship design is based on the logic of the industrial chain and actual business scenarios: inter-enterprise relationships reflect market structure and competitive dynamics; enterprise-person relationships reveal governance structure and allocation of power and responsibility; enterprise-location relationships map regional economic layout; and enterprise-project relationships depict resource allocation and business development.
3. The method for constructing a cross-border industry knowledge graph based on prompting engineering according to claim 1, characterized in that: Step 2 includes: Step 2.1: For the entity extraction task, extract entity e_j from text X. The entity type E is predefined by analyzing the dataset. Construct prompts to enable the large model to perform entity extraction. The prompts include task instructions, relevant examples, entity types, and test text. Step 2.2: After extracting entities in the first step, extract the relationships between entities as triples; based on the given entity relationship type, after extracting entities from the text in the first step, the large model extracts the relationships between entities based on the entity relationship type. Step 2.3: Introduce dynamic example selection to build suggestions by selecting the k examples that are most similar to the input text; Build a sample library to select samples, which contains sample text as well as entities and triples contained in the sample text; First, in the preprocessing stage, the example library is filtered for validity to remove invalid data with missing core fields. Then, text features are extracted based on the TF-IDF vector space model: the term frequency-inverse document frequency TF-IDF weighting strategy is used to strengthen the representation of domain keywords, N-gram features are integrated to capture the semantics of Thai compound words, and a stop word filtering mechanism is integrated to eliminate redundant interference. Finally, a K-nearest neighbor index structure that supports cosine similarity retrieval is constructed. During the real-time query phase, the input text is mapped to a common-origin vector and then a K-nearest neighbor search is performed using a pre-built index.
4. The method for constructing a cross-border industry knowledge graph based on prompting engineering according to claim 1, characterized in that: Step 3 includes: Step 3.1: Relationship validation, using the function Validate(e i ,r,e j Different results are returned based on whether relation r satisfies certain conditions: If the relation r belongs to a predefined set R, i.e. r∈R, the function returns a complete triple (e i ,r,e j ); If the purification condition φ(r|R) of relation r is non-empty, i.e. The function returns an empty set. Indicates an invalid relationship; Step 3.2, triplet verification: count the frequency of triplet t in N samplings; If the frequency exceeds the threshold τ, then F(t) = 1, and it is determined to be a valid triplet; Otherwise, F(t) = 0, and it is determined to be a noise triple; The robustness of the results is improved by ensuring statistical consistency through multiple samplings; Where t is the triple to be verified (e i ,r,e j ), T k It is a set of triples generated by independent sampling, where N is the number of independent samples and I represents the indicator function; Returns 1 if the internal condition is met, otherwise returns 0; t∈T in the inner layer of the formula k It is used to determine whether the triplet t appears in the k-th sampling, and τ is the set consistency threshold.
5. The method for constructing a cross-border industry knowledge graph based on prompting engineering according to claim 1, characterized in that: Step 4 includes: Step 4.1: Construct a cross-lingual semantic space through a multilingual embedding coding model, and achieve deep integration of trilingual knowledge graphs through a hierarchical alignment strategy; the core lies in constructing a hierarchical alignment mechanism that combines cross-lingual semantic space with multimodal verification; based on the deep semantic encoding capability of the multilingual embedding coding model, entities of several languages are mapped to a unified 1024-dimensional semantic space; vector normalization is used to eliminate the difference in the magnitude of entity vectors of different languages, so that cross-lingual similarity calculation has spatial consistency; in, For a standardized vector, v i For the original input, ||v i ||2 represents its L2 norm. Here, ∈ is introduced to avoid the zero vector problem. The max function ensures that the denominator is not zero. Step 4.2: Based on the semantic category of the target entity, filter the type-compatible candidate set in the target language and calculate the cosine similarity. Retain the top-k candidate entities with similar semantics to construct an initial alignment pool. Step 4.3: Integrate the multilingual names, entity types, and contextual descriptions of candidate entities into a structured prompt template to drive the large language model to perform semantic equivalence discrimination; secondly, analyze the language-specific morphological differences by comparing the semantic coverage of the translated text. Step 4.4: Validation and disambiguation steps. Finally, attribute constraint validation is performed to ensure the logical consistency of core attributes. A dynamic confidence threshold is set to classify the validation results. When the model output confidence reaches the preset standard, alignment relationships are automatically generated. The potential of the large model is stimulated by constructing hints. Duplication and errors of entities are disambiguated from two aspects: translation validation and attribute constraint validation. The final result is the aligned entity.
6. A cross-border industry knowledge graph construction system based on prompting engineering, characterized in that, The system includes: a method for constructing a cross-border industry knowledge graph based on prompting engineering as described in any one of claims 1 to 5.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for constructing a cross-border industry knowledge graph based on prompting engineering as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for constructing a cross-border industry knowledge graph based on prompting engineering as described in any one of claims 1 to 5.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for constructing a cross-border industry knowledge graph based on prompting engineering as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Cross-language multi-source vertical domain knowledge graph construction method
CN112199511A
Method and device for searching Chinese and cross-border ethnic text fused with domain knowledge graph
CN115599888A
Method for carrying out relation extraction on specific industry based on multi-modal large language model
CN120011577A
Knowledge graph generating method, apparatus, and terminal, and storage medium
WO2021098491A1
Tax field-oriented knowledge map construction method and system
WO2021196520A1
Cited By
Large-model-driven unmanned aerial vehicle equipment training knowledge graph automatic construction method
CN121581162A
Multilingual knowledge graph-based school history culture intelligent guide system and method
CN121636686A
Problem answering method and device based on context map, equipment and medium
CN121998058A
Method, system and equipment for extracting chapter-level knowledge for long text and medium
CN122174957A