English knowledge graph construction method and system

By constructing a physical thesaurus of multi-source heterogeneous English materials, semantic role annotation and classification, and calculating the co-occurrence probability and inter-word distance between entity words, the problem of inaccurate semantic correlation analysis between English entity words in the prior art is solved, and the in-depth correlation analysis and construction of the English knowledge graph is realized.

CN120045721APending Publication Date: 2025-05-27赵海英
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411996000.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to capture the complex semantic connotation behind vocabulary in the analysis of implicit semantic correlation between English entity words, and it is difficult to deal with semantic differences caused by context changes, resulting in inaccurate extraction of implicit semantic relationships.

Method used

By constructing an entity thesaurus of multi-source heterogeneous English materials, semantic role annotation and classification, semantic role annotation loss of each entity word cluster, calculate the co-occurrence probability and inter-word distance between each pair of entity words to determine the semantic dependence, and then construct an explicit and implicit entity-relationship diagram.

Benefits of technology

The in-depth correlation analysis of semantics between English entity words is realized, the reliability and accuracy of the construction of English knowledge graphs is improved, and the explicit and potential semantic connections between entity words can be better revealed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045721A_ABST
    Figure CN120045721A_ABST
Patent Text Reader

Abstract

The invention provides an English knowledge graph construction method and system, and the method comprises the steps: constructing an English entity lexicon for teaching based on multi-source heterogeneous English data, and carrying out the semantic role marking of each entity word in the English entity lexicon, determining the semantic dependency between every two entity words according to the co-occurrence probability and the inter-word distance between every two entity words in the English entity word bank, and further determining the semantic explicit dependency relationship between every two entity words; semantic association analysis of implicit structures is carried out on the entity words among the entity word clusters according to the context information of the English data to obtain semantic association degrees among the entity word clusters, and then the semantic implicit dependency relationships among the entity words are determined; and constructing a knowledge graph for English teaching through an explicit dependency relationship and an implicit dependency relationship of semantics among the entity words. By adopting the scheme of the invention, deep correlation analysis of semantics between English entity words can be realized, so that the reliability of English knowledge graph construction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of semantic processing technology. More specifically, this application relates to a method and system for constructing an English knowledge graph. Background Art

[0002] Semantic processing is to understand and analyze the vocabulary, sentences and their contexts in language, aiming to extract the underlying deep meaning. In the construction of an English knowledge graph, semantic processing helps the system identify entities, relationships and attributes from a large amount of text, and transform this information into structured knowledge. In English teaching, the knowledge graph can provide students with more systematic and structured language knowledge, and support personalized learning paths. Through the semantic information in the knowledge graph, students can understand vocabulary and grammar rules more deeply, master their contexts and usages. At the same time, the knowledge graph can also provide context-based learning resources to help students better apply the knowledge they have learned, thereby improving language understanding and practical application abilities.

[0003] In the analysis of implicit semantic associations between entity words in the prior art, it usually relies on pattern matching methods based on surface text. This method mainly establishes associations between words through keyword co-occurrence or fixed patterns (such as syntactic structures, word frequency statistics). However, this method lacks an in-depth understanding of the actual semantics of English vocabulary, and only relies on the surface form and local context of words for matching, making it difficult to capture the complex semantic connotations behind words. The meaning of words usually depends on specific contexts and context changes, while the surface matching method cannot effectively identify semantic differences caused by context changes. In addition, the extraction of implicit semantic relationships requires comprehensive consideration of broader context information, including the logical structure, emotional tendency and discourse flow of the discourse, etc. These factors are crucial for revealing the deep connections between words, and the prior art is difficult to handle implicit relationships lacking explicit markers in semantics. For example, some entity words seemingly have no direct connection on the surface, but have profound semantic associations in a specific context. This potential connection often cannot be detected by simple surface pattern detection. Therefore, how to achieve in-depth semantic association analysis between English entity words to improve the reliability of English knowledge graph construction has become a difficult problem faced by the industry. Summary of the Invention

[0004] This application provides a method and system for constructing an English knowledge graph, which can achieve in-depth semantic association analysis between English entity words, thereby improving the reliability of English knowledge graph construction.

[0005] In a first aspect, this application provides a method for constructing an English knowledge graph, including the following steps: Construct an English entity word library for teaching based on multi-source heterogeneous English materials; Perform semantic role labeling on each entity word in the English entity word library, classify the corresponding entity words according to the labeled semantic role information, obtain entity word clusters with different semantic roles, and then determine the labeling loss of the semantic role of each entity word cluster. Determine the semantic dependence between every two entity words in the English entity word library according to the co-occurrence probability and word distance between them, and then determine the explicit semantic dependence relationship between every two entity words from all the semantic dependencies and the semantic role information of each entity word. Conduct an implicit structure semantic association analysis on the entity words between each entity word cluster according to the context information of the English materials, obtain the semantic association degree between each entity word cluster, and then perform a dependence analysis on the implicit semantic relationship between every two entity words in the English entity word library based on all the semantic association degrees and each labeling loss, to obtain the implicit semantic dependence relationship between every two entity words. Construct a knowledge graph for English teaching through the explicit semantic dependence relationship and implicit semantic dependence relationship between every two entity words.

[0006] Preferably, performing semantic role labeling on each entity word in the English entity word library and classifying the corresponding entity words according to the labeled semantic role information to obtain entity word clusters with different semantic roles specifically includes: Build a labeling model for semantic role labeling based on deep learning. Take each entity word in the English entity word library as the input parameter of the labeling model. Label the semantic role information of each entity word through the labeling model. Divide all entity words into entity word clusters with different semantic roles according to the semantic role information of each entity word.

[0007] Preferably, determining the labeling loss of the semantic role of each entity word cluster specifically includes: For each entity word cluster, obtain all the entity words in the entity word cluster, and then evaluate the error of the semantic role information of each entity word to obtain the role labeling error of each entity word. Determine the labeling loss of the semantic role of the entity word cluster through the role labeling errors of all entity words, and then obtain the labeling loss of the semantic role of each entity word cluster.

[0008] Preferably, determining the semantic dependence between every two entity words in the English entity word library according to the co-occurrence probability and word distance between them specifically includes: For every two entity words in the English entity word library; Determine the co-occurrence probability and word distance between the two entity words. Determine the co-occurrence coefficient between two entity words based on the co-occurrence probability and the distance between words; Use the co-occurrence coefficient as a constraint term of the objective function in the preset semantic relation model; Extract the semantic dependence degree between two entity words through the semantic relation model, and then obtain the semantic dependence degree between every two entity words.

[0009] Preferably, determining the explicit semantic dependence relationship between every two entity words based on all semantic dependence degrees and the semantic role information of each entity word specifically includes: For every two entity words, determine the matching relationship of semantic roles between the two entity words according to the semantic role information of the two entity words; Determine the explicit semantic dependence relationship between two entity words through the semantic dependence degree and the matching relationship of semantic roles between the two entity words, and then obtain the explicit semantic dependence relationship between every two entity words.

[0010] Preferably, perform implicit structure semantic association analysis on the entity words between each entity word cluster according to the context information of the English materials, and obtain the semantic association degree between each entity word cluster specifically includes: For every two entity word clusters, extract the semantic relationship of the entity words between the two entity word clusters in the syntactic structure according to the context information of the English materials; Determine the semantic relationship of the entity words between the two entity word clusters in the word order structure according to the context information of the English materials; Determine the semantic association degree between two entity word clusters through the semantic relationship in the syntactic structure and the semantic relationship in the word order structure, and then obtain the semantic association degree between each entity word cluster.

[0011] Preferably, construct a knowledge graph for English teaching through the explicit and implicit dependence relationships of semantics between every two entity words specifically includes: Construct an entity-relationship graph of explicit semantics through the explicit dependence relationship of semantics between every two entity words; Construct an entity-relationship graph of implicit semantics according to the implicit dependence relationship of semantics between every two entity words; Perform associative fusion on the entity-relationship graph of explicit semantics and the entity-relationship graph of implicit semantics to obtain a knowledge graph for English teaching.

[0012] In a second aspect, the present application provides an English knowledge graph construction system, including: A construction module, configured to construct an English entity word library for teaching based on multi-source heterogeneous English materials; A processing module, configured to perform semantic role labeling on each entity word in the English entity word library, classify the corresponding entity words according to the labeled semantic role information to obtain entity word clusters with different semantic roles, and further determine the labeling loss of the semantic role of each entity word cluster; The processing module is further configured to determine the semantic dependence between every two entity words in the English entity word library according to the co-occurrence probability and word distance between every two entity words, and further determine the explicit semantic dependence relationship between every two entity words from all the semantic dependencies and the semantic role information of each entity word; The processing module is further configured to perform semantic association analysis of implicit structures on the entity words between each entity word cluster according to the context information of the English material to obtain the semantic association degree between each entity word cluster, and further perform dependence analysis on the implicit semantic relationship between every two entity words in the English entity word library based on all the semantic association degrees and each labeling loss to obtain the implicit semantic dependence relationship between every two entity words; An execution module, configured to construct a knowledge graph for English teaching through the explicit semantic dependence relationship and implicit semantic dependence relationship between every two entity words.

[0013] In a third aspect, the present application provides a computer device, which includes a memory and a processor. The memory stores code, and the processor is configured to obtain the code and execute the above-mentioned English knowledge graph construction method.

[0014] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned English knowledge graph construction method is implemented.

[0015] The technical solutions provided by the disclosed embodiments of the present application have the following beneficial effects: In an embodiment of the present application, an English entity word bank for teaching is constructed based on multi-source heterogeneous English materials; semantic role annotation is performed on each entity word in the English entity word bank, and the corresponding entity words are classified according to the annotated semantic role information to obtain entity word clusters with different semantic roles, and then the annotation loss of the semantic role of each entity word cluster is determined; the semantic dependence between every two entity words is determined according to the co-occurrence probability and the word distance between every two entity words in the English entity word bank, and then the explicit semantic dependence relationship between every two entity words is determined by all the semantic dependencies and the semantic role information of each entity word; semantic association analysis of the implicit structure is performed on the entity words between each entity word cluster according to the context information of the English materials to obtain the semantic association degree between each entity word cluster, and then the implicit semantic relationship between every two entity words in the English entity word bank is analyzed based on all the semantic association degrees and each annotation loss to obtain the implicit semantic dependence relationship between every two entity words; a knowledge graph for English teaching is constructed through the explicit semantic dependence relationship and the implicit semantic dependence relationship between every two entity words.

[0016] It can be seen that the present application constructs a knowledge graph for English teaching through explicit and implicit semantic dependency relationships between every two entity words. First, semantic role labeling is performed on each entity word, and the corresponding entity words are classified according to the labeled semantic role information to obtain multiple entity word clusters. Through semantic role labeling, the functional attributes of entity words can be clarified, which helps to distinguish the semantic roles of entities in the context. And the classification based on the labeled information can gather entity words with similar semantics or related functions in the same word cluster. This classification method not only makes semantic analysis more organized but also lays a foundation for explicit and implicit semantic association analysis. By focusing on the semantic relationships within and between entity word clusters, the accuracy and efficiency of association analysis are improved. Secondly, the semantic dependency degree is determined by calculating the co-occurrence probability and word distance between every two entity words, and then the explicit dependency relationship between the semantics of every two entity words is mined. This step not only integrates statistical information but also combines the key context factor of word distance, making up for the deficiency of traditional methods in capturing context changes, so as to more accurately reflect the explicit semantic connection between entity words. Then, semantic association analysis is performed on the implicit structure between entity word clusters based on context information to further identify the implicit semantic relationships between words. By integrating context, logical structure, and potential semantic connections, the deficiencies in implicit semantic extraction in the prior art can be made up, making the deep semantic association between entity words clearer. Finally, a knowledge graph for English teaching is constructed through the explicit and implicit dependency relationships between the semantics of entity words, which can not only reveal the explicit semantic logic between entity words but also mine their potential semantic connections, providing structured and hierarchical language knowledge support for English teaching and significantly improving the reliability and teaching application value of the knowledge graph. In summary, the solution of the present application can realize the in-depth association analysis of the semantics between English entity words, thereby improving the reliability of English knowledge graph construction. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is an exemplary flowchart of a method for constructing an English knowledge graph according to some embodiments of the present application; Figure 2 is a schematic structural diagram of constructing an English knowledge graph according to some embodiments of the present application; Figure 3 is a schematic flowchart of determining semantic association degree according to some embodiments of the present application; Figure 4 is a schematic structural diagram of an English knowledge graph construction system according to some embodiments of the present application; Figure 5 is a schematic structural diagram of a computer device for implementing the method for constructing an English knowledge graph according to some embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] In order to better understand the technical solution of the present application, the technical solution of the present application will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0019] refer to Figure 1 , which is an exemplary flow chart of a method for constructing an English knowledge graph according to some embodiments of the present application. The English knowledge graph construction method 100 mainly includes the following steps: In step 101, an English entity vocabulary for teaching is constructed based on multi-source heterogeneous English materials.

[0020] In the specific implementation, first, English materials suitable for English teaching in different grades are collected, among which English materials can specifically include electronic text materials such as English textbooks, extracurricular English articles, teaching materials and dictionaries; then, named entity recognition tools such as SpaCy and Stanford NER are used to extract entity words (for example, place names, names of people, objects and concepts) in the English materials, and noun phrases (NP) and verb phrases (VP) can be further extracted through part-of-speech tagging (POS tagging) and dependency parsing (Dependency Parsing), among which noun phrases and verb phrases are both regarded as entity words. Furthermore, all the extracted entity words are combined into a set as an English entity vocabulary for English teaching.

[0021] It should be noted that the entity words in the English entity vocabulary described in this application cover all nouns with specific and unique identifiers in a specific context, and usually refer to objects with clear reference functions such as things, places, people, time, organizations, etc. in the real world.

[0022] It should also be noted that the reference Figure 2 As shown in the figure, this is a schematic diagram of the structure of constructing the English knowledge graph in this application. The structure includes the following key parts: entity words, which are the basis for constructing the knowledge graph and include all entity words that need to be analyzed; explicit semantic relationships, which construct explicit entity-relationship graphs by analyzing the direct semantic dependency between every two entity words in the English entity vocabulary; implicit semantic relationships, which are opposite to explicit semantic relationships. Implicit semantic relationships are obtained through deeper analysis. These relationships are not directly expressed but implied in the text, such as the potential connection between "apple" and "health"; fusion weights, which are used to weigh between explicit and implicit semantic relationships to determine their relative importance in the final knowledge graph; loss metrics, which are used to measure the accuracy of the model in the process of constructing the English knowledge graph. By minimizing the loss metrics, the performance of the model can be improved.

[0023] In step 102, semantic role labeling is performed on each entity word in the English entity word library, and the corresponding entity words are classified according to the labeled semantic role information to obtain entity word clusters with different semantic roles, and then the labeling loss of the semantic role of each entity word cluster is determined.

[0024] In some embodiments, semantic role labeling is performed on each entity word in the English entity word library, and the corresponding entity words are classified according to the labeled semantic role information to obtain entity word clusters with different semantic roles, which can be implemented by the following steps: Construct a labeling model for semantic role labeling based on deep learning; Take each entity word in the English entity word library as the input parameter of the labeling model; Label the semantic role information of each entity word through the labeling model; Divide all entity words into entity word clusters with different semantic roles according to the semantic role information of each entity word.

[0025] Specifically, first, a labeling model for semantic role labeling (SRL, Semantic Role Labeling) can be constructed based on an existing deep learning architecture (such as the BERT architecture); then, all entity words in the English entity word library are input into the trained labeling model, and the semantic roles of each entity word are further learned and analyzed through the labeling model, and the output role information is used as the semantic role information of each entity word, such as the subject (Agent), object (Patient), and instrument (Instrument). Specifically, the labeling model identifies the role played by each entity word one by one through sequence labeling and assigns the most frequently occurring semantic role to each entity word; finally, it can be implemented in the form of a dictionary, that is, the semantic role information of each entity word is labeled, and all entity words with the same label are grouped into one category to obtain entity word clusters with different semantic roles.

[0026] It should be noted that in the embodiments of the present application, the semantic roles of entity words can be labeled as three types of roles: subject, object, and instrument. In other embodiments, the semantic roles of entity words can also be labeled as others, which are not specifically limited here.

[0027] In some embodiments, the labeling loss of the semantic role of each entity word cluster can be implemented by the following steps: For each entity word cluster, obtain all the entity words in the entity word cluster, and then evaluate the error of the semantic role information of each entity word to obtain the role labeling error of each entity word; The annotation loss of the semantic role of the entity word cluster is determined by the role annotation errors of all entity words, and then the annotation loss of the semantic role of each entity word cluster is obtained.

[0028] In specific implementation, for each entity word cluster, first, all entity words in the entity word cluster are obtained, then an entity word is selected as the selected entity word, and then the difference value between the actual annotation information and the model prediction annotation information of the selected entity word is calculated through the cross-entropy loss function, and this difference value is used as the role annotation error of the selected entity word. Repeating the above steps can obtain the role annotation errors of the remaining entity words; then, the average value of the role annotation errors of all entity words can be used as the annotation loss of the semantic role of the entity word cluster. Repeating the method of the above embodiment can obtain the annotation loss of the semantic role of each entity word cluster.

[0029] It should be noted that the role annotation error in this application refers to the difference between the semantic role information predicted by the model and the true annotation information in the semantic role annotation task; in addition, the annotation loss is an index to measure the accuracy of the semantic role annotation model, which reflects the overall difference between the semantic role annotation results of all entity words in the entity word cluster and the true annotation.

[0030] In step 103, the semantic dependence degree between every two entity words is determined according to the co-occurrence probability and the word distance between every two entity words in the English entity word library, and then the explicit dependence relationship between the semantics of every two entity words is determined by all the semantic dependence degrees and the semantic role information of each entity word.

[0031] In some embodiments, the determination of the semantic dependence degree between every two entity words according to the co-occurrence probability and the word distance between every two entity words in the English entity word library can be implemented by the following steps: For every two entity words in the English entity word library; Determine the co-occurrence probability and the word distance between the two entity words; Determine the co-occurrence coefficient between the two entity words through the co-occurrence probability and the word distance; Take the co-occurrence coefficient as the constraint term of the objective function in the preset semantic relationship model; Extract the semantic dependence degree between the two entity words through the semantic relationship model, and then obtain the semantic dependence degree between every two entity words.

[0032] In specific implementation, for every two entity words in the English entity word library, first, the co-occurrence probability can be determined by calculating the frequency of the two entity words co-occurring within the same context window (such as a sentence, a paragraph, or a text block of a specific length). Specifically, assume that entity words E1 and E2 appear n1 and n2 times respectively in the English entity word library, and the number of times they co-occur within the same context window is c(E1, E2). Then, the co-occurrence probability P(E1, E2) between entity words E1 and E2 can be expressed as: P(E1, E2) = c(E1, E2) / min(n1, n2). Further, the word distance refers to the positional difference between two entity words in the same text, usually quantified by calculating the difference in their relative positions in a sentence. For example, if two entity words appear in the same sentence, the word distance is the difference in their lexical positions in that sentence. Secondly, the product of the reciprocal of the sum of the word distance and the value 1 and the co-occurrence probability can be used as the co-occurrence coefficient between the two entity words. Here, adding 1 is to avoid the problem of dealing with zero distance (i.e., the same position). By the above method, it can be ensured that when the word distance is smaller (i.e., the lexical positions are closer), the co-occurrence coefficient is larger, and vice versa. Then, the co-occurrence coefficient can be introduced into the objective function of the semantic relationship model, and the co-occurrence coefficient is used as a constraint parameter of the semantic relationship model. Through the constraint parameter, the weight of the semantic relationship model in learning the semantic dependence between words can be adjusted, thereby improving the accuracy of the semantic relationship model in extracting the semantic relationship between entity words. Finally, the semantic dependence relationship between the two entity words is derived through the network structure in the semantic relationship model, and this semantic dependence relationship is mapped to a value between 0 and 1, and this value is used as the semantic dependence degree between the two entity words. The magnitude of this dependence degree can reflect the strength of the semantic dependence relationship between the entity words. The larger the value, the stronger the semantic dependence relationship between the entity words. Eventually, through the above method, the semantic dependence degree between every two entity words can be obtained.

[0033] It should be noted that the semantic dependence degree in the embodiments of the present application is an index to measure the degree of mutual dependence between two entity words semantically. The higher the dependence degree, the closer the semantic relationship between the two entity words, the stronger the dependence, and vice versa, the weaker the relationship.

[0034] In some embodiments, the explicit semantic dependence relationship between every two entity words determined by all the semantic dependence degrees and the semantic role information of each entity word can be implemented by the following steps: For every two entity words, determine the matching relationship of semantic roles between the two entity words according to the semantic role information of the two entity words; Determine the explicit semantic dependence relationship between the two entity words through the semantic dependence degree and the matching relationship of semantic roles between the two entity words, and then obtain the explicit semantic dependence relationship between every two entity words.

[0035] In specific implementation, for each pair of entity words E1 and E2, first, a semantic role matching model is preset. This matching model can be trained based on a deep learning framework. Then, the semantic role information of entity words E1 and E2 is input into the matching model. The matching model processes the matching relationship of semantic roles between entity words E1 and E2 through preset rules and can reflect the matching relationship of semantic roles between entity words E1 and E2 through the matching label. For example, in the matching model, if entity word E1 is labeled as the subject (Agent) and entity word E2 is labeled as the object (Patient), then there is a clear "Agent-Patient" matching relationship between them. The matching model processes this matching relationship to generate the matching label of entity words E1 and E2. The magnitude of the value in the matching label represents the intensity of the match. The larger the value, the higher the matching intensity of the matching relationship corresponding to the matching label, and vice versa. Then, the semantic dependency and the matching relationship of semantic roles (i.e., the matching label) between entity words E1 and E2 are input into a pre-trained inference engine (usually a fully connected layer). Through the inference of the inference engine, the semantic relationship between the two entity words E1 and E2 is inferred, and this semantic relationship is used as the explicit dependency relationship of the semantics between entity words E1 and E2. For example, if E1 and E2 have a high value in semantic dependency and their semantic roles (such as Agent-Patient) are a typical "action-receptor" relationship in the grammatical structure, then their semantic relationship can be inferred. Finally, by repeating the above method steps, the explicit dependency relationship of the semantics between each two entity words can be obtained.

[0036] It should be noted that the explicit dependency relationship of semantics in this application refers to the semantic relationship that can be directly deduced through clear language structures or rules, which reflects the direct semantic connection between two entity words in a specific language context.

[0037] In step 104, according to the context information of the English material, implicit structural semantic association analysis is performed on the entity words between each entity word cluster to obtain the semantic association degree between each entity word cluster. Furthermore, based on all the semantic association degrees and each annotation loss, dependency analysis is performed on the implicit semantic relationship between each two entity words in the English entity word library to obtain the implicit dependency relationship of the semantics between each two entity words.

[0038] In some embodiments, as shown in Figure 3 This figure is a schematic flowchart for determining the semantic association degree in some embodiments of this application. In this embodiment, the implicit structural semantic association analysis is performed on the entity words between each entity word cluster according to the context information of the English material. The steps to obtain the semantic association degree between each entity word cluster can be implemented as follows: In step 1041, for every two entity clusters, extract the semantic relationship of the entities between the two entity clusters in terms of syntactic structure based on the context information of the English materials; In step 1042, determine the semantic relationship of the entities between the two entity clusters in terms of word order structure according to the context information of the English materials; In step 1043, determine the semantic correlation degree between the two entity clusters through the semantic relationship in syntactic structure and the semantic relationship in word order structure, and then obtain the semantic correlation degrees between each entity cluster.

[0039] In specific implementation, for every two entity clusters, first, construct a syntactic analysis model based on a deep learning framework, such as SpaCy, use the context information of the English materials as the initialization parameters of the syntactic analysis model, and mark the entities in the two entity clusters in the context information of the English materials. Then, perform semantic analysis of the entities in the entity clusters in terms of syntactic structure through the syntactic analysis model, and extract the semantic relationship of the entities between the two entity clusters in terms of syntactic structure. For example, select an entity from one entity cluster as the selected entity, extract the matching degree in syntactic structure between the selected entity and each entity in the other entity cluster through the syntactic analysis model, and then take the average value of all matching degrees as the semantic relationship of the entities between the selected entity and the other entity cluster in terms of syntactic structure. Continue to determine the semantic relationship of the entities between the remaining selected entities and the other entity cluster in terms of syntactic structure, and then take the semantic relationship of the entities between all selected entities and the other entity cluster in terms of syntactic structure as the semantic relationship of the entities between the two entity clusters in terms of syntactic structure; then, the word order structure refers to the influence of the arrangement order of words in a sentence on semantics. In the analysis of semantic relationship based on word order, it usually depends on a language model (such as BERT, GPT and other Transformer architectures) to capture the word order information of the sentence. That is, the context information of the English materials can be used as the initialization parameters of the existing language model, and then analyze whether the word order information between the two entity clusters is relatively close through the language model, and take the quantization value of the closeness as the semantic relationship of the entities between the two entity clusters in terms of word order structure. For example, if the two entity clusters are close to each other in the sentence and their connection can be efficiently captured through the self-attention mechanism, it can be considered that there is a strong semantic association between them; finally, after determining the semantic relationship in syntactic structure and the semantic relationship in word order structure, the semantic correlation degree between the two entity clusters can be calculated by means of weighted fusion, that is, semantic correlation degree = w1 * semantic relationship in syntactic structure + w2 * semantic relationship in word order structure, where w1 and w2 are weights that can be adjusted according to the actual corpus or task, and w1 + w2 = 1; finally, the semantic correlation degrees between each entity cluster can be obtained through the above method.

[0040] It should be noted that the semantic correlation degree in this application is an index to measure the strength of the semantic association between two entity word clusters. The higher the semantic correlation degree, the closer the semantic connection between the two entity word clusters, the more frequent the information interaction, and vice versa, the weaker the correlation.

[0041] In some embodiments, based on all the semantic correlation degrees and each annotation loss, a dependency analysis is performed on the implicit semantic relationship between every two entity words in the English entity word library, and the implicit semantic dependency relationship between every two entity words can be realized by the following steps: For every two entity words, determine the association attention between the two entity words according to all the semantic correlation degrees and each annotation loss; Construct a machine model based on a neural network to capture the implicit semantic relationship between entity words; Set the association attention as the connection weight in the forward propagation of the machine model; Capture the implicit semantic dependency relationship between two entity words through the machine model, and then obtain the implicit semantic dependency relationship between every two entity words.

[0042] In specific implementation, for every two entity words, first, when the two entity words come from the same entity word cluster, obtain the annotation loss corresponding to the entity word cluster where the entity words are located, and use the natural exponential function value of the opposite number of the annotation loss as the association attention between the two entity words. When the two entity words come from different entity word clusters, obtain the semantic correlation degree between the two different entity word clusters and the annotation losses corresponding to the entity word clusters where the two entity words are located, and record the average value of the two annotation losses as the average loss. Then, use the sum of the natural exponential function value of the opposite number of the average loss and the semantic correlation degree as the association attention between the two entity words. Then, construct a machine model of a neural network model (such as a bidirectional LSTM or Transformer) to capture the implicit semantic relationship between words. In this process, the association attention is used as the connection weight of the machine model, which can guide the network to learn the deep semantic association between entity words. Finally, through forward propagation, the machine model generates a relationship matrix of the implicit semantics between two entity words to represent the dependency strength (i.e., the implicit dependency relationship) between the two entity words. Through the above steps, the implicit semantic dependency relationship between every two entity words can be obtained.

[0043] It should be noted that the implicit semantic dependency relationship in this application refers to the potential dependency relationship between two entity words at the semantic level. It reflects a deeper correlation at the semantic level and plays an important role in understanding implicit information and non-explicit semantic relationships.

[0044] In step 105, a knowledge graph for English teaching is constructed through the explicit and implicit semantic dependency relationships between every two entity words.

[0045] In some embodiments, constructing a knowledge graph for English teaching through the explicit and implicit semantic dependency relationships between every two entity words can be implemented by the following steps: Construct an entity-relationship graph with explicit semantics through the explicit semantic dependency relationships between every two entity words; Construct an entity-relationship graph with implicit semantics according to the implicit semantic dependency relationships between every two entity words; Associate and fuse the entity-relationship graph with explicit semantics and the entity-relationship graph with implicit semantics to obtain a knowledge graph for English teaching.

[0046] When specifically implemented, first, each entity word can be used as a node of the graph, and the explicit semantic dependency relationship between every two entity words can be used as the edge weight between the corresponding nodes. Then, the visualized graph structure composed of all nodes is used as the entity-relationship graph with explicit semantics; then, through tensor operations, the implicit semantic dependency relationship between every two entity words can be mapped to the relationship edges between the corresponding nodes, and the visualized graph structure composed of all nodes and all relationship edges is used as the entity-relationship graph with implicit semantics; finally, using graph embedding techniques (such as Node2Vec or GraphSAGE), the entity-relationship graph with explicit semantics and the entity-relationship graph with implicit semantics are respectively embedded into the vector space. By calculating the similarity of the node embedding vectors, the nodes and edges of the two graphs are matched and merged. At the same time, the weights of the explicit and implicit edges are integrated using the method of weighted averaging to generate a comprehensive knowledge graph. This fused knowledge graph can not only support English teaching through entity query and relationship path analysis, but also expand the connectivity and understanding depth of knowledge through implicit semantic relationships.

[0047] On the other hand, in some embodiments, the present application provides an English knowledge graph construction system. Refer to Figure 4 , which is a schematic structural diagram of the English knowledge graph construction system shown in some embodiments of the present application. The English knowledge graph construction system 400 includes: a construction module 401, a processing module 402, and an execution module 403, which are described as follows: Construction module 401. In the present application, the construction module 401 is mainly used to construct an English entity word library for teaching based on multi-source heterogeneous English materials; Processing module 402. In the present application, the processing module 402 is used to perform semantic role annotation on each entity word in the English entity word library, classify the corresponding entity words according to the annotated semantic role information to obtain entity word clusters with different semantic roles, and further determine the annotation loss of the semantic role of each entity word cluster; In this application, the processing module 402 is further configured to determine the semantic dependency between every two entity words according to the co-occurrence probability and word distance between every two entity words in the English entity word library, and then determine the explicit semantic dependency relationship between every two entity words from the semantic dependencies of all semantics and the semantic role information of each entity word; In this application, the processing module 402 is further configured to perform implicit structural semantic association analysis on the entity words between each entity word cluster according to the context information of the English material, obtain the semantic association degree between each entity word cluster, and then perform dependency analysis on the implicit semantic relationship between every two entity words in the English entity word library based on all the semantic association degrees and each annotation loss to obtain the implicit semantic dependency relationship between every two entity words; An execution module 403. In this application, the execution module 403 is mainly configured to construct a knowledge graph for English teaching through the explicit dependency relationship and implicit dependency relationship of semantics between every two entity words.

[0048] In addition, this application also provides a computer device, which includes a memory and a processor. The memory stores code, and the processor is configured to obtain the code and execute the above-mentioned English knowledge graph construction method.

[0049] In some embodiments, refer to Figure 5 , this figure is a schematic structural diagram of a computer device for implementing the English knowledge graph construction method according to some embodiments of this application. The English knowledge graph construction method in the above embodiments can be implemented by Figure 5 the computer device shown. The computer device 500 includes at least one processor 501, a communication bus 502, a memory 503, and at least one communication interface 504.

[0050] The processor 501 can be a general-purpose central processing unit (CPU) or an application-specific integrated circuit (ASIC).

[0051] The communication bus 502 can be used to transmit information between the above components.

[0052] The memory 503 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disks or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 503 can exist independently and be connected to the processor 501 through the communication bus 502. The memory 503 can also be integrated with the processor 501.

[0053] Among them, the memory 503 is used to store the program code for executing the solution of this application and is controlled by the processor 501 for execution. The processor 501 is used to execute the program code stored in the memory 503. The program code can include one or more software modules. The above-mentioned method for constructing an English knowledge graph in the embodiment can be implemented by one or more software modules in the processor 501 and the program code in the memory 503.

[0054] The communication interface 504 uses any device such as a transceiver to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.

[0055] In a specific implementation, as an embodiment, the computer device can include multiple processors, and each of these processors can be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0056] The computer device described above can be a general-purpose computer device or a special-purpose computer device. In specific implementations, the computer device can be a desktop computer, a laptop computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. The embodiments of the present application do not limit the type of the computer device.

[0057] In addition, the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned English knowledge graph construction method is implemented.

[0058] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application.

[0059] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A method for constructing an English knowledge graph, characterized in that: The steps include: Constructing English entity word bank for teaching based on multi-source heterogeneous English materials; Annotate the semantic roles of the entity words in the English entity word library, and classify the corresponding entity words according to the annotated semantic role information to obtain entity word clusters with different semantic roles, and then determine the annotation loss of the semantic role of each entity word cluster; Determine the semantic dependency between each two entity words according to the co-occurrence probability and the inter-word distance between each two entity words in the English entity word library, and then determine the explicit semantic dependency relationship between each two entity words according to all the semantic dependencies and the semantic role information of each entity word; Performing semantic association analysis of implicit structures on entity words between entity word clusters according to the context information of the English data to obtain the semantic association between entity word clusters, and then performing dependency analysis on the implicit semantic relationship between every two entity words in the English entity word library based on all the semantic associations and each annotation loss to obtain the implicit semantic dependency relationship between every two entity words; A knowledge graph for English teaching is constructed through the explicit and implicit semantic dependencies between every two entity words.

2. The method according to claim 1, characterized in that Each entity word in the English entity word library is annotated with a semantic role, and the corresponding entity words are classified according to the annotated semantic role information, and entity word clusters with different semantic roles are obtained, which specifically include: Build an annotation model for semantic role labeling based on deep learning; Using each entity word in the English entity word library as an input parameter of the annotation model; Annotating the semantic role information of each entity word by using the annotation model; According to the semantic role information of each entity word, all entity words are divided into entity word clusters with different semantic roles.

3. The method according to claim 1, characterized in that The annotation loss for determining the semantic role of each entity cluster specifically includes: For each entity word cluster, all entity words in the entity word cluster are obtained, and then the semantic role information of each entity word is evaluated for error to obtain the role labeling error of each entity word; The labeling loss of the semantic role of the entity word cluster is determined by the role labeling error of all entity words, and then the labeling loss of the semantic role of each entity word cluster is obtained.

4. The method according to claim 1, characterized in that Determining the semantic dependency between each two entity words according to the co-occurrence probability and the inter-word distance between each two entity words in the English entity word library specifically includes: For every two entity words in the English entity lexicon; Determine the co-occurrence probability and inter-word distance between two entity words; Determine the co-occurrence coefficient between two entity words by using the co-occurrence probability and the inter-word distance; Using the co-occurrence coefficient as a constraint item of an objective function in a preset semantic relationship model; The semantic dependency between two entity words is extracted through the semantic relationship model, and then the semantic dependency between every two entity words is obtained.

5. The method according to claim 1, characterized in that The explicit semantic dependency between each two entity words is determined by all semantic dependencies and the semantic role information of each entity word, including: For every two entity words, the matching relationship of the semantic roles between the two entity words is determined according to the semantic role information of the two entity words; The explicit semantic dependency between two entity words is determined by the semantic dependency between the two entity words and the matching relationship between the semantic roles, and then the explicit semantic dependency between every two entity words is obtained.

6. The method according to claim 1, characterized in that According to the context information of the English data, the entity words between each entity word cluster are subjected to implicit structure semantic association analysis, and the semantic association degree between each entity word cluster is obtained, which specifically includes: For each two entity word clusters, extracting the semantic relationship of the entity words in the syntactic structure between the two entity word clusters according to the context information of the English data; Determining the semantic relationship of entity words in word order structure between two entity word clusters according to the context information of the English material; The semantic relevance between two entity word clusters is determined by the semantic relationship in syntactic structure and the semantic relationship in word order structure, and then the semantic relevance between each entity word cluster is obtained.

7. The method according to claim 1, characterized in that The knowledge graph for English teaching is constructed through the explicit and implicit semantic dependencies between every two entity words, including: Construct an entity-relationship graph with explicit semantics through the explicit semantic dependency between every two entity words; Construct an entity-relationship graph with implicit semantics based on the implicit semantic dependency between every two entity words; The entity-relationship graph of explicit semantics and the entity-relationship graph of implicit semantics are associated and fused to obtain a knowledge graph for English teaching.

8. An English knowledge graph construction system, characterized in that: include: A construction module is used to construct an English entity vocabulary for teaching based on multi-source heterogeneous English materials; A processing module, used to annotate the semantic roles of the entity words in the English entity word library, and classify the corresponding entity words according to the annotated semantic role information to obtain entity word clusters with different semantic roles, and then determine the annotation loss of the semantic role of each entity word cluster; The processing module is further used to determine the semantic dependency between each two entity words according to the co-occurrence probability and the inter-word distance between each two entity words in the English entity word library, and then determine the explicit semantic dependency relationship between each two entity words according to all the semantic dependencies and the semantic role information of each entity word; The processing module is further used to perform semantic association analysis of the implicit structure of the entity words between each entity word cluster according to the context information of the English material, to obtain the semantic association between each entity word cluster, and then perform dependency analysis on the implicit semantic relationship between every two entity words in the English entity word library based on all the semantic associations and each annotation loss, to obtain the implicit semantic dependency relationship between every two entity words; The execution module is used to construct a knowledge graph for English teaching through the explicit and implicit semantic dependencies between every two entity words.

9. A computer device, comprising a memory and a processor, wherein the memory stores a code, characterized in that: The processor is configured to obtain the code and execute the English knowledge graph construction method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for constructing an English knowledge graph as described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Association method and system for implicit semantics and explicit terms of traditional Chinese medicine ancient books and medium

    CN120611717A

  • Teaching resource recommendation method and system based on natural language processing

    CN121724811A