Knowledge graph construction method for English vocabulary semantic comprehension and extension
By constructing semantic models and context models, combining semantic association networks, and using historical semantic sequences and theoretical semantic parameters for multi-dimensional analysis and parameter adjustment, the problems of semantic ambiguity and inaccurate association when dealing with polysemous words or words with strong contextual dependence in knowledge graphs are solved, thereby improving construction efficiency and semantic expression capabilities.
Patent Information
- Application Number
- CN202510647898.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-09-19
AI Technical Summary
Existing knowledge graphs have problems with semantic ambiguity and inaccurate associations when dealing with polysemous words or words with strong context dependence, and their construction efficiency is low.
By acquiring the semantic features and contextual features of the target vocabulary set, constructing semantic models and contextual models, and combining them with the semantic association network, a comprehensive semantic model is generated. Multi-dimensional analysis and parameter adjustment are performed using historical semantic sequences and theoretical semantic parameters to optimize the semantic expansion of vocabulary.
It improves the semantic expression ability and construction efficiency of knowledge graphs when dealing with polysemous words or words with strong context dependence, and solves the problems of semantic ambiguity and inaccurate associations.
Smart Images

Figure CN120671672A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of natural language processing and knowledge graph technology, and specifically provides a method for constructing a knowledge graph for semantic understanding and expansion of English vocabulary. Background Art
[0002] A knowledge graph is a semantic network used to describe entities and their relationships. It is widely used in fields such as information retrieval, natural language processing, and data mining. In English vocabulary learning, knowledge graphs can help learners understand the meaning of words and expand their application scenarios through semantic associations. Existing knowledge graph construction methods typically rely on dictionary data, corpora, or machine learning algorithms. They analyze word co-occurrence relationships, contextual features, and semantic similarity to generate a network of associations between words.
[0003] However, in practical applications, the effectiveness of knowledge graph construction is closely related to the data quality and algorithm accuracy used. High-quality data and high-precision algorithms can significantly improve the semantic expressiveness of knowledge graphs, but over-reliance on complex algorithms and large-scale data can increase the system's computational burden and affect construction efficiency. Furthermore, existing methods can suffer from semantic ambiguity or inaccurate associations when dealing with polysemous words or words with strong contextual dependencies, thus limiting the further application of knowledge graphs in English vocabulary learning. Summary of the Invention
[0004] In order to make up for the deficiencies of the prior art, at least one technical problem raised in the background technology is solved.
[0005] The technical solution adopted by the present invention to solve its technical problems is: the purpose of the present invention is to provide a knowledge graph construction method and system for semantic understanding and expansion of English vocabulary, aiming to solve the problems of semantic ambiguity and inaccurate association in the knowledge graph in the existing technology when processing polysemous words or words with strong context dependence, as well as the shortcomings of low construction efficiency caused by data quality and algorithm complexity.
[0006] The present invention is implemented as follows. In the first aspect, the present invention provides a method for constructing a knowledge graph for semantic understanding and expansion of English vocabulary, comprising: obtaining vocabulary and its contextual information in a target vocabulary set, extracting semantic features and contextual features of each vocabulary, and constructing a semantic model and a contextual model of the vocabulary according to the semantic features and contextual features respectively; fusing the semantic model and the contextual model to generate a comprehensive semantic model of the vocabulary, and establishing a semantic association network between the vocabulary according to the similarity calculation results between the comprehensive semantic models; wherein the semantic model is used to describe the core semantic attributes of the vocabulary, the contextual model is used to describe the semantic change rules of the vocabulary in different contexts, and the comprehensive semantic model is used to fully express the semantic characteristics of the vocabulary; continuously collecting the frequency of occurrence and semantic distribution of the vocabulary in the target vocabulary set in different corpora, and sorting the collected vocabulary in chronological order. The frequency of occurrence and semantic distribution of words in different corpora are combined and sorted to obtain the historical semantic sequence of each word; the current semantic extension demand is obtained, and the current semantic extension demand is distributed through the semantic association network to obtain the theoretical semantic parameters of each comprehensive semantic model corresponding to the current semantic extension demand; each theoretical semantic parameter is substituted into the corresponding comprehensive semantic model, and each comprehensive semantic model is made to perform multi-dimensional semantic analysis on each theoretical semantic parameter according to the corresponding historical semantic sequence to obtain the semantic analysis parameters of each comprehensive semantic model; the semantic analysis parameters of each comprehensive semantic model are substituted into the semantic association network, and the semantic association network is made to perform parameter adjustment on each theoretical semantic parameter to obtain the parameter adjustment result, and the words in the target vocabulary set are semantically expanded according to the parameter adjustment result.
[0007] Preferably, the steps of obtaining the current semantic extension demand and allocating and processing the current semantic extension demand through a semantic association network to obtain theoretical semantic parameters of each comprehensive semantic model corresponding to the current semantic extension demand include: collecting data from a corpus related to a target vocabulary set to obtain corpus data in the corpus, and analyzing and processing the corpus data on semantic topics and semantic extension conditions of semantic topics to obtain semantic extension conditions of semantic topics in the corpus, and using the semantic extension conditions as the current semantic extension demand; splitting the current semantic extension demand through each comprehensive semantic model in the semantic association network to obtain split semantic requirements of each comprehensive semantic model corresponding to the current semantic extension demand; and reversely deducing each corresponding split semantic requirement according to each comprehensive semantic model to obtain the theoretical semantic parameters required for each comprehensive semantic model to execute the split semantic requirement.
[0008] Preferably, each theoretical semantic parameter is substituted into the corresponding comprehensive semantic model, and each comprehensive semantic model is made to perform multi-dimensional semantic analysis on each theoretical semantic parameter according to the corresponding historical semantic sequence to obtain the semantic analysis parameters of each comprehensive semantic model. The steps include: substituting each theoretical semantic parameter into the corresponding comprehensive semantic model, and making the comprehensive semantic model search for the same parameters in the corresponding historical semantic sequence according to the theoretical semantic parameters to obtain the semantic parameters that are the same as the theoretical semantic parameters in the historical semantic sequence, and marking the semantic parameters and the corresponding semantic distribution as first reference data; integrating each continuous first reference data into a first reference paragraph according to the time corresponding to each first reference data to obtain a plurality of first reference paragraphs, and performing a search for each first reference paragraph. Perform multi-dimensional analysis respectively to obtain analysis parameters of the first reference data; take the theoretical semantic parameters as a benchmark, obtain several neighboring semantic parameters adjacent to the theoretical semantic parameters, and let the comprehensive semantic model search for the same parameters in the corresponding historical semantic sequence according to the neighboring semantic parameters to obtain semantic parameters identical to the neighboring semantic parameters in the historical semantic sequence, and mark the semantic parameters and the corresponding semantic distribution as second reference data; according to the time corresponding to each second reference data, integrate the consecutive second reference data into a second reference paragraph to obtain several second reference paragraphs, and perform multi-dimensional analysis on each second reference paragraph to obtain analysis parameters of the second reference data; the analysis parameters of the first reference data and the analysis parameters of the second reference data together constitute the semantic analysis parameters of the comprehensive semantic model.
[0009] Preferably, the step of performing a multi-dimensional analysis on each first reference paragraph to obtain analysis parameters of the first reference data includes: performing mean calculation on the first reference data of each first reference paragraph in the same time period, and performing difference calculation on the result of the mean calculation and the semantic distribution of the corresponding theoretical semantic parameter under an ideal state to obtain the degree of deviation between the semantic distribution of the theoretical semantic parameter generated by the time period and the semantic distribution under the ideal state, combining each deviation degree with the time period of the corresponding first reference data, and arranging each combination in chronological order to obtain a first stable feature sequence of the semantic parameter; performing a comparative analysis on the first reference data of each first reference paragraph in the same time period to obtain the degree of deviation of the first reference data of each first reference paragraph in the same time period, combining each deviation degree with the time period of the corresponding first reference data, and arranging each combination in chronological order to obtain a second stable feature sequence of the semantic parameter; superimposing the first stable feature sequence and the second stable feature sequence to summarize the deviation degrees corresponding to the same time period in the first stable feature sequence and the second stable feature sequence into the same combination as the first analysis parameter of the first reference data.
[0010] Preferably, the step of performing a multi-dimensional analysis on each second reference paragraph to obtain analysis parameters of the second reference data includes: performing mean calculation on the first reference data of each second reference paragraph in the same time period, and performing difference calculation on the result of the mean calculation and the semantic distribution of the corresponding theoretical semantic parameter under an ideal state to obtain the degree of deviation between the semantic distribution of the theoretical semantic parameter generated by the time period and the semantic distribution under the ideal state, combining each deviation degree with the time period of the corresponding first reference data, and arranging each combination in chronological order to obtain a third stable feature sequence of the semantic parameter; performing a comparative analysis on the first reference data of each first reference paragraph in the same time period to obtain the degree of deviation of the first reference data of each first reference paragraph in the same time period, combining each deviation degree with the time period of the corresponding first reference data, and arranging each combination in chronological order to obtain a fourth stable feature sequence of the semantic parameter; superimposing the third stable feature sequence and the fourth stable feature sequence to summarize the deviation degrees corresponding to the same time period in the third stable feature sequence and the fourth stable feature sequence into the same combination as the second analysis parameter of the first reference data.
[0011] Preferably, the steps of substituting the semantic analysis parameters of each comprehensive semantic model into the semantic association network and making the semantic association network perform parameter adjustment processing on each theoretical semantic parameter include: substituting the first analysis parameter of the first reference data of each comprehensive semantic model into the semantic association network, making the semantic association network perform overall semantic operation according to the first analysis parameter of the first reference data of each comprehensive semantic model to obtain a baseline semantic graph of the semantic association network; taking the baseline semantic graph of the semantic association network as a benchmark, substituting the second analysis parameters of each second reference data of each comprehensive semantic model into the semantic association network respectively, making the semantic association network perform overall semantic operation to obtain each neighboring semantic graph of the semantic association network; performing comparative analysis on the baseline semantic graph and each neighboring semantic graph to obtain the result of the comparative analysis, determining the baseline semantic graph or neighboring semantic graph with the best effect according to the result of the comparative analysis, adjusting the semantic parameters corresponding to the baseline semantic graph or the neighboring semantic graph to the theoretical semantic parameters to obtain the parameter adjustment result.
[0012] In the second aspect, the present invention provides a knowledge graph construction system for semantic understanding and expansion of English vocabulary, including: a model construction module, which is used to obtain vocabulary and its context information in the target vocabulary set, extract the semantic features and context features of each vocabulary, and respectively construct a semantic model and a context model of the vocabulary according to the semantic features and context features, fuse the semantic model and the context model to generate a comprehensive semantic model of the vocabulary, and establish a semantic association network between the vocabulary according to the similarity calculation results between the comprehensive semantic models; wherein the semantic model is used to describe the core semantic attributes of the vocabulary, the context model is used to describe the semantic change rules of the vocabulary in different contexts, and the comprehensive semantic model is used to fully express the semantic characteristics of the vocabulary; a data acquisition module is used to continuously collect the occurrence frequency and semantic distribution of the vocabulary in the target vocabulary set in different corpora, and chronologically sort the collected vocabulary in different corpora. The frequency of occurrence and semantic distribution are combined and sorted to obtain the historical semantic sequence of each word; a preliminary allocation module is used to obtain the current semantic extension demand, and allocate the current semantic extension demand through the semantic association network to obtain the theoretical semantic parameters of each comprehensive semantic model corresponding to the current semantic extension demand; a semantic analysis module is used to substitute each theoretical semantic parameter into the corresponding comprehensive semantic model, and make each comprehensive semantic model perform multi-dimensional semantic analysis on each theoretical semantic parameter according to the corresponding historical semantic sequence to obtain the semantic analysis parameters of each comprehensive semantic model; a parameter adjustment module is used to substitute the semantic analysis parameters of each comprehensive semantic model into the semantic association network, and make the semantic association network perform parameter adjustment on each theoretical semantic parameter to obtain the parameter adjustment result, and semantically expand the words in the target vocabulary set according to the parameter adjustment result.
[0013] The present invention provides a method for constructing a knowledge graph for semantic understanding and expansion of English vocabulary, which has the following beneficial effects: the present invention constructs a semantic model, a context model and a comprehensive semantic model for the vocabulary in the target vocabulary set, and combines the semantic association network to generate semantic relationships between the vocabulary, which can comprehensively reflect the core semantic attributes of the vocabulary and its changing patterns in different contexts. By collecting the frequency of occurrence and semantic distribution of vocabulary in different corpora, a historical semantic sequence is formed, and the actual semantic performance of the vocabulary is further analyzed. On this basis, the semantic association network is used to adjust the theoretical semantic parameters and optimize the semantic expansion effect of the vocabulary, which solves the problems of semantic ambiguity and inaccurate association in the knowledge graph in the prior art when dealing with polysemous words or vocabulary with strong context dependence, and at the same time improves the construction efficiency and semantic expression ability of the knowledge graph. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1This is a flow chart of the knowledge graph construction method for English vocabulary semantic understanding and expansion in an embodiment of the present invention, showing the overall steps from data collection to semantic expansion.
[0015] Figure 2 This is a logical diagram of the generation of a comprehensive semantic model and the construction of a semantic association network in an embodiment of the present invention, focusing on the process of generating a comprehensive semantic model after the fusion of the semantic model and the context model and their association relationships in the network.
[0016] Figure 3 This is a schematic diagram of the structure of the historical semantic sequence analysis process in an embodiment of the present invention, which describes the specific implementation method of combining and sorting the semantic distribution of words in chronological order.
[0017] Figure 4 Schematic diagram of the module for adjusting the theoretical semantic parameters in an embodiment of the present invention, showing the process of analyzing and optimizing the theoretical semantic parameters by the semantic association network.
[0018] The accompanying drawings are numbered as follows:
[0019] 1. Target vocabulary set; 2. Semantic model; 3. Context model; 4. Comprehensive semantic model; 5. Semantic association network; 6. Historical semantic sequence; 7. Theoretical semantic parameters; 8. Semantic analysis parameters; 9. Parameter adjustment results. DETAILED DESCRIPTION
[0020] The present invention provides a method and system for constructing a knowledge graph for semantic understanding and expansion of English vocabulary. The core of the method is to expand the vocabulary semantically by combining the semantic association network through the generation and fusion of semantic models, context models and comprehensive semantic models. Figure 1 To the attached Figure 4 and the component numbers in the figure marks are described in detail.
[0021] The first step is to obtain target vocabulary set 1. Target vocabulary set 1 forms the foundation for the entire knowledge graph construction and contains the English vocabulary to be processed and its associated contextual information. In practical applications, target vocabulary set 1 can be formed by extracting high-frequency vocabulary or specialized vocabulary from a specific field from a corpus. For example, in an educational scenario, target vocabulary set 1 may include core vocabulary that students need to master and the contextual sentences in which they appear in textbooks or exercises. This vocabulary and contextual information are passed to the model construction module for the subsequent generation of semantic model 2 and context model 3. The model construction module first extracts features for each word in target vocabulary set 1. The extracted features include the core semantic attributes of the word and the semantic variation patterns in context. Core semantic attributes refer to the basic meaning of a word in isolation, while the semantic variation patterns in context reflect the dynamic behavior of the word in different contexts. After feature extraction, semantic model 2 and context model 3 are generated. Semantic model 2 primarily describes the core semantic attributes of a word, while context model 3 captures the semantic variation patterns of a word in different contexts. The generation process of these two models relies on word vector representation and context encoding technology in natural language processing technology, such as using pre-trained language models to embed vocabulary and analyzing the semantic distribution of vocabulary through context windows.
[0022] Next, the semantic model 2 and the context model 3 are further fused to generate a comprehensive semantic model 4. Figure 2 As shown, the generation process of the comprehensive semantic model 4 is achieved by weighted fusion of the feature vectors of the semantic model 2 and the context model 3. Specifically, the weight of the feature vector of the semantic model 2 is determined by the semantic stability of the vocabulary in an isolated state, while the weight of the feature vector of the context model 3 is determined by the frequency of semantic changes of the vocabulary in different contexts. This weighted fusion method can ensure that the comprehensive semantic model 4 can not only reflect the core semantic attributes of the vocabulary, but also reflect its dynamic semantic expression in different contexts. After the comprehensive semantic model 4 is generated, the semantic association network 5 is established by calculating the similarity between each vocabulary. The semantic association network 5 is a graph structure in which the nodes represent vocabulary and the edges represent the semantic association strength between vocabulary. The calculation of the semantic association strength is based on the cosine similarity or other similarity measurement methods between the comprehensive semantic models 4. In this way, the semantic association network 5 can fully reflect the semantic relationship between vocabulary, thereby providing a basis for subsequent semantic expansion.
[0023] In order to further improve the semantic expression ability of the knowledge graph, the data collection module continuously collects the frequency and semantic distribution of the words in the target vocabulary set 1 in different corpora. This process is achieved through crawler technology or corpus retrieval tools. The collected data includes the frequency of occurrence, semantic distribution and context information of the words in different time periods. The collected data is combined and sorted in chronological order to generate a historical semantic sequence 6. As shown in the attached figure, Figure 3 As shown, the generation process of historical semantic sequence 6 involves two steps: first, the frequency of occurrence and semantic distribution of words in different time periods are segmented, with each segment corresponding to a time period; second, these segmented data are arranged in chronological order to form a continuous time series. The purpose of historical semantic sequence 6 is to record the semantic evolution of words in different time periods, providing a reference for subsequent semantic expansion.
[0024] The preliminary allocation module is responsible for acquiring the current semantic extension demand and allocating it through the semantic association network 5 to obtain the theoretical semantic parameters 7 corresponding to each comprehensive semantic model 4. Specifically, the current semantic extension demand is obtained by collecting and analyzing data from relevant corpora. The analysis includes the semantic topics within the corpus and their semantic extension conditions. For example, in a news corpus, semantic topics may include "economy" and "technology," while semantic extension conditions may involve the semantic change trends of these topics over different time periods. By analyzing the semantic topics and semantic extension conditions within the corpus, the specific content of the current semantic extension demand can be determined. Subsequently, the current semantic extension demand is split using the various comprehensive semantic models 4 in the semantic association network 5 to obtain the split semantic requirements corresponding to each comprehensive semantic model 4. The purpose of splitting semantic requirements is to decompose complex semantic extension requirements into multiple simple sub-requirements for easier subsequent processing. Finally, the corresponding split semantic requirements are reverse-derived based on each comprehensive semantic model 4 to obtain the theoretical semantic parameters 7 required for each comprehensive semantic model 4 to implement the split semantic requirements. Theoretical semantic parameters 7 are the basis for subsequent semantic analysis, and their accuracy directly affects the effect of semantic expansion.
[0025] The semantic analysis module substitutes each theoretical semantic parameter 7 into the corresponding comprehensive semantic model 4, and makes each comprehensive semantic model 4 perform multi-dimensional semantic analysis on each theoretical semantic parameter 7 according to the corresponding historical semantic sequence 6 to obtain the semantic analysis parameter 8 of each comprehensive semantic model 4. Figure 4As shown, the semantic analysis process includes the following steps: First, the theoretical semantic parameter 7 is substituted into the comprehensive semantic model 4, and semantic parameters identical to the theoretical semantic parameter 7 are searched in the historical semantic sequence 6. These semantic parameters and their corresponding semantic distributions are recorded and marked as first reference data. The purpose of the first reference data is to reflect the actual performance of the theoretical semantic parameter 7 in the historical semantic sequence 6. Second, based on the time corresponding to the first reference data, the continuous first reference data are integrated into a first reference paragraph to obtain several first reference paragraphs. Each first reference paragraph is subjected to a multi-dimensional analysis to obtain analysis parameters for the first reference data. The multi-dimensional analysis includes the mean semantic distribution of the first reference paragraph within the same time period and the degree of deviation from the ideal semantic distribution. These analysis parameters are summarized into a first stable feature sequence and a second stable feature sequence, which ultimately constitute the first analysis parameters of the first reference data. Similarly, using the theoretical semantic parameter 7 as a benchmark, several neighboring semantic parameters adjacent to the theoretical semantic parameter 7 are obtained. Semantic parameters identical to the neighboring semantic parameters are searched in the historical semantic sequence 6 and marked as second reference data. By performing a similar multi-dimensional analysis on the second reference data, a third stable feature sequence and a fourth stable feature sequence of the second reference data are obtained, which ultimately together constitute the second analysis parameters of the second reference data. The first analysis parameters of the first reference data and the second analysis parameters of the second reference data together constitute semantic analysis parameters 8 of the comprehensive semantic model 4.
[0026] The parameter adjustment module substitutes the semantic analysis parameters 8 of each comprehensive semantic model 4 into the semantic association network 5, and instructs the semantic association network 5 to perform parameter adjustment processing on each theoretical semantic parameter 7 to obtain parameter adjustment results 9. The parameter adjustment process includes the following steps: First, the first analysis parameter of the first reference data of each comprehensive semantic model 4 is substituted into the semantic association network 5, and the semantic association network 5 is instructed to perform overall semantic operations based on the first analysis parameters to obtain a baseline semantic graph of the semantic association network 5. The baseline semantic graph is the basic state of the semantic association network 5 before parameter adjustment. Second, based on the baseline semantic graph, the second analysis parameter of the second reference data of each comprehensive semantic model 4 is substituted into the semantic association network 5, and the semantic association network 5 is instructed to perform overall semantic operations to obtain each neighboring semantic graph of the semantic association network 5. The neighboring semantic graph reflects the semantic performance of the semantic association network 5 under different parameter adjustment conditions. Finally, a comparative analysis is performed between the baseline semantic graph and each neighboring semantic graph to obtain the comparative analysis results. Based on the results of the comparative analysis, the optimal baseline semantic map or neighboring semantic map is determined, and the semantic parameters corresponding to the optimal semantic map are adjusted against the theoretical semantic parameters 7 to obtain parameter adjustment results 9. Parameter adjustment results 9 are the core output of the entire semantic expansion process and directly determine the semantic expansion effect of the vocabulary in the target vocabulary set 1.
[0027] Through the above steps, the present invention achieves the construction of a knowledge graph for the semantic understanding and expansion of English vocabulary. Throughout this process, the target vocabulary set 1 serves as input data. Through the generation of the semantic model 2, context model 3, and comprehensive semantic model 4, combined with the construction of the semantic association network 5, a comprehensive understanding and expansion of lexical semantics is ultimately achieved. The introduction of the historical semantic sequence 6 enables the recording and analysis of the semantic evolution patterns of vocabulary, while the generation of theoretical semantic parameters 7, semantic analysis parameters 8, and parameter adjustment results 9 ensures the accuracy and effectiveness of semantic expansion.
[0028] In order to better enable relevant personnel in this technical field to fully understand and implement the present invention, the specific implementation principle of the present invention is supplemented below with reference to a specific application scenario.
[0029] In practical applications, taking an English vocabulary learning system as an example, target vocabulary set 1 can be formed by extracting high-frequency words and their contextual information from student textbooks or corpora on online learning platforms. For example, target vocabulary set 1 might include the polysemous word "bank" and record its usage in different contexts, such as "river bank" and "moneybank." This vocabulary and its contextual information are passed to the model building module to generate semantic model 2 and context model 3. Semantic model 2 analyzes the core semantic properties of the words and determines that the basic meaning of "bank" in isolation is "river bank" or "bank." Context model 3 analyzes the semantic variation of words in different contexts and captures the tendency of "bank" to mean "river bank" in "river bank" and "bank" in "money bank." The generation of these two models relies on word vector representation and context encoding techniques in natural language processing. For example, a pre-trained language model is used to embed words and analyze the semantic distribution of words through a context window.
[0030] Next, semantic model 2 and context model 3 are weightedly fused to generate a comprehensive semantic model 4. Specifically, the feature vector weights of semantic model 2 are determined by the semantic stability of the vocabulary in isolation, while the feature vector weights of context model 3 are determined by the frequency of semantic changes of the vocabulary in different contexts. For example, "bank" is more likely to mean "bank" in isolation, but may be more likely to mean "river bank" in a specific context. This weighted fusion approach ensures that comprehensive semantic model 4 can reflect both the core semantic attributes of the vocabulary and its dynamic performance in different contexts. Subsequently, a semantic association network 5 is established by calculating the cosine similarity between each vocabulary. In semantic association network 5, the semantic association strength between "bank" and "river" is high, and the semantic association strength between "bank" and "money" is also high, thus comprehensively reflecting the semantic relationship between the vocabulary.
[0031] To further enhance the semantic expression capabilities of the knowledge graph, the data collection module continuously collects the frequency and semantic distribution of the words in the target vocabulary set 1 across different corpora. For example, crawler technology is used to collect the frequency and semantic distribution of the word "bank" across different time periods from news websites, social media, and academic papers. The collected data is segmented and processed chronologically, with each segment corresponding to a time period. The data is then arranged chronologically to form a historical semantic sequence 6. The historical semantic sequence 6 records the semantic evolution of "bank" across different time periods. For example, during the economic crisis, "bank" was more commonly associated with "money," while in environmentally relevant corpora, "bank" was more commonly associated with "river."
[0032] The preliminary allocation module is responsible for obtaining the current semantic extension requirements and allocating the requirements through the semantic association network 5. For example, in an educational scenario, the current semantic extension requirements may include helping students understand the specific meaning of "bank" in different contexts. By collecting and analyzing data from the relevant corpus, the semantic theme of "bank" and its extension conditions can be determined. Subsequently, the current semantic extension requirements are split through the comprehensive semantic model 4 in the semantic association network 5 to obtain split semantic requirements. For example, the semantic extension requirements of "bank" are split into two sub-requirements: "river bank" and "bank". Finally, the split semantic requirements are reversely deduced based on the comprehensive semantic model 4 to obtain the theoretical semantic parameters 7. Theoretical semantic parameters 7 are the basis for subsequent semantic analysis, and their accuracy directly affects the effect of semantic extension.
[0033] The semantic analysis module substitutes the theoretical semantic parameters 7 into the comprehensive semantic model 4 and performs a multi-dimensional analysis on the theoretical semantic parameters 7 based on the historical semantic sequence 6. For example, the theoretical semantic parameters 7 of "bank" are substituted into the comprehensive semantic model 4, and the historical semantic sequence 6 is searched for semantic parameters identical to the theoretical semantic parameters 7, which are marked as the first reference data. The first reference data reflects the actual performance of "bank" in the historical semantic sequence 6. Subsequently, the continuous first reference data are integrated into the first reference paragraph, and a multi-dimensional analysis is performed on this. For example, the semantic distribution mean of the first reference paragraph within the same time period and its degree of deviation from the ideal semantic distribution are analyzed to obtain the first stable feature sequence and the second stable feature sequence. Similarly, using the theoretical semantic parameters 7 as a benchmark, adjacent semantic parameters are obtained, and the historical semantic sequence 6 is searched for semantic parameters identical to the adjacent semantic parameters, which are marked as the second reference data. A similar multi-dimensional analysis is performed on the second reference data to obtain the third and fourth stable feature sequences. Finally, the first analysis parameters of the first reference data and the second analysis parameters of the second reference data together constitute the semantic analysis parameters 8 of the comprehensive semantic model 4.
[0034] The parameter adjustment module substitutes the semantic analysis parameters 8 into the semantic association network 5, and performs parameter adjustment processing on the theoretical semantic parameters 7. For example, the first analysis parameter of the first reference data of the comprehensive semantic model 4 is substituted into the semantic association network 5 to obtain a baseline semantic map. The baseline semantic map reflects the basic state of the semantic association network 5 before the parameter adjustment. Subsequently, the second analysis parameter of the second reference data of the comprehensive semantic model 4 is substituted into the semantic association network 5 to obtain a neighboring semantic map. The neighboring semantic map reflects the semantic performance of the semantic association network 5 under different parameter adjustment conditions. Finally, a comparative analysis is performed on the baseline semantic map and the neighboring semantic map to determine the semantic map with the best effect, and the semantic parameters corresponding to the best semantic map are adjusted to the theoretical semantic parameters 7 to obtain a parameter adjustment result 9. The parameter adjustment result 9 directly determines the semantic expansion effect of the vocabulary in the target vocabulary set 1.
[0035] Through the above steps, the present invention realizes the construction of a knowledge graph for the semantic understanding and expansion of English vocabulary. In the whole process, the target vocabulary set 1 is used as input data, and through the generation of semantic model 2, context model 3 and comprehensive semantic model 4, combined with the construction of semantic association network 5, a comprehensive understanding and expansion of vocabulary semantics is finally achieved. The introduction of historical semantic sequence 6 enables the semantic evolution law of vocabulary to be recorded and analyzed, while the generation of theoretical semantic parameters 7, semantic analysis parameters 8 and parameter adjustment results 9 ensures the accuracy and effectiveness of semantic expansion. The contents not described in detail in the specification belong to the prior art known to those skilled in the art, and the model parameters of each electrical appliance are not specifically limited, and conventional equipment can be used. In this technical solution, the electrical control components not mentioned are not shown in the figure because they belong to the prior art, and they are not described here.
[0036] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for constructing a knowledge graph for semantic understanding and expansion of English vocabulary, characterized by: include: Acquire the words and their contextual information in the target vocabulary set, extract the semantic features and contextual features of each word, and construct a semantic model and a contextual model of the word based on the semantic features and contextual features, respectively; fuse the semantic model and the contextual model to generate a comprehensive semantic model of the word, and establish a semantic association network between the words based on the similarity calculation results between the comprehensive semantic models; wherein the semantic model is used to describe the core semantic attributes of the word, the contextual model is used to describe the semantic change rules of the word in different contexts, and the comprehensive semantic model is used to comprehensively express the semantic characteristics of the word; Continuously collect the frequency and semantic distribution of the words in the target vocabulary set in different corpora, and combine and sort the collected frequency and semantic distribution of the words in different corpora in chronological order to obtain the historical semantic sequence of each word; Obtain the current semantic extension requirements and distribute them through the semantic association network to obtain the theoretical semantic parameters of each comprehensive semantic model corresponding to the current semantic extension requirements; Substituting each theoretical semantic parameter into the corresponding comprehensive semantic model, and making each comprehensive semantic model perform multi-dimensional semantic analysis on each theoretical semantic parameter according to the corresponding historical semantic sequence, so as to obtain the semantic analysis parameters of each comprehensive semantic model; The semantic analysis parameters of each comprehensive semantic model are substituted into the semantic association network, and the semantic association network is made to perform parameter adjustment processing on each theoretical semantic parameter to obtain the parameter adjustment result, and the words in the target vocabulary set are semantically expanded according to the parameter adjustment result.
2. The method for constructing a knowledge graph for English vocabulary semantic understanding and expansion according to claim 1, characterized in that: The steps of obtaining current semantic extension requirements and distributing the current semantic extension requirements through a semantic association network to obtain theoretical semantic parameters of each comprehensive semantic model corresponding to the current semantic extension requirements include: Data collection is performed on a corpus related to the target vocabulary set to obtain corpus data in the corpus, and the corpus data is analyzed and processed for semantic topics and semantic extension conditions of the semantic topics to obtain semantic extension conditions of the semantic topics in the corpus, and the semantic extension conditions are used as current semantic extension requirements; The current semantic extension demand is split through each comprehensive semantic model in the semantic association network to obtain the split semantic demand of each comprehensive semantic model corresponding to the current semantic extension demand; According to each comprehensive semantic model, the corresponding split semantic requirements are reversely deduced to obtain the theoretical semantic parameters required by each comprehensive semantic model to execute the split semantic requirements.
3. The method for constructing a knowledge graph for understanding and expanding English vocabulary semantics according to claim 1, characterized in that: Substituting each theoretical semantic parameter into a corresponding comprehensive semantic model, and allowing each comprehensive semantic model to perform multi-dimensional semantic analysis on each theoretical semantic parameter according to the corresponding historical semantic sequence to obtain the semantic analysis parameters of each comprehensive semantic model includes the following steps: Substituting each theoretical semantic parameter into the corresponding comprehensive semantic model, the comprehensive semantic model is made to search for the same parameter in the corresponding historical semantic sequence according to the theoretical semantic parameter, so as to obtain the semantic parameter in the historical semantic sequence that is the same as the theoretical semantic parameter, and marking the semantic parameter and the corresponding semantic distribution as the first reference data; Integrating the consecutive first reference data into a first reference segment according to the time corresponding to each first reference data to obtain a plurality of first reference segments, and performing multi-dimensional analysis on each first reference segment to obtain analysis parameters of the first reference data; Taking the theoretical semantic parameter as a benchmark, obtaining several neighboring semantic parameters adjacent to the theoretical semantic parameter, and making the comprehensive semantic model search for identical parameters in the corresponding historical semantic sequence based on the neighboring semantic parameters to obtain semantic parameters identical to the neighboring semantic parameters in the historical semantic sequence, and marking the semantic parameters and the corresponding semantic distribution as second reference data; Integrating the consecutive second reference data into a second reference segment according to the time corresponding to each second reference data to obtain a plurality of second reference segments, and performing multi-dimensional analysis on each second reference segment to obtain analysis parameters of the second reference data; The analysis parameters of the first reference data and the analysis parameters of the second reference data are combined to form semantic analysis parameters of the comprehensive semantic model.
4. A method for constructing a knowledge graph for understanding and expanding English vocabulary semantics according to claim 3, characterized in that: The step of performing multi-dimensional analysis on each first reference paragraph to obtain analysis parameters of the first reference data includes: Calculating the mean of the first reference data of each first reference paragraph in the same time period, and performing a difference calculation between the mean calculation result and the semantic distribution of the corresponding theoretical semantic parameter under an ideal state to obtain the degree of deviation between the semantic distribution of the theoretical semantic parameter generated over the time period and the semantic distribution under the ideal state, combining each degree of deviation with the time period of the corresponding first reference data, and arranging each combination in chronological order to obtain a first stable feature sequence of the semantic parameter; performing a comparative analysis on the first reference data in the same time period for each first reference paragraph to obtain a degree of deviation of the first reference data in the same time period for each first reference paragraph, combining each degree of deviation with the corresponding time period of the first reference data, and arranging each combination in chronological order to obtain a second stable feature sequence of the semantic parameter; The first stable feature sequence and the second stable feature sequence are superimposed to summarize the deviation degrees corresponding to the same time period in the first stable feature sequence and the second stable feature sequence into the same combination as the first analysis parameter of the first reference data.
5. The method for constructing a knowledge graph for English vocabulary semantic understanding and expansion according to claim 3, characterized in that: The step of performing multi-dimensional analysis on each second reference paragraph to obtain analysis parameters of the second reference data includes: Calculating the mean of the first reference data of each second reference paragraph in the same time period, and performing a difference calculation between the mean calculation result and the semantic distribution of the corresponding theoretical semantic parameter under an ideal state to obtain the degree of deviation between the semantic distribution of the theoretical semantic parameter generated over the time period and the semantic distribution under the ideal state, combining each degree of deviation with the time period of the corresponding first reference data, and arranging each combination in chronological order to obtain a third stable feature sequence of the semantic parameter; performing a comparative analysis on the first reference data in the same time period for each first reference paragraph to obtain a degree of deviation of the first reference data in the same time period for each first reference paragraph, combining each degree of deviation with the time period of the corresponding first reference data, and arranging each combination in chronological order to obtain a fourth stable feature sequence of the semantic parameter; The third stable feature sequence and the fourth stable feature sequence are superimposed to classify the deviation degrees corresponding to the same time period in the third stable feature sequence and the fourth stable feature sequence into the same combination as the second analysis parameter of the first reference data.
6. The method for constructing a knowledge graph for understanding and expanding English vocabulary semantics according to claim 1, characterized in that: Substituting the semantic analysis parameters of each comprehensive semantic model into the semantic association network, and allowing the semantic association network to perform parameter adjustment processing on each theoretical semantic parameter includes the following steps: Substituting the first analysis parameter of the first reference data of each comprehensive semantic model into the semantic association network, and causing the semantic association network to perform an overall semantic operation according to the first analysis parameter of the first reference data of each comprehensive semantic model, so as to obtain a baseline semantic graph of the semantic association network; Taking the benchmark semantic graph of the semantic association network as a benchmark, the second analysis parameters of each second reference data of each comprehensive semantic model are respectively substituted into the semantic association network, and the semantic association network is subjected to an overall semantic operation to obtain each neighboring semantic graph of the semantic association network; A comparative analysis is performed on the baseline semantic map and each neighboring semantic map to obtain the comparative analysis results. Based on the comparative analysis results, the baseline semantic map or neighboring semantic map with the best effect is determined, and the semantic parameters corresponding to the baseline semantic map or the neighboring semantic map are adjusted to the theoretical semantic parameters to obtain the parameter adjustment results.
7. A knowledge graph construction system for English vocabulary semantic understanding and expansion, characterized by: include: The model building module is used to obtain the vocabulary and its contextual information in the target vocabulary set, extract the semantic features and contextual features of each vocabulary, and respectively build a semantic model and a contextual model of the vocabulary based on the semantic features and contextual features. The semantic model and the contextual model are integrated to generate a comprehensive semantic model of the vocabulary, and a semantic association network between the vocabulary is established based on the similarity calculation results between the comprehensive semantic models. Among them, the semantic model is used to describe the core semantic attributes of the vocabulary, the context model is used to describe the semantic change rules of the vocabulary in different contexts, and the comprehensive semantic model is used to comprehensively express the semantic characteristics of the vocabulary. The data collection module is used to continuously collect the occurrence frequency and semantic distribution of the words in the target vocabulary set in different corpora, and to combine and sort the collected occurrence frequency and semantic distribution of the words in different corpora in chronological order to obtain the historical semantic sequence of each word; A preliminary allocation module is used to obtain the current semantic extension requirements and allocate the current semantic extension requirements through the semantic association network to obtain the theoretical semantic parameters of each comprehensive semantic model corresponding to the current semantic extension requirements; A semantic analysis module is used to substitute each theoretical semantic parameter into the corresponding comprehensive semantic model, and to make each comprehensive semantic model perform multi-dimensional semantic analysis on each theoretical semantic parameter according to the corresponding historical semantic sequence to obtain the semantic analysis parameters of each comprehensive semantic model; The parameter adjustment module is used to substitute the semantic analysis parameters of each comprehensive semantic model into the semantic association network, so that the semantic association network performs parameter adjustment processing on each theoretical semantic parameter to obtain the parameter adjustment result, and semantically expand the vocabulary in the target vocabulary set according to the parameter adjustment result.
8. A knowledge graph construction system for English vocabulary semantic understanding and expansion according to claim 7, characterized in that: The preliminary allocation module is also used to: Data collection is performed on a corpus related to the target vocabulary set to obtain corpus data in the corpus, and the corpus data is analyzed and processed for semantic topics and semantic extension conditions of the semantic topics to obtain semantic extension conditions of the semantic topics in the corpus, and the semantic extension conditions are used as current semantic extension requirements; The current semantic extension demand is split through each comprehensive semantic model in the semantic association network to obtain the split semantic demand of each comprehensive semantic model corresponding to the current semantic extension demand; According to each comprehensive semantic model, the corresponding split semantic requirements are reversely deduced to obtain the theoretical semantic parameters required by each comprehensive semantic model to execute the split semantic requirements.
9. The knowledge graph construction system for English vocabulary semantic understanding and expansion according to claim 7, characterized in that: The semantic analysis module is also used to: Substituting each theoretical semantic parameter into the corresponding comprehensive semantic model, the comprehensive semantic model is made to search for the same parameter in the corresponding historical semantic sequence according to the theoretical semantic parameter, so as to obtain the semantic parameter in the historical semantic sequence that is the same as the theoretical semantic parameter, and marking the semantic parameter and the corresponding semantic distribution as the first reference data; Integrating the consecutive first reference data into a first reference segment according to the time corresponding to each first reference data to obtain a plurality of first reference segments, and performing multi-dimensional analysis on each first reference segment to obtain analysis parameters of the first reference data; Taking the theoretical semantic parameter as a benchmark, obtaining several neighboring semantic parameters adjacent to the theoretical semantic parameter, and making the comprehensive semantic model search for identical parameters in the corresponding historical semantic sequence based on the neighboring semantic parameters to obtain semantic parameters identical to the neighboring semantic parameters in the historical semantic sequence, and marking the semantic parameters and the corresponding semantic distribution as second reference data; Integrating the consecutive second reference data into a second reference segment according to the time corresponding to each second reference data to obtain a plurality of second reference segments, and performing multi-dimensional analysis on each second reference segment to obtain analysis parameters of the second reference data; The analysis parameters of the first reference data and the analysis parameters of the second reference data are combined to form semantic analysis parameters of the comprehensive semantic model.
10. The knowledge graph construction system for English vocabulary semantic understanding and expansion according to claim 7, characterized in that: The parameter adjustment module is also used to: Substituting the first analysis parameter of the first reference data of each comprehensive semantic model into the semantic association network, and causing the semantic association network to perform an overall semantic operation according to the first analysis parameter of the first reference data of each comprehensive semantic model, so as to obtain a baseline semantic graph of the semantic association network; Taking the benchmark semantic graph of the semantic association network as a benchmark, the second analysis parameters of each second reference data of each comprehensive semantic model are respectively substituted into the semantic association network, and the semantic association network is subjected to an overall semantic operation to obtain each neighboring semantic graph of the semantic association network; A comparative analysis is performed on the baseline semantic map and each neighboring semantic map to obtain the comparative analysis results. Based on the comparative analysis results, the baseline semantic map or neighboring semantic map with the best effect is determined, and the semantic parameters corresponding to the baseline semantic map or the neighboring semantic map are adjusted to the theoretical semantic parameters to obtain the parameter adjustment results.