AI-based Method for Mining and Expanding Semantic Coherence of English Corpus

The AI-based method for English corpora semantic coherence mining and expansion addresses the challenge of deep semantic understanding in NLP by improving precision and coherence in language data processing.

CN119990141BActive Publication Date: 2025-07-15BEIJING UNION UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510222646.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-15
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The existing technology cannot deeply understand the deep semantic relationships and context information in text, resulting in insufficient accuracy and consistency of corpus semantic analysis.

Method used

Through AI-based methods, classify text data of the English corpus, determine text category data, perform word segmentation processing and semantic analysis, build semantic maps, conduct coherent analysis and generate extension evaluation reports to ensure the accuracy and depth of semantic understanding.

Benefits of technology

It improves the accuracy and depth of semantic understanding, enhances the accuracy and flexibility of semantic expansion, ensures the semantic coherence between the expanded content and the English corpus, and improves the intelligence level of natural language processing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990141B_ABST
    Figure CN119990141B_ABST
Patent Text Reader

Abstract

The present invention provides a method for semantic coherence mining and expansion of an English corpus based on AI, belonging to the technical field of data processing, including: Step 1: Classify all text data in the English corpus to determine text category data, and determine text tokenization data based on the text category data; Step 2: Perform semantic analysis on the text tokenization data to determine the semantic graph of the English corpus; Step 3: Perform coherence analysis on the semantic graph to determine semantic coherence data, and determine semantic expansion data based on the text tokenization data and the semantic coherence data; Step 4: Evaluate the semantic expansion data and generate an expansion evaluation report. It can improve the accuracy and depth of semantic understanding, enhance the accuracy and flexibility of semantic expansion, ensure that the generated expanded content is consistent with the semantic coherence of the English corpus, achieve a systematic enhancement of the semantic depth in the English corpus, and improve the intelligence of natural language processing tasks in the practical application of the English corpus.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a method for mining and expanding semantic coherence of an English corpus based on AI. Background Art

[0002] With the continuous development of natural language processing (NLP) technology, the construction of corpora and semantic analysis have become one of the core tasks in the field of artificial intelligence. Early text processing technologies mainly relied on simple statistical models and rule-based methods, such as word frequency-based lexical analysis and word segmentation processing. However, these methods usually cannot deeply understand the deep semantic relationships and context information in the text.

[0003] With the introduction of deep learning and machine learning technologies, especially in the late 2010s, neural network-based language models (such as Word2Vec, BERT, etc.) have shown significant advantages in processing language data. These models can capture more complex semantic relationships between words in the text, promoting the development of corpus semantic analysis. Semantic coherence analysis and semantic expansion technologies have begun to become key technologies for enhancing the depth and consistency of corpus semantic understanding.

[0004] Therefore, the present invention provides a method for mining and expanding semantic coherence of an English corpus based on AI. Summary of the Invention

[0005] The present invention provides a method for mining and expanding semantic coherence of an English corpus based on AI. By determining text tokenization data based on determined text category data, analyzing text analysis data to determine a semantic graph, performing coherence analysis on the semantic graph to determine semantic coherence data, determining semantic expansion data based on the text tokenization data and the semantic coherence data, evaluating the semantic expansion data, and generating an expansion evaluation report, the accuracy and depth of semantic understanding can be improved, the accuracy and flexibility of semantic expansion can be enhanced, ensuring that the generated expanded content is consistent with the semantic coherence of the English corpus, realizing a systematic enhancement of the semantic depth in the English corpus, and enhancing the intelligence of natural language processing tasks in the actual application of the English corpus.

[0006] The present invention provides a method for mining and expanding semantic coherence of an English corpus based on AI, including:

[0007] Step 1: Classify all text data of the English corpus to determine text category data, and determine text tokenization data based on the text category data;

[0008] Step 2: Perform semantic analysis on the text tokenization data to determine the semantic graph of the English corpus;

[0009] Step 3: Conduct a coherence analysis on the semantic graph to determine semantic coherence data, and determine semantic expansion data based on the text segmentation data and the semantic coherence data;

[0010] Step 4: Evaluate the semantic expansion data and generate an expansion evaluation report.

[0011] According to the AI-based method for mining and expanding semantic coherence of an English corpus provided by the present invention, classify all text data of the English corpus to determine text category data, including:

[0012] Determine the text characteristics of each text data in the English corpus, where the text characteristics include text theme, text structure, language style, and target audience;

[0013] Based on the text characteristics of all text data in the English corpus, determine text classification rules, and match each text data in the English corpus with the text classification rules to determine the category label of each text data in the English corpus;

[0014] Based on all text data with the same category label, determine the text category sub-data of each category label;

[0015] Based on the text category sub-data of all category labels, determine the text category data of the English corpus.

[0016] According to the AI-based method for mining and expanding semantic coherence of an English corpus provided by the present invention, determine text segmentation data based on the text category data, including:

[0017] Perform word segmentation processing and stop word removal on each text data in each text category sub-data in the text category data to determine the word segmentation set of each text data in each text category sub-data in the text category data, where the word segmentation set includes multiple word segments in the corresponding text data;

[0018] Based on the word segmentation set of each text data in each text category sub-data in the text category data, determine the word segmentation data of each text data;

[0019] Based on the word segmentation sets of all text data in each text category sub-data in the text category data, determine the sub-category word segmentation set of each text category sub-data;

[0020] Based on the sub-category word segmentation sets of all text category sub-data, determine the text segmentation data of the English corpus.

[0021] According to the AI-based method for mining and expanding semantic coherence of an English corpus provided by the present invention, perform semantic analysis on the text segmentation data to determine the semantic graph of the English corpus, including:

[0022] Based on all the text data in each text category sub - data of the text category data, perform part - of - speech analysis and sentence analysis on each word segment in the sub - category word segment set of each text category sub - data in the text word - segmented data to determine the word vectors of each word segment in the sub - category word segment set of each text category sub - data;

[0023] Input the sub - category word segment set of each text category sub - data in the text word - segmented data and the word vectors of each word segment in the sub - category word segment set into the word - segment semantic recognition model, and determine the entity word segment set and sub - category word segment relationship of each text category sub - data based on the output structure of the word - segment semantic recognition model;

[0024] Construct the sub - semantic graph of each text category sub - data based on the sub - category word segment set, entity word segment set, and sub - category word segment relationship of each text category sub - data;

[0025] Determine the semantic graph of the English corpus based on the sub - semantic graphs of all text category sub - data.

[0026] According to the method for mining and expanding semantic coherence of an English corpus based on AI provided by the present invention, perform coherence analysis on the semantic graph to determine semantic coherence data, and determine semantic expansion data based on the text word - segmented data and semantic coherence data, including:

[0027] Based on the text category data, text word - segmented data, and semantic graph, calculate the word - level semantic coherence value and text - level semantic coherence value of each word segment;

[0028] Calculate the semantic coherence value of each word segment based on the word - level semantic coherence value and text - level semantic coherence value of each word segment;

[0029] Based on the semantic coherence values of all word segments in the text word - segmented data, determine the semantic coherence data of the English corpus;

[0030] Input the semantic coherence data and the text word - segmented data into the semantic coherence model, and determine the semantic expansion sub - data of each analysis in the text word - segmented data based on the output result of the semantic coherence model;

[0031] Based on all the semantic expansion sub - data of the analysis in the text word - segmented data, determine the semantic expansion data of the English corpus.

[0032] According to the method for mining and expanding semantic coherence of an English corpus based on AI provided by the present invention, based on the text category data, text word - segmented data, and semantic graph, calculate the word - level semantic coherence value and text - level semantic coherence value of each word segment, including:

[0033] ;

[0034] ;

[0035] ; ;

[0036] wherein, represents the semantic coherence value of the $i$-th word segmentation in the text segmentation data in the $a$-th text category sub-data of the text category data, represents the coherence value of the $i$-th word segmentation in the text segmentation data in the $b$-th text data of the $a$-th text category sub-data of the text category data, represents the word vector of the $i$-th word segmentation in the sub-category word segmentation set of all text category sub-data in the text segmentation data, represents the word vector of the $k$-th word segmentation connected to the $i$-th word segmentation in the sub-semantic graph of the $a$-th text category sub-data in the semantic graph, The word vector of the $j$-th word segmentation not connected to the $i$-th word segmentation in the sub-semantic graph of the $a$-th text category sub-data in the semantic graph, represents the number of word segmentations connected to the $i$-th word segmentation in the sub-semantic graph of the $a$-th text category sub-data in the semantic graph, represents the number of word segmentations in the sub-semantic graph of the $a$-th text category sub-data in the semantic graph, represents the first adjustment parameter, 2 represents the second adjustment parameter, represents the first indicator function of the $i$-th word segmentation in the text segmentation data for the $a$-th text category sub-data of the text category data, represents the sub-category word segmentation set of the $a$-th text category sub-data in the text segmentation data, represents the word segmentation position of the $i$-th word segmentation in the text segmentation data in the $b$-th text data of the $a$-th text category sub-data of the text category data, represents the number of word segmentations in the word segmentation window of the $i$-th word segmentation in the text segmentation data in the $b$-th text data of the $a$-th text category sub-data of the text category data, represents the word vectors of the $(i - c)$-th word segmentation and the $(i + c)$-th word segmentation of the $i$-th word segmentation in the text segmentation data in the $b$-th text data of the $a$-th text category sub-data of the text category data, represents the second indicator function of the $i$-th word segmentation in the text segmentation data for the $b$-th text data of the $a$-th text category sub-data of the text category data, represents the word segmentation window decay parameter, represents the word segmentation data of the $b$-th text data of the $a$-th text category sub-data of the text category data.

[0037] The method for mining and expanding semantic coherence of an English corpus based on AI provided by the present invention calculates the semantic coherence value of each word segment based on the word semantic coherence value and the text semantic coherence value of each word segment, including:

[0038] ;

[0039] wherein, represents the semantic coherence value of the i-th word segment in the sub-category word segment set of all text category sub-data in the text word segment data, N3 represents the number of text category sub-data in the text category data, the number of text data in the a-th text category sub-data in the text category data.

[0040] The method for mining and expanding semantic coherence of an English corpus based on AI provided by the present invention evaluates the semantic expansion data and generates an expansion evaluation report, including:

[0041] Based on the text word segment data and the semantic expansion data of each piece of text data in each text category sub-data in the text category data, determine the text expansion data of each piece of text data in each text category sub-data in the text category data;

[0042] Evaluate the consistency between each piece of text data and the corresponding text expansion data in each text category sub-data in the text category data;

[0043] Generate an expansion evaluation report based on the consistency of all text data in all text category sub-data in the text category data.

[0044] Compared with the prior art, the beneficial effects of the present application are as follows:

[0045] By determining the text word segment data based on the determined text category data, analyzing the text analysis data to determine the semantic graph, performing coherence analysis on the semantic graph to determine the semantic coherence data, determining the semantic expansion data according to the text word segment data and the semantic coherence data, evaluating the semantic expansion data, and generating an expansion evaluation report, the accuracy and depth of semantic understanding can be improved, the accuracy and flexibility of semantic expansion can be enhanced, ensuring that the generated expanded content is consistent with the semantic coherence of the English corpus, realizing a systematic enhancement of the semantic depth in the English corpus, and improving the intelligence of natural language processing tasks in the actual application of the English corpus. Description of the Drawings

[0046] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0047] Figure 1 It is a schematic flowchart of the method for semantic coherence mining and expansion of an English corpus based on AI provided by an embodiment of the present invention. Detailed implementation manners

[0048] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0049] Embodiment 1:

[0050] The embodiment of the present invention provides a method for semantic coherence mining and expansion of an English corpus based on AI, as Figure 1 shown, including:

[0051] Step 1: Classify all text data in the English corpus to determine text category data, and determine text segmentation data based on the text category data;

[0052] Step 2: Perform semantic analysis on the text segmentation data to determine the semantic graph of the English corpus;

[0053] Step 3: Perform coherence analysis on the semantic graph to determine semantic coherence data, and determine semantic expansion data based on the text segmentation data and the semantic coherence data;

[0054] Step 4: Evaluate the semantic expansion data and generate an expansion evaluation report.

[0055] In this embodiment, all text data in the English corpus is analyzed to identify its text characteristics (such as theme, text structure, language style, audience, etc.), and the text is classified based on these features. After classification, text category data is generated, and each text category sub-data represents a specific text type.

[0056] In this embodiment, according to the text category data, text segmentation processing is performed to break each text into several words or phrases. The segmented data provides a basis for subsequent semantic analysis.

[0057] In this embodiment, semantic analysis aims to deeply understand the semantic relationships between words in the text, and determine the semantic graph through language processing and the word segmentation semantic recognition model.

[0058] In this embodiment, analyze the coherence of word segmentation in the semantic graph, identify the semantic consistency between word segmentation and text data, and generate semantically coherent data.

[0059] In this embodiment, combine the text word segmentation data with the semantically coherent data, explore the potential of semantic expansion, and generate new extended data.

[0060] Beneficial effects of the above technical solution: Determine the text word segmentation data through the determined text category data, analyze the text analysis data to determine the semantic graph, perform coherence analysis on the semantic graph to determine the semantically coherent data, determine the semantic expansion data according to the text word segmentation data and the semantically coherent data, evaluate the semantic expansion data, and generate an extended evaluation report, which can improve the accuracy and depth of semantic understanding, enhance the accuracy and flexibility of semantic expansion, ensure that the generated extended content is consistent with the semantic coherence of the English corpus, achieve a systematic enhancement of the semantic depth in the English corpus, and improve the intelligence of natural language processing tasks in the practical application of the English corpus.

[0061] Embodiment 2:

[0062] The embodiment of the present invention provides an AI-based method for mining and expanding semantic coherence of an English corpus, classifying all text data of the English corpus to determine text category data, including:

[0063] Determine the text characteristics of each text data in the English corpus, where the text characteristics include text theme, text structure, language style, and target audience;

[0064] Based on the text characteristics of all text data in the English corpus, determine the text classification rules, match each text data in the English corpus with the text classification rules, and determine the category label of each text data in the English corpus;

[0065] Based on all text data with the same category label, determine the text category sub-data of each category label;

[0066] Based on the text category sub-data of all category labels, determine the text category data of the English corpus.

[0067] In this embodiment, each text in the corpus is analyzed to extract text features, which include: text theme: the main content or topic discussed in the text (e.g., technology, education, politics, etc.); text structure: the organization of the text, such as paragraph structure, sentence construction, etc., analyzing the typesetting, paragraph division, etc. of the text; language style: the language features used in the text, such as formal or informal language, academic or popular language, etc.; target audience: the group of readers the text is targeted at, such as academic people, general readers or experts in a specific industry.

[0068] In this embodiment, according to the text features obtained from the analysis, text classification rules are defined, that is, classification criteria are set for text data according to factors such as the theme, structure, style and audience of the text.

[0069] In this embodiment, according to the set classification rules, each text is matched with the rules to determine its category label, and the category label reflects the core content and features of the text.

[0070] In this embodiment, for each category label, all texts that match the label are collected to form a text category sub-data, that is, a set of texts containing the same category label.

[0071] In this embodiment, the text category sub-data under all category labels are integrated to form complete text category data, which is a corpus data sorted and classified by category.

[0072] The beneficial effects of the above technical solution: Classifying all text data in the English corpus to determine text category data can achieve efficient automatic classification, improve the accuracy and depth of semantic understanding, and provide a data basis for determining text analysis data.

[0073] Embodiment 3:

[0074] The embodiment of the present invention provides an AI-based method for mining and expanding semantic coherence of an English corpus. Based on the text category data, text tokenization data is determined, including:

[0075] Perform word segmentation processing and stop word removal on each text data in each text category sub-data in the text category data to determine the word segmentation set of each text data in each text category sub-data in the text category data. Among them, the word segmentation set of each text includes multiple word segments in the corresponding text data;

[0076] Based on the word segmentation sets of each text data in each text category sub-data in the text category data, determine the word segmentation data of each text data.

[0077] Based on the word segmentation sets of all text data in each text category sub-data in the text category data, determine the sub-category word segmentation set of each text category sub-data.

[0078] Determine the text tokenization data of the English corpus based on the sub-category tokenization sets of all text category sub-data.

[0079] In this embodiment, each text is tokenized and split into independent tokens. Then, stop words (such as common but meaningless words like "the", "and", "is", etc.) are removed, and only meaningful words are retained to obtain the tokenization set of each text, that is, the set of all meaningful tokens in the text.

[0080] In this embodiment, the text data is compared according to the tokenization set of each text to determine the tokenization data of the text after removing stop words and only retaining meaningful tokens.

[0081] In this embodiment, for all text data in each text category sub-data, the tokenization sets of each text are collected, and the common features of the texts in this category are analyzed to extract the sub-category tokenization set representing this category, that is, the tokens with higher frequencies in the texts of this category.

[0082] In this embodiment, the sub-category tokenizations in all category sub-data are integrated to finally form a global text tokenization data, which represents the tokenization characteristics of the entire English corpus and is convenient for further applications such as semantic analysis and information extraction.

[0083] Beneficial effects of the above technical solution: Determining text tokenization data based on text category data can refine the lexical analysis of the text, improve the accuracy of classification, enhance the ability to capture text semantics, and provide a more accurate data basis for determining the semantic map.

[0084] Embodiment 4:

[0085] The embodiment of the present invention provides a method for mining and expanding semantic coherence of an English corpus based on AI, which performs semantic analysis on text tokenization data to determine the semantic map of the English corpus, including:

[0086] Based on all text data in each text category sub-data in the text category data, perform part-of-speech analysis and sentence analysis on each token in the sub-category tokenization set of each text category sub-data in the text tokenization data to determine the word vector of each token in the sub-category tokenization set of each text category sub-data;

[0087] Input the sub-category tokenization set of each text category sub-data in the text tokenization data and the word vector of each token in the sub-category tokenization set into the token semantic recognition model, and determine the entity tokenization set and sub-category token relationship of each text category sub-data based on the output structure of the token semantic recognition model;

[0088] Construct a sub-semantic graph for each text category sub-data based on the sub-category word segmentation set, entity word segmentation set, and sub-category word segmentation relationship of each text category sub-data;

[0089] Determine the semantic graph of the English corpus based on the sub-semantic graphs of all text category sub-data.

[0090] In this embodiment, perform part-of-speech analysis on the word segments in each sub-category word segmentation, that is, identify the part of speech of each word segment (such as noun, verb, adjective, etc.). At the same time, perform sentence analysis to examine the structure and grammatical relationship of the vocabulary in the sentence. Based on these analyses, generate word vectors to capture the semantic information of the vocabulary.

[0091] In this embodiment, input the sub-category word segmentation of each text category sub-data and the word vectors of each word segment into a word segmentation semantic recognition model. The model identifies the role of each word segment based on context and semantic understanding, and outputs entity word segments (such as person names, place names, time, etc.) and the relationships between the vocabulary (such as "belong to", "contain", etc.).

[0092] In this embodiment, the word segmentation semantic recognition model extracts semantic information from all text data in each text category sub-data, and identifies entities and establishes mappings between entities, relationships, and attributes.

[0093] In this embodiment, construct a sub-semantic graph according to the sub-category word segmentation, entity word segmentation, and word segmentation relationship of each text category sub-data. This sub-semantic graph shows the relationships between different vocabulary and entities in the corresponding text category sub-data, reflecting the semantic network structure inside the text.

[0094] In this embodiment, merge the sub-semantic graphs of all text category sub-data to form a global semantic graph, which shows the vocabulary, entities, and their mutual relationships in the entire English corpus.

[0095] The beneficial effects of the above technical solution: Perform semantic analysis on the text word segmentation data, determine the semantic graph of the English corpus, can deeply explore the complex semantic relationships between vocabulary and entities in the text, improve the accuracy of text understanding and semantic reasoning, and enhance the comprehensive understanding of the semantic structure of the English corpus.

[0096] Embodiment 5:

[0097] The embodiment of the present invention provides a method for mining and expanding semantic coherence of an English corpus based on AI. Perform coherence analysis on the semantic graph to determine semantic coherence data, and determine semantic expansion data based on the text word segmentation data and semantic coherence data, including:

[0098] Based on the text category data, text word segmentation data, and semantic graph, calculate the word semantic coherence value and text semantic coherence value of each word segment;

[0099] Calculate the semantic coherence value of each word segment based on the word semantic coherence value and the text semantic coherence value of each word segment;

[0100] Determine the semantic coherence data of the English corpus based on the semantic coherence values of all word segments in the text word segmentation data;

[0101] Input the semantic coherence data and the text word segmentation data into the semantic coherence model, and based on the output result of the semantic coherence model, determine the semantic expansion sub-data of each analysis in the text word segmentation data;

[0102] Determine the semantic expansion data of the English corpus based on the semantic expansion sub-data of all analyses in the text word segmentation data.

[0103] In this embodiment, by combining the word semantic coherence value and the text semantic coherence value of each word segment, the final semantic coherence value of the word segment is calculated, which represents the semantic coherence and consistency of the word in the text.

[0104] In this embodiment, according to the semantic coherence values of the word segments in all texts, the semantic coherence data of the entire corpus is constructed, which reflects the semantic consistency and coherence of the entire English corpus.

[0105] In this embodiment, the determined semantic coherence data and the text word segmentation data are input into the semantic coherence model. The model can analyze which word segments or concepts in the text are expandable and output semantic expansion sub-data, that is, in the semantic coherence framework, the model identifies the possible expanded semantic information.

[0106] In this embodiment, according to the semantic expansion sub-data of each analysis output by the model, summarize and integrate, and finally determine the semantic expansion data of the entire corpus, which is the set of expanded semantic information in the corpus.

[0107] Beneficial effects of the above technical solution: Conduct coherent analysis on the semantic graph to determine the semantic coherence data, which can quantify the semantic coherence, improve the understanding of the internal semantic structure of the text, perform effective semantic expansion within the global scope of the English corpus, and ensure that the generated expanded content is consistent with the semantic coherence of the English corpus.

[0108] Embodiment 6:

[0109] The embodiment of the present invention provides a method for mining and expanding the semantic coherence of an English corpus based on AI. Based on the text category data, the text word segmentation data, and the semantic graph, calculate the word semantic coherence value and the text semantic coherence value of each word segment, including:

[0110] ;

[0111] ;

[0112] ; ;

[0113] wherein, represents the semantic coherence value of the i-th word segmentation in the text segmentation data in the a-th text category sub-data of the text category data, represents the semantic coherence value of the i-th word segmentation in the text segmentation data in the b-th text data of the a-th text category sub-data of the text category data, represents the word vector of the i-th word segmentation in the sub-category word segmentation set of all text category sub-data in the text segmentation data, represents the word vector of the k-th word segmentation connected to the i-th word segmentation in the sub-semantic graph of the a-th text category sub-data in the semantic graph, The word vector of the j-th word segmentation not connected to the i-th word segmentation in the sub-semantic graph of the a-th text category sub-data in the semantic graph, represents the number of word segmentations connected to the i-th word segmentation in the sub-semantic graph of the a-th text category sub-data in the semantic graph, represents the number of word segmentations in the sub-semantic graph of the a-th text category sub-data in the semantic graph, represents the first adjustment parameter, 2 represents the second adjustment parameter, represents the first indication function of the i-th word segmentation in the text segmentation data for the a-th text category sub-data of the text category data, represents the sub-category word segmentation set of the a-th text category sub-data in the text segmentation data, represents the word segmentation position of the i-th word segmentation in the text segmentation data in the b-th text data of the a-th text category sub-data of the text category data in the text segmentation data of the b-th text, represents the number of word segmentations in the word segmentation window of the word segmentation data of the b-th text data of the a-th text category sub-data of the text category data where the i-th word segmentation is located in the text segmentation data, represents the word vectors of the (i - c)-th word segmentation and the (i + c)-th word segmentation of the i-th word segmentation in the word segmentation data of the b-th text data of the a-th text category sub-data of the text category data in the text segmentation data, represents the second indication function of the i-th word segmentation in the text segmentation data for the b-th text data of the a-th text category sub-data of the text category data, represents the word segmentation window decay parameter, represents the word segmentation data of the b-th text data of the a-th text category sub-data of the text category data.

[0114] In this embodiment, the semantic coherence value of the word Indicates the coherence of all the word segmentations of the $i$-th word segmentation in the text segmentation data within the sub-category word segmentation set of the $a$-th text category sub-data in the text category data.

[0115] In this embodiment, the semantic coherence value of an article Indicates the coherence of all the analyses within the word segmentation window in the article word segmentation data of the $b$-th article data in the $a$-th text category sub-data of the text category data for the $i$-th word segmentation in the text segmentation data.

[0116] In this embodiment, the word segmentation window is determined according to the text length of the corresponding word segmentation in the article word segmentation data of the text data. The longer the word segmentation length of the article word segmentation data, the larger the word segmentation window; the shorter the word segmentation length of the article word segmentation data, the smaller the word segmentation window.

[0117] Advantages of the above technical solution: Based on the text category data, text segmentation data, and semantic graph, calculating the word semantic coherence value and article semantic coherence value for each word segmentation can provide a data basis for quantifying the semantic coherence of each word segmentation, enhance the understanding of the internal semantic structure of the text, and provide a data basis for effective semantic expansion within the global scope of the English corpus.

[0118] Embodiment 7:

[0119] The embodiment of the present invention provides an AI-based method for mining and expanding semantic coherence in an English corpus. Calculating the semantic coherence value for each word segmentation based on the word semantic coherence value and article semantic coherence value of each word segmentation includes:

[0120] ;

[0121] wherein, represents the semantic coherence value of the $i$-th word segmentation in the sub-category word segmentation sets of all text category sub-data in the text segmentation data, and $N3$ represents the number of text category sub-data in the text category data, the number of text data in the $a$-th text category sub-data in the text category data.

[0122] In this embodiment, the semantic coherence value represents the integration of the word semantic coherence values of all text category sub-data in the text category data and the article semantic coherence values of all text data of all text category sub-data in the text category data.

[0123] Advantages of the above technical solution: Calculating the semantic coherence value for each word segmentation based on the word semantic coherence value and article semantic coherence value of each word segmentation can provide a data basis for determining semantic coherence data, enhance the understanding of the internal semantic structure of the text, and provide a data basis for effective semantic expansion within the global scope of the English corpus.

[0124] Embodiment 8:

[0125] The embodiment of the present invention provides a method for semantic coherence mining and expansion of an English corpus based on AI, which evaluates semantic expansion data and generates an expansion evaluation report, including:

[0126] Determine the text expansion data of each text data in each text category sub-data in the text category data based on the word segmentation data and semantic expansion data of each text data in each text category sub-data in the text category data;

[0127] Evaluate the consistency between each text data and its corresponding text expansion data in each text category sub-data in the text category data;

[0128] Generate an expansion evaluation report based on the consistency of all text data in all text category sub-data in the text category data.

[0129] In this embodiment, in each text data in each text category sub-data in the text category data, by combining its word segmentation data and semantic expansion data, the corresponding text expansion data is generated. The text expansion data is an information set after semantic expansion of the text. Based on the word segmentation in the original text and the newly added data obtained through semantic expansion, the semantic depth and breadth of the text are enhanced.

[0130] In this embodiment, by comparing the text data and its corresponding text expansion data, their consistency is evaluated. The consistency evaluation mainly measures whether the expansion data accurately and effectively supplements the semantics of the original text and whether it maintains the original theme and structure of the text.

[0131] In this embodiment, in all text category sub-data, determine the consistency between each text and its expansion data, and summarize the evaluation results to generate an expansion evaluation report. The report provides feedback on the effect of the text expansion process, demonstrating the accuracy, quality, and consistency issues of semantic expansion.

[0132] The beneficial effects of the above technical solution: Evaluating semantic expansion data and generating an expansion evaluation report can ensure that the generated expansion content is consistent with the semantic coherence of the English corpus, improve the accuracy and flexibility of language understanding and generation, achieve a systematic enhancement of the semantic depth in the English corpus, and improve the intelligence of natural language processing tasks in the practical application of the English corpus.

[0133] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0134] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0135] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An AI-based method for semantic coherence mining and expansion of English corpora, characterized in that Including: Step 1: Classify all text data in the English corpus to determine text category data, and determine text tokenization data based on the text category data; Step 2: Conduct semantic analysis on the text tokenization data to determine the semantic graph of the English corpus; Step 3: Conduct coherence analysis on the semantic graph to determine semantic coherence data, and determine semantic expansion data based on the text tokenization data and the semantic coherence data; Step 4: Evaluate the semantic expansion data and generate an expansion evaluation report; Among them, Step 3 includes: Based on the text category data, text tokenization data, and semantic graph, calculate the word semantic coherence value and the passage semantic coherence value of each token; Calculate the semantic coherence value of each token based on the word semantic coherence value and the passage semantic coherence value of each token; Based on the semantic coherence values of all tokens in the text tokenization data, determine the semantic coherence data of the English corpus; Input the semantic coherence data and the text tokenization data into the semantic coherence model, and determine the semantic expansion sub-data of each analysis in the text tokenization data based on the output result of the semantic coherence model; Based on the semantic expansion sub-data of all analyses in the text tokenization data, determine the semantic expansion data of the English corpus; Among them, based on the text category data, text tokenization data, and semantic graph, calculating the word semantic coherence value and the passage semantic coherence value of each token includes: Among them, C1 ai represents the word semantic coherence value of the i-th word segmentation in the text segmentation data in the a-th text category sub-data of the text category data, C2 abi represents the text semantic coherence value of the i-th word segmentation in the text segmentation data in the b-th text data of the a-th text category sub-data of the text category data, V i represents the word vector of the i-th word segmentation in the sub-category word segmentation set of all text category sub-data in the text segmentation data, V aik represents the word vector of the k-th word segmentation connected to the i-th word segmentation in the sub-semantic graph of the a-th text category sub-data in the semantic graph, V aj The word vector of the j-th word segmentation not connected to the i-th word segmentation in the sub-semantic graph of the a-th text category sub-data in the semantic graph, aiN1 represents the number of word segmentations connected to the i-th word segmentation in the sub-semantic graph of the a-th text category sub-data in the semantic graph, aN2 represents the number of word segmentations in the sub-semantic graph of the a-th text category sub-data in the semantic graph, τ1 represents the first adjustment parameter, τ2 represents the second adjustment parameter, I1 ai represents the first indicator function of the i-th word segmentation in the text segmentation data for the a-th text category sub-data of the text category data, Ca a represents the sub-category word segmentation set of the a-th text category sub-data in the text segmentation data, P abi represents the word segmentation position of the i-th word segmentation in the text segmentation data in the b-th text data of the a-th text category sub-data of the text category data, wi abi represents the number of word segmentations in the word segmentation window of the i-th word segmentation in the text segmentation data in the b-th text data of the a-th text category sub-data of the text category data, represents the word vectors of the (i - c)-th and (i + c)-th word segmentations of the i-th word segmentation in the text segmentation data in the b-th text data of the a-th text category sub-data of the text category data, I2 abi represents the second indicator function of the i-th word segmentation in the text segmentation data for the b-th text data of the a-th text category sub-data of the text category data, δ represents the word segmentation window decay parameter, Ct ab represents the text segmentation data of the b-th text data of the a-th text category sub-data of the text category data.

2. The method for semantic coherence mining and expansion of an English corpus based on AI according to claim 1, wherein Classify all text data in the English corpus to determine text category data, including: Determine the text characteristics of each text data in the English corpus, where the text characteristics include text theme, text structure, language style, and target audience; Determine text classification rules based on the text characteristics of all text data in the English corpus, match each text data in the English corpus with the text classification rules, and determine the category label of each text data in the English corpus; Based on all text data with the same category label, determine the text category sub-data of each category label; Based on the text category sub-data of all category labels, determine the text category data of the English corpus.

3. The AI-based method for semantic coherence mining and expansion of English corpora according to claim 1, wherein Determine text tokenization data based on the text category data, including: Perform tokenization processing and stop word removal on each text data in each text category sub-data in the text category data, and determine the passage token set of each text data in each text category sub-data in the text category data, where the passage token set includes multiple tokens in the corresponding text data; Based on the passage token set of each text data in each text category sub-data in the text category data, determine the passage tokenization data of each text data; Based on the passage token sets of all text data in each text category sub-data in the text category data, determine the sub-category token set of each text category sub-data; Based on the sub-category token sets of all text category sub-data, determine the text tokenization data of the English corpus.

4. The method for semantic coherence mining and expansion of an English corpus based on AI according to claim 1, characterized in that Conduct semantic analysis on the text tokenization data to determine the semantic graph of the English corpus, including: Based on all the text data in each text category sub-data of the text category data, perform part-of-speech analysis and sentence analysis on each token in the sub-category token set of each text category sub-data in the text token data to determine the word vector of each token in the sub-category token set of each text category sub-data; Input the sub-category token set of each text category sub-data in the text token data and the word vector of each token in the sub-category token set into the token semantic recognition model, and determine the entity token set and sub-category token relationship of each text category sub-data based on the output structure of the token semantic recognition model; Construct the sub-semantic graph of each text category sub-data based on the sub-category token set, entity token set, and sub-category token relationship of each text category sub-data; Determine the semantic graph of the English corpus based on the sub-semantic graphs of all text category sub-data; 5. The method for mining and expanding semantic coherence of an English corpus based on AI according to claim 1, wherein Calculate the semantic coherence value of each token based on the word semantic coherence value and text semantic coherence value of each token, including: Among them, C i represents the semantic coherence value of the i-th word segmentation in the sub-category word segmentation set of all text category sub-data in the text word segmentation data, N3 represents the number of text category sub-data in the text category data, and aN4 represents the number of text data in the a-th text category sub-data in the text category data.

6. The method for semantic coherence mining and expansion of an English corpus based on AI according to claim 3, wherein Evaluate the semantic extension data and generate an extended evaluation report, including: Based on the text token data and semantic extension data of each text data in each text category sub-data of the text category data, determine the text extension data of each text data in each text category sub-data of the text category data; Evaluate the consistency between each text data and the corresponding text extension data in each text category sub-data of the text category data; Generate an extended evaluation report based on the consistency of all text data in all text category sub-data of the text category data.

Citation Information

Patent Citations

  • Medical text classification method based on knowledge graph and multi-head pooling graph convolution

    CN116975282A

  • Advertisement material data annotation method and system based on AI

    CN118761812A