English corpus semantic coherence mining and expanding method based on AI
Through the coherent semantic mining and expansion method of the English corpus based on AI, the problem of insufficient semantic understanding accuracy and depth in the existing technology is solved, the accuracy and flexibility of semantic expansion are achieved, and the intelligence of natural language processing tasks is improved.
Patent Information
- Application Number
- CN202510222646.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-27
AI Technical Summary
The prior art is difficult to deeply understand the deep semantic relationships and context information in text, resulting in insufficient accuracy and depth of semantic analysis of corpus.
Through the semantic coherent mining and expansion method of the English corpus based on AI, text category data is determined to obtain text word segmentation data, text word segmentation data is analyzed to construct a semantic map, semantic map is performed to determine semantic coherent data, and semantic expansion data is determined based on these data, and finally the semantic expansion data is evaluated to generate an extension evaluation report.
It improves the accuracy and depth of semantic understanding, improves the accuracy and flexibility of semantic expansion, ensures that the generated extended content is consistent with the semantic coherence of the English corpus, realizes systematic enhancement of semantic depth in the English corpus, and improves the intelligence of natural language processing tasks.
Smart Images

Figure FT_1 
Figure QLYQS_1 
Figure QLYQS_2
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to an AI-based English corpus semantic coherence mining and expansion method. Background Art
[0002] With the continuous development of natural language processing (NLP) technology, corpus construction and semantic analysis have become one of the core tasks in the field of artificial intelligence. Early text processing technologies mainly relied on simple statistical models and rule-based methods, such as word frequency-based vocabulary analysis and word segmentation. However, these methods usually cannot deeply understand the deep semantic relationships and contextual information in the text.
[0003] With the introduction of deep learning and machine learning technologies, especially in the late 2010s, language models based on neural networks (such as Word2Vec, BERT, etc.) have shown significant advantages in processing language data. These models can capture more complex semantic relationships between words in the text, which has promoted the development of corpus semantic analysis. Semantic coherence analysis and semantic extension techniques have begun to become key technologies for improving the depth and consistency of corpus semantic understanding.
[0004] Therefore, the present invention provides an AI-based English corpus semantic coherence mining and expansion method. Summary of the invention
[0005] The present invention provides an AI-based English corpus semantic coherence mining and expansion method, which determines text segmentation data by determining text category data, analyzes the text analysis data to determine a semantic graph, performs coherence analysis on the semantic graph to determine semantic coherence data, determines semantic expansion data based on the text segmentation data and the semantic coherence data, evaluates the semantic expansion data, and generates an expansion evaluation report, which can improve the precision and depth of semantic understanding, improve the accuracy and flexibility of semantic expansion, ensure that the generated expansion content is consistent with the semantic coherence of the English corpus, achieve systematic enhancement of the semantic depth in the English corpus, and improve the intelligence of natural language processing tasks of the English corpus in practical applications.
[0006] The present invention provides an AI-based English corpus semantic coherence mining and expansion method, comprising:
[0007] Step 1: Classify all text data in the English corpus to determine text category data, and determine text segmentation data based on the text category data;
[0008] Step 2: Perform semantic analysis on the text segmentation data to determine the semantic graph of the English corpus;
[0009] Step 3: Perform a coherent analysis on the semantic graph to determine the semantic coherence data, and determine the semantic expansion data based on the text segmentation data and the semantic coherence data;
[0010] Step 4: Evaluate the semantic extension data and generate an extension evaluation report.
[0011] According to the AI-based English corpus semantic coherence mining and expansion method provided by the present invention, all text data in the English corpus are classified to determine text category data, including:
[0012] Determine the text characteristics of each text data in the English corpus, where the text characteristics include text topic, text structure, language style and target audience;
[0013] Determine text classification rules based on the text features of all text data in the English corpus, match each text data in the English corpus with the text classification rules, and determine the category label of each text data in the English corpus;
[0014] Based on all text data with the same category label, determine text category sub-data for each category label;
[0015] Based on the text category sub-data of all category labels, text category data of the English corpus is determined.
[0016] According to the AI-based English corpus semantic coherence mining and expansion method provided by the present invention, text segmentation data is determined based on text category data, including:
[0017] Performing word segmentation processing and stop word removal on each text data in each text category sub-data in the text category data, and determining a word segmentation set for each text data in each text category sub-data in the text category data, wherein the word segmentation set includes a plurality of word segments in the corresponding text data;
[0018] Determine the segment word data of each piece of text data based on the segment word set of each piece of text data in each text category sub-data in the text category data;
[0019] Determine a subcategory segmentation set for each text category subdata based on the article segmentation set of all text data in each text category subdata in the text category data;
[0020] The text segmentation data of the English corpus is determined based on the subcategory segmentation set of all the text category subdata.
[0021] According to the AI-based English corpus semantic coherence mining and expansion method provided by the present invention, semantic analysis is performed on text segmentation data to determine the semantic map of the English corpus, including:
[0022] Based on all text data in each text category sub-data in the text category data, perform part-of-speech analysis and word-sentence analysis on each segmentation in the sub-category segmentation set of each text category sub-data in the text segmentation data, and determine the word vector of each segmentation in the sub-category segmentation set of each text category sub-data;
[0023] Input the subcategory segmentation set of each text category sub-data in the text segmentation data and the word vector of each segmentation in the subcategory segmentation set into the segmentation semantic recognition model, and determine the entity segmentation set and subcategory segmentation relationship of each text category sub-data based on the output structure of the segmentation semantic recognition model;
[0024] Construct a sub-semantic graph for each text category sub-data based on the sub-category segmentation set, entity segmentation set, and sub-category segmentation relationship of each text category sub-data;
[0025] The semantic graph of the English corpus is determined based on the sub-semantic graphs of all text category sub-data.
[0026] According to the AI-based English corpus semantic coherence mining and expansion method provided by the present invention, a semantic graph is coherently analyzed to determine semantic coherence data, and semantic expansion data is determined based on text segmentation data and semantic coherence data, including:
[0027] Based on the text category data, text segmentation data and semantic graph, the word semantic coherence value and the text semantic coherence value of each segmentation are calculated;
[0028] Calculate the semantic coherence value of each segmentation based on the word semantic coherence value and the text semantic coherence value of each segmentation;
[0029] Determine semantic coherence data of an English corpus based on semantic coherence values of all segmented words in the text segmented data;
[0030] Inputting the semantic coherence data and the text segmentation data into the semantic coherence model, and determining the semantic extension sub-data for each analysis in the text segmentation data based on the output result of the semantic coherence model;
[0031] Based on all analyzed semantic extension sub-data in the text segmentation data, the semantic extension data of the English corpus is determined.
[0032] According to the AI-based English corpus semantic coherence mining and expansion method provided by the present invention, based on text category data, text segmentation data and semantic graph, the word semantic coherence value and the text semantic coherence value of each segmentation are calculated, including:
[0033] ;
[0034] ;
[0035] ; ;
[0036] in, Represents the word semantic coherence value of the i-th word segmentation data in the a-th text category sub-data of the text category data, It represents the semantic coherence value of the i-th word segmentation data in the b-th text data of the a-th text category sub-data of the text category data, Represents the word vector of the i-th word in the subcategory word set of all text category subdata in the text word segmentation data, Represents the word vector of the kth word segment connected to the ith word segment in the sub-semantic graph of the ath text category sub-data in the semantic graph, The word vector of the jth word segmentation in the sub-semantic graph of the ath text category sub-data in the semantic graph that is not connected to the i-th word segmentation, Represents the number of words connected to the i-th word in the sub-semantic graph of the a-th text category sub-data in the semantic graph, Represents the number of word segments of the sub-semantic graph of the a-th text category sub-data in the semantic graph, represents the first adjustment parameter, 2 represents the second adjustment parameter, Represents the first indicator function of the i-th word segmentation in the text segmentation data to the a-th text category sub-data of the text category data, Represents the subcategory segmentation set of the a-th text category subdata in the text segmentation data, Indicates the word segmentation position of the i-th word segmentation data in the b-th piece of text data in the a-th text category sub-data of the text category data, It represents the number of word segments of the i-th word segmentation data in the word segmentation window of the b-th text data of the a-th text category sub-data of the text category data. Represents the word vector of the icth word and the i+cth word of the i-th word segmentation data in the b-th text data of the a-th text category sub-data of the text category data, The second indicator function representing the bth text data of the ath text category sub-data of the i-th word segmentation in the text segmentation data, represents the word segmentation window attenuation parameter, Indicates the piece segmentation data of the b-th piece of text data of the a-th text category sub-data of the text category data.
[0037] According to the AI-based English corpus semantic coherence mining and expansion method provided by the present invention, the semantic coherence value of each segmentation is calculated based on the word semantic coherence value and the text semantic coherence value of each segmentation, including:
[0038] ;
[0039] in, represents the semantic coherence value of the ith segmentation of the subcategory segmentation set of all text category subdata in the text segmentation data, N3 represents the number of text category subdata in the text category data, The number of text data in the ath text category sub-data in the text category data.
[0040] According to the AI-based English corpus semantic coherence mining and expansion method provided by the present invention, the semantic extension data is evaluated and an expansion evaluation report is generated, including:
[0041] Determine the article expansion data of each piece of text data in each text category sub-data in the text category data based on the article segmentation data and the semantic expansion data of each piece of text data in each text category sub-data in the text category data;
[0042] Evaluate the consistency of each text data and the corresponding extended data in each text category sub-data in the text category data;
[0043] Based on the consistency of all text data in all text category sub-data in the text category data, an extended evaluation report is generated.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] By determining the text segmentation data through the determined text category data, analyzing the text analysis data to determine the semantic map, conducting coherence analysis on the semantic map to determine the semantic coherence data, determining the semantic extension data based on the text segmentation data and the semantic coherence data, evaluating the semantic extension data, and generating an extension evaluation report, the precision and depth of semantic understanding can be improved, the accuracy and flexibility of semantic extension can be improved, and it can be ensured that the generated extension content is consistent with the semantic coherence of the English corpus, so as to achieve a systematic enhancement of the semantic depth in the English corpus and enhance the intelligence of natural language processing tasks of the English corpus in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0047] Figure 1 It is a flow chart of the AI-based English corpus semantic coherence mining and expansion method provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0049] Embodiment 1:
[0050] The embodiment of the present invention provides an AI-based English corpus semantic coherence mining and expansion method, such as Figure 1 As shown, including:
[0051] Step 1: Classify all text data in the English corpus to determine text category data, and determine text segmentation data based on the text category data;
[0052] Step 2: Perform semantic analysis on the text segmentation data to determine the semantic graph of the English corpus;
[0053] Step 3: Perform a coherent analysis on the semantic graph to determine the semantic coherence data, and determine the semantic expansion data based on the text segmentation data and the semantic coherence data;
[0054] Step 4: Evaluate the semantic extension data and generate an extension evaluation report.
[0055] In this embodiment, all text data in the English corpus are analyzed to identify its text features (such as topic, text structure, language style, audience, etc.), and the text is classified based on these features. After classification, text category data is generated, and each text category sub-data represents a specific text type.
[0056] In this embodiment, text segmentation processing is performed based on text category data to break each text into several words or phrases. The segmented data provides a basis for subsequent semantic analysis.
[0057] In this embodiment, semantic analysis aims to deeply understand the semantic relationship between words in the text and determine the semantic map by means of language processing and word segmentation semantic recognition model.
[0058] In this embodiment, the coherence of the word segments in the semantic graph is analyzed, the semantic consistency between the word segments and the text data is identified, and semantic coherence data is generated.
[0059] In this embodiment, the text segmentation data and the semantic coherence data are combined to explore the potential of semantic extension and generate new extended data.
[0060] The beneficial effects of the above technical solution are as follows: determining text segmentation data by determining text category data, analyzing text analysis data to determine a semantic map, performing coherence analysis on the semantic map to determine semantic coherence data, determining semantic extension data based on text segmentation data and semantic coherence data, evaluating the semantic extension data, and generating an extension evaluation report, which can improve the accuracy and depth of semantic understanding, improve the accuracy and flexibility of semantic extension, ensure that the generated extension content is consistent with the semantic coherence of the English corpus, achieve a systematic enhancement of the semantic depth in the English corpus, and improve the intelligence of natural language processing tasks of the English corpus in practical applications.
[0061] Embodiment 2:
[0062] The embodiment of the present invention provides an AI-based English corpus semantic coherence mining and expansion method, which classifies all text data in the English corpus to determine text category data, including:
[0063] Determine the text characteristics of each text data in the English corpus, where the text characteristics include text topic, text structure, language style and target audience;
[0064] Determine text classification rules based on the text features of all text data in the English corpus, match each text data in the English corpus with the text classification rules, and determine the category label of each text data in the English corpus;
[0065] Based on all text data with the same category label, determine text category sub-data for each category label;
[0066] Based on the text category sub-data of all category labels, text category data of the English corpus is determined.
[0067] In this embodiment, each text in the corpus is analyzed to extract text features, which include: text subject: the main content or topic discussed in the text (for example, technology, education, politics, etc.); text structure: the organization of the text, such as paragraph structure, sentence construction, etc., analyzing the typesetting and paragraph division of the text; language style: the language features used in the text, such as formal or informal language, academic or popular, etc.; target audience: the reader group targeted by the text, such as academics, general readers, or experts in a specific industry.
[0068] In this embodiment, text classification rules are defined based on the analyzed text features, that is, classification standards are set for text data based on factors such as the subject, structure, style, and audience of the text.
[0069] In this embodiment, each text is matched with the rules according to the set classification rules to determine the category label to which it belongs. The category label reflects the core content and characteristics of the text.
[0070] In this embodiment, for each category label, all texts matching the label are collected to form a text category sub-data, that is, a set of texts containing the same category label.
[0071] In this embodiment, the text category sub-data under all category labels are integrated to form complete text category data, which is a corpus data organized and classified according to categories.
[0072] The beneficial effects of the above technical solution are: by classifying all text data in the English corpus to determine the text category data, efficient automatic classification can be achieved, the accuracy and depth of semantic understanding can be improved, and a data basis can be provided for determining text analysis data.
[0073] Embodiment 3:
[0074] The embodiment of the present invention provides an AI-based English corpus semantic coherence mining and expansion method, which determines text segmentation data based on text category data, including:
[0075] Performing word segmentation processing and stop word removal on each text data in each text category sub-data in the text category data, and determining a word segmentation set for each text data in each text category sub-data in the text category data, wherein the word segmentation set includes a plurality of word segments in the corresponding text data;
[0076] Determine the segment word data of each piece of text data based on the segment word set of each piece of text data in each text category sub-data in the text category data;
[0077] Determine a subcategory segmentation set for each text category subdata based on the article segmentation set of all text data in each text category subdata in the text category data;
[0078] The text segmentation data of the English corpus is determined based on the subcategory segmentation set of all the text category subdata.
[0079] In this embodiment, each text is segmented into independent segments, and then stop words (such as "the", "and", "is", and other common but meaningless words) are removed, and only meaningful words are retained to obtain a segmentation set for each text, that is, a set of all meaningful segmentations in the text.
[0080] In this embodiment, the text data is compared according to the segmented word set of each text, and the segmented word data after removing the stop words and retaining only the meaningful segmented words is determined.
[0081] In this embodiment, for all text data in each text category sub-data, its article segmentation set is collected, and the common features of the text under the category are analyzed to extract the sub-category segmentation set representing the category, that is, the segmentations that appear more frequently in the text of the category.
[0082] In this embodiment, the sub-category segmentations in all category sub-data are integrated to eventually form a global text segmentation data, which represents the segmentation features of the entire English corpus, facilitating further semantic analysis, information extraction and other applications.
[0083] The beneficial effects of the above technical solution are as follows: determining text segmentation data based on text category data can refine the lexical analysis of the text, improve the accuracy of classification, enhance the ability to capture the semantics of the text, and provide a more accurate data basis for determining the semantic map.
[0084] Embodiment 4:
[0085] The embodiment of the present invention provides an AI-based English corpus semantic coherence mining and expansion method, which performs semantic analysis on text segmentation data and determines the semantic map of the English corpus, including:
[0086] Based on all text data in each text category sub-data in the text category data, perform part-of-speech analysis and word-sentence analysis on each segmentation in the sub-category segmentation set of each text category sub-data in the text segmentation data, and determine the word vector of each segmentation in the sub-category segmentation set of each text category sub-data;
[0087] Input the subcategory segmentation set of each text category sub-data in the text segmentation data and the word vector of each segmentation in the subcategory segmentation set into the segmentation semantic recognition model, and determine the entity segmentation set and subcategory segmentation relationship of each text category sub-data based on the output structure of the segmentation semantic recognition model;
[0088] Construct a sub-semantic graph for each text category sub-data based on the sub-category segmentation set, entity segmentation set, and sub-category segmentation relationship of each text category sub-data;
[0089] The semantic graph of the English corpus is determined based on the sub-semantic graphs of all text category sub-data.
[0090] In this embodiment, part-of-speech analysis is performed on the words in each subcategory of the word segmentation, that is, the part of speech of each word (such as noun, verb, adjective, etc.) is identified. At the same time, word and sentence analysis is performed to examine the structure and grammatical relationship of the vocabulary in the sentence. Based on these analyses, word vectors are generated to capture the semantic information of the vocabulary.
[0091] In this embodiment, the subcategory segmentation of each text category sub-data and the word vector of each segmentation are input into a segmentation semantic recognition model. The model identifies the role of each segmentation based on context and semantic understanding, and outputs entity segmentations (such as names of people, places, time, etc.) and the relationship between words (such as "belongs to", "contains", etc.).
[0092] In this embodiment, the word segmentation semantic recognition model extracts semantic information from all text data in each text category sub-data, identifies entities, and establishes mappings between entities, relations, and attributes.
[0093] In this embodiment, a sub-semantic graph is constructed based on the sub-category segmentation, entity segmentation and segmentation relationship of each text category sub-data. This sub-semantic graph shows the relationship between different words and entities in the corresponding text category sub-data, reflecting the semantic network structure within the text.
[0094] In this embodiment, the sub-semantic graphs of all text category sub-data are merged to form a global semantic graph, which shows the vocabulary, entities and their mutual relationships in the entire English corpus.
[0095] The beneficial effects of the above technical solution are: semantic analysis of text segmentation data and determination of the semantic map of the English corpus can deeply explore the complex semantic relationship between words and entities in the text, improve the accuracy of text comprehension and semantic reasoning, and enhance the comprehensive understanding of the semantic structure of the English corpus.
[0096] Embodiment 5:
[0097] The embodiment of the present invention provides an AI-based English corpus semantic coherence mining and expansion method, which performs coherence analysis on a semantic graph to determine semantic coherence data, and determines semantic expansion data based on text segmentation data and semantic coherence data, including:
[0098] Based on the text category data, text segmentation data and semantic graph, the word semantic coherence value and the text semantic coherence value of each segmentation are calculated;
[0099] Calculate the semantic coherence value of each segmentation based on the word semantic coherence value and the text semantic coherence value of each segmentation;
[0100] Determine semantic coherence data of an English corpus based on semantic coherence values of all segmented words in the text segmented data;
[0101] Inputting the semantic coherence data and the text segmentation data into the semantic coherence model, and determining the semantic extension sub-data for each analysis in the text segmentation data based on the output result of the semantic coherence model;
[0102] Based on all analyzed semantic extension sub-data in the text segmentation data, the semantic extension data of the English corpus is determined.
[0103] In this embodiment, the word semantic coherence value and the text semantic coherence value of each word segment are combined to calculate the final semantic coherence value of the word segment, which represents the semantic coherence and consistency of the word in the text.
[0104] In this embodiment, semantic coherence data of the entire corpus is constructed according to the semantic coherence values of the segmented words in all the texts, which reflects the semantic consistency and coherence of the entire English corpus.
[0105] In this embodiment, the determined semantic coherence data and text segmentation data are input into a semantic coherence model, which can analyze which segmentations or concepts in the text can be expanded and output semantic expansion sub-data, that is, under the semantic coherence framework, the model identifies semantic information that may be expanded.
[0106] In this embodiment, the semantic extension sub-data of each analysis is outputted according to the model, summarized and integrated, and finally the semantic extension data of the entire corpus is determined, that is, the expanded semantic information set in the corpus.
[0107] The beneficial effects of the above technical solution are as follows: coherence analysis of the semantic graph is performed to determine the semantic coherence data, which can quantify the semantic coherence, improve the understanding of the internal semantic structure of the text, perform effective semantic expansion in the global scope of the English corpus and ensure that the generated expansion content is consistent with the semantic coherence of the English corpus.
[0108] Embodiment 6:
[0109] The embodiment of the present invention provides an AI-based English corpus semantic coherence mining and expansion method, which calculates the word semantic coherence value and the text semantic coherence value of each segmentation based on text category data, text segmentation data and semantic graph, including:
[0110] ;
[0111] ;
[0112] ; ;
[0113] in, Represents the word semantic coherence value of the i-th word segmentation data in the a-th text category sub-data of the text category data, It represents the semantic coherence value of the i-th word segmentation data in the b-th text data of the a-th text category sub-data of the text category data, Represents the word vector of the i-th word in the subcategory word set of all text category subdata in the text word segmentation data, Represents the word vector of the kth word segment connected to the ith word segment in the sub-semantic graph of the ath text category sub-data in the semantic graph, The word vector of the jth word segmentation in the sub-semantic graph of the ath text category sub-data in the semantic graph that is not connected to the i-th word segmentation, Represents the number of words connected to the i-th word in the sub-semantic graph of the a-th text category sub-data in the semantic graph, Represents the number of word segments of the sub-semantic graph of the a-th text category sub-data in the semantic graph, represents the first adjustment parameter, 2 represents the second adjustment parameter, Represents the first indicator function of the i-th word segmentation in the text segmentation data to the a-th text category sub-data of the text category data, Represents the subcategory segmentation set of the a-th text category subdata in the text segmentation data, Indicates the word segmentation position of the i-th word segmentation data in the b-th piece of text data in the a-th text category sub-data of the text category data, It represents the number of word segments of the i-th word segmentation data in the word segmentation window of the b-th text data of the a-th text category sub-data of the text category data. Represents the word vector of the icth word and the i+cth word of the i-th word segmentation data in the b-th text data of the a-th text category sub-data of the text category data, The second indicator function representing the bth text data of the ath text category sub-data of the i-th word segmentation in the text segmentation data, represents the word segmentation window attenuation parameter, Indicates the piece segmentation data of the b-th piece of text data of the a-th text category sub-data of the text category data.
[0114] In this embodiment, the word semantic coherence value It represents the coherence of all the segmentations of the i-th segmentation of the text segmentation data in the subcategory segmentation set of the a-th text category subdata of the text category data.
[0115] In this embodiment, the semantic coherence value of the article It represents the coherence of all analyses within the word segmentation window of the i-th word segmentation of the text segmentation data in the b-th piece of text data of the a-th text category sub-data of the text category data.
[0116] In this embodiment, the word segmentation window is determined according to the text length of the corresponding word in the word segmentation data of the text data. The longer the word segmentation length of the word segmentation data, the larger the word segmentation window, and the shorter the word segmentation length of the word segmentation data, the smaller the word segmentation window.
[0117] The beneficial effects of the above technical solution are as follows: based on text category data, text segmentation data and semantic maps, the word semantic coherence value and the paragraph semantic coherence value of each segmentation are calculated, which can provide a data basis for quantifying the semantic coherence of each segmentation, so as to enhance the understanding of the internal semantic structure of the text and provide a data basis for effective semantic expansion in the global scope of the English corpus.
[0118] Embodiment 7:
[0119] The embodiment of the present invention provides an AI-based English corpus semantic coherence mining and expansion method, which calculates the semantic coherence value of each segmentation based on the word semantic coherence value and the text semantic coherence value of each segmentation, including:
[0120] ;
[0121] in, represents the semantic coherence value of the ith segmentation of the subcategory segmentation set of all text category subdata in the text segmentation data, N3 represents the number of text category subdata in the text category data, The number of text data in the ath text category sub-data in the text category data.
[0122] In this embodiment, the semantic coherence value It represents integrating the word semantic coherence values of all text category sub-data in the text category data and the article semantic coherence values of all text data of all text category sub-data in the text category data.
[0123] The beneficial effects of the above technical solution are as follows: the semantic coherence value of each word segment is calculated based on the word semantic coherence value and the paragraph semantic coherence value of each word segment, which can provide a data basis for determining the semantic coherence data, improve the understanding of the internal semantic structure of the text, and provide a data basis for effective semantic expansion in the global scope of the English corpus.
[0124] Embodiment 8:
[0125] The embodiment of the present invention provides an AI-based English corpus semantic coherence mining and expansion method, which evaluates semantic extension data and generates an expansion evaluation report, including:
[0126] Determine the article expansion data of each piece of text data in each text category sub-data in the text category data based on the article segmentation data and the semantic expansion data of each piece of text data in each text category sub-data in the text category data;
[0127] Evaluate the consistency of each text data and the corresponding extended data in each text category sub-data in the text category data;
[0128] Based on the consistency of all text data in all text category sub-data in the text category data, an extended evaluation report is generated.
[0129] In this embodiment, in each text data in each text category sub-data in the text category data, the corresponding article expansion data is generated by combining its article segmentation data and semantic expansion data. The article expansion data is a collection of information after the text is semantically expanded. The semantic depth and breadth of the text are enhanced based on the word segmentation in the original text and the new data obtained through semantic expansion.
[0130] In this embodiment, the consistency of text data and its corresponding extended data is evaluated by comparing them. The consistency evaluation mainly measures whether the extended data accurately and effectively supplements the semantics of the original text and whether the original theme and structure of the text are maintained.
[0131] In this embodiment, the consistency between each text and its extended data is determined in all text category sub-data, and the evaluation results are summarized to generate an extension evaluation report. The report provides feedback on the effect of the text extension process and shows the accuracy, quality and consistency of the semantic extension.
[0132] The beneficial effects of the above technical solution are: evaluating the semantic extension data and generating an extension evaluation report, which can ensure that the generated extension content is consistent with the semantic coherence of the English corpus, improve the accuracy and flexibility of language understanding and generation, achieve a systematic enhancement of the semantic depth in the English corpus, and improve the intelligence of natural language processing tasks of the English corpus in practical applications.
[0133] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0134] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0135] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An AI-based English corpus semantic coherence mining and expansion method, characterized in that: include: Step 1: Classify all text data in the English corpus to determine text category data, and determine text segmentation data based on the text category data; Step 2: Perform semantic analysis on the text segmentation data to determine the semantic graph of the English corpus; Step 3: Perform a coherent analysis on the semantic graph to determine the semantic coherence data, and determine the semantic expansion data based on the text segmentation data and the semantic coherence data; Step 4: Evaluate the semantic extension data and generate an extension evaluation report.
2. The AI-based English corpus semantic coherence mining and expansion method according to claim 1, characterized in that: Classify all text data in the English corpus to determine text category data, including: Determine the text characteristics of each text data in the English corpus, where the text characteristics include text topic, text structure, language style and target audience; Determine text classification rules based on the text features of all text data in the English corpus, match each text data in the English corpus with the text classification rules, and determine the category label of each text data in the English corpus; Based on all text data with the same category label, determine text category sub-data for each category label; Based on the text category sub-data of all category labels, text category data of the English corpus is determined.
3. The AI-based English corpus semantic coherence mining and expansion method according to claim 1, characterized in that: Determine text segmentation data based on text category data, including: Performing word segmentation processing and stop word removal on each text data in each text category sub-data in the text category data, and determining a word segmentation set for each text data in each text category sub-data in the text category data, wherein the word segmentation set includes a plurality of word segments in the corresponding text data; Determine the segment word data of each piece of text data based on the segment word set of each piece of text data in each text category sub-data in the text category data; Determine a subcategory segmentation set for each text category subdata based on the article segmentation set of all text data in each text category subdata in the text category data; The text segmentation data of the English corpus is determined based on the subcategory segmentation set of all the text category subdata.
4. The AI-based English corpus semantic coherence mining and expansion method according to claim 1, characterized in that: Perform semantic analysis on the text segmentation data to determine the semantic graph of the English corpus, including: Based on all text data in each text category sub-data in the text category data, perform part-of-speech analysis and word-sentence analysis on each segmentation in the sub-category segmentation set of each text category sub-data in the text segmentation data, and determine a word vector for each segmentation in the sub-category segmentation set of each text category sub-data; Inputting the subcategory segmentation set of each text category sub-data in the text segmentation data and the word vector of each segmentation in the subcategory segmentation set into the segmentation semantic recognition model, and determining the entity segmentation set and the subcategory segmentation relationship of each text category sub-data based on the output structure of the segmentation semantic recognition model; Construct a sub-semantic graph for each text category sub-data based on the sub-category segmentation set, entity segmentation set, and sub-category segmentation relationship of each text category sub-data; The semantic graph of the English corpus is determined based on the sub-semantic graphs of all text category sub-data.
5. The AI-based English corpus semantic coherence mining and expansion method according to claim 1, characterized in that: Perform a coherent analysis on the semantic graph to determine the semantic coherence data, and determine the semantic expansion data based on the text segmentation data and the semantic coherence data, including: Based on the text category data, text segmentation data and semantic graph, the word semantic coherence value and the text semantic coherence value of each segmentation are calculated; Calculate the semantic coherence value of each segmentation based on the word semantic coherence value and the text semantic coherence value of each segmentation; Determine semantic coherence data of an English corpus based on semantic coherence values of all segmented words in the text segmented data; Inputting the semantic coherence data and the text segmentation data into the semantic coherence model, and determining the semantic extension sub-data for each analysis in the text segmentation data based on the output result of the semantic coherence model; Based on all analyzed semantic extension sub-data in the text segmentation data, the semantic extension data of the English corpus is determined.
6. The AI-based English corpus semantic coherence mining and expansion method according to claim 5 is characterized in that: Based on the text category data, text segmentation data and semantic graph, the word semantic coherence value and the text semantic coherence value of each segmentation are calculated, including: ; ; ; ; in, Represents the word semantic coherence value of the i-th word segmentation data in the a-th text category sub-data of the text category data, It represents the semantic coherence value of the i-th word segmentation data in the b-th text data of the a-th text category sub-data of the text category data, Represents the word vector of the i-th word in the subcategory word set of all text category subdata in the text word segmentation data, Represents the word vector of the kth word segment connected to the ith word segment in the sub-semantic graph of the ath text category sub-data in the semantic graph, The word vector of the jth word segmentation in the sub-semantic graph of the ath text category sub-data in the semantic graph that is not connected to the i-th word segmentation, Represents the number of words connected to the i-th word in the sub-semantic graph of the a-th text category sub-data in the semantic graph, Represents the number of word segments of the sub-semantic graph of the a-th text category sub-data in the semantic graph, represents the first adjustment parameter, 2 represents the second adjustment parameter, Represents the first indicator function of the i-th word segmentation in the text segmentation data to the a-th text category sub-data of the text category data, Represents the subcategory segmentation set of the a-th text category subdata in the text segmentation data, Indicates the word segmentation position of the i-th word segmentation data in the b-th piece of text data in the a-th text category sub-data of the text category data, It represents the number of word segments of the i-th word segmentation data in the word segmentation window of the b-th text data of the a-th text category sub-data of the text category data. Represents the word vector of the icth word and the i+cth word of the i-th word segmentation data in the b-th text data of the a-th text category sub-data of the text category data, The second indicator function representing the bth text data of the ath text category sub-data of the i-th word segmentation in the text segmentation data, represents the word segmentation window attenuation parameter, Indicates the piece segmentation data of the b-th piece of text data of the a-th text category sub-data of the text category data.
7. The AI-based English corpus semantic coherence mining and expansion method according to claim 6, characterized in that: The semantic coherence value of each segmented word is calculated based on the word semantic coherence value and the text semantic coherence value of each segmented word, including: ; in, represents the semantic coherence value of the ith segmentation of the subcategory segmentation set of all text category subdata in the text segmentation data, N3 represents the number of text category subdata in the text category data, The number of text data in the ath text category sub-data in the text category data.
8. The AI-based English corpus semantic coherence mining and expansion method according to claim 3, characterized in that: Evaluate semantic extension data and generate an extension evaluation report, including: Determine the article expansion data of each piece of text data in each text category sub-data in the text category data based on the article segmentation data and the semantic expansion data of each piece of text data in each text category sub-data in the text category data; Evaluate the consistency of each text data and the corresponding extended data in each text category sub-data in the text category data; Based on the consistency of all text data in all text category sub-data in the text category data, an extended evaluation report is generated.
Citation Information
Patent Citations
Attribute word mining method and related product
CN116484079A
Medical text classification method based on knowledge graph and multi-head pooling graph convolution
CN116975282A
Advertisement material data annotation method and system based on AI
CN118761812A
Systems and methods for text based knowledge mining
US20210049169A1