English learning corpus dynamic updating method and system

By dynamically updating the English learning corpus, the problem of insufficient data diversity and balance in the existing technology has been solved, efficient updates and quality improvements of the corpus have been achieved, and the development of natural language processing technology has been promoted.

CN119938894APending Publication Date: 2025-05-06UNIV FOR SCI & TECH ZHENGZHOU
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510015530.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing English corpus has problems with data diversity and balance, resulting in poor performance of models in specific fields or topics, limiting the development of natural language processing technology.

Method used

A dynamic update method of English learning corpus is adopted. By assigning the acquisition task to multiple nodes and performing preliminary processing, deeply cleaning and standardizing the corpus data, multi-dimensional classification model is used for multi-dimensional classification, knowledge graph is constructed, and incremental evaluation and dynamic update are performed through information entropy changes.

Benefits of technology

It realizes dynamic balanced update of the corpus, improves the quality of the corpus, enhances the generalization ability of the model in specific fields and topics, and promotes the comprehensive progress of natural language processing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938894A_ABST
    Figure CN119938894A_ABST
Patent Text Reader

Abstract

The invention provides an English learning corpus dynamic updating method and system, and the method comprises the steps: distributing English corpus data collection tasks to a plurality of nodes based on a preset rule, and carrying out the primary processing of the English corpus data collected by the plurality of nodes; performing deep cleaning and standardization processing on the preliminarily processed English corpus to obtain standardized English corpus data; performing multi-dimensional classification on the standardized English corpus data based on a preset multi-label classification model, generating a series of labels, and obtaining a classification result; entities and relations are extracted from the classified English corpora, and a knowledge graph is formed by establishing semantic links and concept mapping; according to the method, incremental evaluation is carried out on the acquired different classified English corpora through the knowledge graph, the evaluation result is obtained, and the English corpora with the evaluation result meeting the preset threshold value are dynamically updated to the corresponding classified English corpus data, so that the quality of the English corpora can be efficiently improved, and dynamic balanced updating of the English corpora is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of dynamic updating of English learning corpus, and in particular to a method and system for dynamic updating of English learning corpus. Background Art

[0002] With the continuous advancement of natural language processing technology, data diversity and balance issues have gradually become one of the key factors affecting model performance, especially in the application of English corpora. Currently, many English corpora tend to focus too much on popular content when they are constructed, such as social media, entertainment news, etc. This type of content is easily collected and included in the corpus due to its wide spread and high frequency of updates. However, the result of this bias is that the distribution of different categories of data in the corpus is extremely uneven, and the corpus of certain professional fields or niche topics is relatively scarce.

[0003] This class imbalance not only limits the model's ability to generalize to specific fields or topics, but may also cause the model to perform poorly when dealing with non-popular or highly professional texts. For example, the model may be better at recognizing and understanding words and expressions related to entertainment and sports, but may be unable to cope with professional terms in fields such as science and law.

[0004] In current natural language processing tasks, especially in machine translation and cross-language retrieval, English is often the central language, and the data size of other languages ​​is relatively small. This data size imbalance may lead to a decline in model performance when processing non-English data, especially in complex tasks that require a deep understanding of semantics and context.

[0005] In summary, the problems with data diversity and balance in the existing English corpus not only affect the performance of the model in specific fields and topics, but also restrict the development of cross-language tasks. In order to promote the overall progress of natural language processing technology, we need to pay more attention to the diversity and balance of data and strive to build a multilingual corpus covering a wide range of topics and fields to better meet the needs of practical applications.

[0006] Therefore, how to overcome the above-mentioned technical problems and defects becomes a key issue that needs to be addressed. Summary of the invention

[0007] In order to overcome the above problems existing in the prior art, the present application provides a method and system for dynamically updating an English learning corpus, which adopts the following technical solutions:

[0008] In a first aspect, the present application provides a method for dynamically updating an English learning corpus, comprising:

[0009] Distribute the English corpus data collection task to multiple nodes based on preset rules, and perform preliminary processing on the English corpus data collected by multiple nodes;

[0010] Perform deep cleaning and standardization on the initially processed English corpus to obtain standardized English corpus data;

[0011] Based on the preset multi-label classification model, the standardized English corpus data is preliminarily classified in multiple dimensions, a series of labels are generated, and the classification results are obtained;

[0012] Extract entities and relations from classified English corpus and form knowledge graph by establishing semantic links and concept mapping;

[0013] The collected English corpora of different categories are incrementally evaluated through the knowledge graph to obtain the evaluation results. The English corpora whose evaluation results meet the preset threshold are dynamically updated to the corresponding classified English corpus data.

[0014] Furthermore, the English corpus data collection task is distributed to multiple nodes based on preset rules, and the English corpus data collected by the multiple nodes is preliminarily processed, including:

[0015] In the process of data collection at multiple nodes, English corpus data is extracted, preliminary data cleaning is performed on the English corpus data through preset regular expressions, and preliminary standardization is performed on the cleaned English corpus data;

[0016] The initially standardized English corpus data is deduplicated to obtain the initially processed English corpus.

[0017] Furthermore, based on the update frequencies of different nodes, corresponding collection times are set for data collection tasks of different nodes.

[0018] Furthermore, before distributing the English corpus data collection task to multiple nodes based on preset rules and performing preliminary processing on the English corpus data collected by the multiple nodes, it includes:

[0019] Conduct statistical analysis on the quantity and usage frequency of corpora in various fields, obtain scarce fields, obtain the entropy values ​​of scarce fields, sort them from high to low according to the entropy values, and update the collection priority of English corpora.

[0020] Furthermore, based on the preset multi-label classification model, the standardized English corpus data is preliminarily classified in multiple dimensions to generate a series of labels and obtain classification results, including:

[0021] Preprocess the standardized English corpus data to obtain preprocessed data;

[0022] The preprocessed English corpus data is input into the preset multi-label classification model. Based on the feature and label mapping relationship learned by the preset multi-label classification model and each dimension of the corpus, the corresponding classification labels are output respectively to form a multi-dimensional label combination as the classification result.

[0023] Furthermore, before the preset multi-label classification model outputs a multi-dimensional label combination, the classification confidence of each label is obtained through the Softmax activation function, that is, the reliability of the classification result of the label is estimated.

[0024] Furthermore, entities and relations are extracted from the classified English corpus, and a knowledge graph is formed by establishing semantic links and concept mapping, including:

[0025] Identify entities in the preliminarily classified English corpus through the preset NER tool, extract the relationship between entities through syntactic analysis and semantic role labeling; connect the extracted entities through semantic relationships to form semantic connections between entities;

[0026] The extracted entities, entity relationships, and semantic connections are stored in the form of a graph database to build the basic architecture of the knowledge graph.

[0027] Furthermore, the collected English corpora of different categories are incrementally evaluated through the knowledge graph to obtain evaluation results. The English corpora whose evaluation results meet the preset threshold are dynamically updated to the corresponding classified English corpus data, including:

[0028] Incremental evaluation of collected English corpora of different categories is performed through the knowledge graph, including: obtaining the change in information entropy of collected English corpora of different categories, performing incremental evaluation on the knowledge graph, and when the change in information entropy of collected English corpora of different categories meets a preset threshold, dynamically updating the collected English corpora to the corresponding classified English corpus data.

[0029] In a second aspect, the present application also provides a system for dynamically updating an English learning corpus, comprising:

[0030] A preliminary processing module, used to distribute the English corpus data collection task to multiple nodes based on preset rules, and perform preliminary processing on the English corpus data collected by multiple nodes;

[0031] A standardized English corpus data acquisition module is used to perform in-depth cleaning and standardization on the initially processed English corpus to obtain standardized English corpus data;

[0032] The hierarchical annotation module is used to perform preliminary multi-dimensional classification of the standardized English corpus data based on a preset multi-label classification model, generate a series of labels, and obtain classification results;

[0033] The knowledge graph construction module is used to extract entities and relations from the classified English corpus and form a knowledge graph by establishing semantic links and concept mapping;

[0034] The dynamic update expansion module is used to perform incremental evaluation on the collected English corpora of different categories through the knowledge graph, obtain the evaluation results, and dynamically update the English corpora whose evaluation results meet the preset threshold to the corresponding classified English corpus data.

[0035] In a third aspect, the present application provides an electronic device, including:

[0036] One or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions, which, when executed by the device, cause the device to perform the method as described in the first aspect.

[0037] In a fourth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes the method described in the first aspect.

[0038] In a fifth aspect, the present application provides a computer program, which, when executed by a computer, is used to execute the method described in the first aspect.

[0039] In one possible design, the program in the fifth aspect may be stored in whole or in part on a storage medium packaged together with the processor, or may be stored in whole or in part on a memory not packaged together with the processor.

[0040] This application has the following beneficial effects:

[0041] 1. This application can make full use of multiple computing resources or multiple collection channels to work in parallel by allocating the collection tasks to multiple nodes at the same time, greatly shortening the time required to collect a large amount of English corpus data.

[0042] 2. This application conducts statistical analysis on the quantity and usage frequency of corpora in various fields, obtains scarce fields, obtains the entropy value of scarce fields, sorts them from high to low according to the entropy value, and updates the collection priority of English corpora. It can effectively improve the quality of English corpora and realize dynamic and balanced updating of the English corpus.

[0043] 3. This application can more finely remove various errors and non-standard content in the corpus through deep cleaning, such as spelling errors, grammatical errors, improper use of punctuation, etc. At the same time, standardization processing can standardize and unify the text in terms of vocabulary, sentence structure, etc., making the corpus more accurate and standardized in language expression, meeting the requirements of high-quality data.

[0044] 4. This application can use a multi-label classification model to simultaneously label English corpus with multiple relevant labels, which better reflects the true attributes and applicable scenarios of the corpus.

[0045] 5. This application integrates the knowledge scattered in various English corpora in the form of entities, relationships, semantic links and concept mappings to construct a structured knowledge representation form of a knowledge graph, which can clearly show the mutual connections and hierarchical structures between different knowledge elements.

[0046] 6. This application uses changes in information entropy to incrementally evaluate and dynamically update the collected English corpus of different categories, which helps to ensure that the knowledge graph can continuously and high-quality absorb new knowledge content, so that it always maintains a good knowledge structure and adaptability to different application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is an exemplary system architecture diagram to which the embodiments of the present application can be applied;

[0048] Figure 2 This is a flow chart of a method for dynamically updating an English learning corpus according to an embodiment of the present application;

[0049] Figure 3 This is a flowchart of the preliminary processing of English corpus data in an embodiment of the present application;

[0050] Figure 4 This is a flowchart of the preliminary classification result processing of the embodiment of the present application;

[0051] Figure 5 This is a system flow chart of an embodiment of the present application;

[0052] Figure 6 It is a schematic diagram of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by technicians in the technical field of the present application; the terms used in the specification of the application herein are only for the purpose of describing specific embodiments and are not intended to limit the present application; the terms "including" and "having" and any variations thereof in the specification and claims of the present application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of the present application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.

[0054] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0055] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0056] like Figure 1 As shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0057] Users can use terminal devices 101, 102, 103 to interact with server 105 through network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0058] Terminal devices 101, 102, 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III, Moving Picture Experts Group Audio Layer 3), MP4 (Moving Picture Experts Group Audio Layer IV, Moving Picture Experts Group Audio Layer 4) players, laptop computers and desktop computers, etc.

[0059] The server 105 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal devices 101 , 102 , and 103 .

[0060] It should be noted that the method for dynamically updating the English learning corpus provided in the embodiment of the present application is generally executed by a server / terminal device, and accordingly, the system for dynamically updating the English learning corpus is generally arranged in a server / terminal device.

[0061] It should be understood that Figure 1The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0062] Continue to refer Figure 2 , the figure shows a flow chart of a method for dynamically updating an English learning corpus of the present application, the method comprising the following steps:

[0063] Step 201 : distribute the English corpus data collection task to multiple nodes based on preset rules, and perform preliminary processing on the English corpus data collected by the multiple nodes.

[0064] In a possible implementation, the English corpus data collection task is distributed to multiple nodes based on preset rules, and the English corpus data collected by the multiple nodes are preliminarily processed. Please refer to Figure 3 , the specific contents include:

[0065] Step 31, in the process of data collection at multiple nodes, extracting English corpus data, performing preliminary data cleaning on the English corpus data by using a preset regular expression, and performing preliminary standardization on the cleaned English corpus data;

[0066] For example, English corpus captured from web pages may contain a large number of HTML tags, such as 、 , Etc. The regular expression <.*?> can be used to match and remove these tags, so that the text content only retains pure text information. For example, Thisisa sample text. Cleaned to This is a sample text.

[0067] In some web pages, the main text is mixed with a lot of script codes, style codes and other irrelevant HTML elements. You can use regular expressions to extract the main text according to the characteristics of the web page structure. For example, for common web page layouts, you can match the content between .*? and further remove non-text elements to obtain the main text information.

[0068] There may be some unnecessary special characters in the corpus, such as too many punctuation marks (such as multiple consecutive exclamation marks!!! or question marks???), garbled characters (such as \x00, etc.), and some non-English characters in specific fields (such as mathematical symbols, chemical symbols, etc., if the corpus is mainly used for general English learning). The regular expression [^\w\s.,:;? ! \"] can match and remove these special characters, and clean Hello!!!World??This is a test.\x00 into Hello World.Thisis atest., making the text more standardized.

[0069] If the corpus involves the expression of numbers and units, there may be inconsistencies, such as 10kg, 10kgs, 10kilograms, etc. Regular expressions can be used to unify them into a standard format. For example, the above expressions can be converted into 10kilograms, making the expression of numbers and units in the corpus more standardized and unified, which is convenient for subsequent analysis and processing.

[0070] Standardized date and time formats: For corpora containing date and time, there may be multiple ways of expressing them, such as 01 / 02 / 2024 (which may be month / day / year or day / month / year), Jan 2, 2024, 2024-01-02, etc. Different date and time formats can be identified through regular expressions and converted into a standard format, such as YYYY-MM-DD, which is more convenient and accurate when processing time-related corpora.

[0071] Step 32, performing deduplication processing on the initially standardized English corpus data to obtain initially processed English corpus.

[0072] In a possible implementation, the words of the English corpus are mapped to vectors, and the similarity between the two corpus vectors is obtained by cosine similarity. When the similarity exceeds a preset threshold, one of the English corpus is removed. Example: "The brown dog is lying on the grass.", "A dog which is brown is resting on the grass." After calculating the word vector and cosine similarity, it is found that the semantic similarity of these two corpora is high. According to actual needs, for example, the first corpus can be retained and the second can be removed, so as to reduce the semantic duplication in the corpus.

[0073] In a possible implementation, before distributing the English corpus data collection task to multiple nodes based on a preset rule and preliminarily processing the English corpus data collected by the multiple nodes, the process includes:

[0074] Conduct statistical analysis on the quantity and usage frequency of corpora in various fields, obtain scarce fields, obtain the entropy value of scarce fields, sort them from high to low according to the entropy value, update the collection priority of English corpus, and make the distribution of English corpus in various fields gradually balanced.

[0075] In a possible implementation, the formula for obtaining the entropy value of the scarce field is: Where X represents a scarce field, x i represents different categories or types within the scarce field, p(x i ) represents the proportion of different categories or types. For example, in the cultural field, historical culture corpus accounts for 30%, folk culture corpus accounts for 50%, and art culture corpus accounts for 20%. Then the entropy value of this scarce field is calculated as:

[0076] H(culture)=-(0.3*log 2 0.3+0.5*log 2 0.5+0.2log 2 0.2)≈1.47

[0077] Sort each scarce field from high to low according to the calculated entropy value. The higher the entropy value, the more uneven the distribution of English corpus in the field, and the more priority is needed to supplement the corpus to make it balanced, so the priority of the scarce field in English corpus collection is increased accordingly. For example, after sorting, it is found that the entropy value of the cultural field is the highest, followed by the science and technology field. Then the priority of the corpus collection task in the cultural field is set to the highest, and various high-quality corpus resources in the cultural field are collected first. Later, the corpus of other fields is collected in order according to the priority, so as to gradually make the distribution of English corpus in various fields gradually balanced.

[0078] Different English learners have different needs for knowledge in various fields. Some focus on academic English, while others focus more on business communication English, etc. This application obtains the entropy values ​​of scarce fields for sorting, and prioritizes the collection of English corpora in scarce fields, so that the distribution of English corpora in various fields tends to be balanced, which can provide rich and comprehensive learning materials for all types of learners, and avoid limiting learners' knowledge acquisition and ability improvement due to too little corpus in certain fields. At the same time, a balanced corpus structure helps to improve the efficiency of resource utilization, allowing each corpus to play a role in the right place, while also enhancing the adaptability of the corpus to different learning scenarios and needs, and overall improving its quality and value as an English learning resource.

[0079] In a possible implementation, corresponding collection times are set for data collection tasks at different nodes based on the update frequencies of different nodes, to ensure timely acquisition of newly released English learning corpora from different data sources.

[0080] For example, different data sources include but are not limited to academic journal libraries, educational institution resources, government public documents, international organization documents, professional blogs, technical communities, online education platforms, multilingual news websites, etc.

[0081] Step 202: perform deep cleaning and standardization on the initially processed English corpus to obtain standardized English corpus data.

[0082] For example, the English corpus collected from different data sources is deeply cleaned and standardized, including: unifying the rules for the use of punctuation marks, correcting common grammatical errors and spelling errors, and converting non-standard English expressions into standard forms. Through standardization, words and phrases with similar meanings but different expressions are unified into standard forms (for example, "USA", "United States", "The United States of America", etc. are unified into "United States"), which can reduce the ambiguity of semantic understanding caused by the diversity of language expressions.

[0083] Step 203, based on a preset multi-label classification model, a preliminary multi-dimensional classification is performed on the standardized English corpus data, a series of labels are generated, and a classification result is obtained.

[0084] In a possible implementation, a preliminary multi-dimensional classification is performed on the standardized English corpus data based on a preset multi-label classification model to generate a series of labels and obtain the classification results. Please refer to Figure 4 , the specific contents include:

[0085] Step 41, preprocessing the standardized English corpus data to obtain preprocessed data. The preprocessing step includes converting the text into lowercase letters, performing word segmentation (using natural language processing tools such as NLTK, spaCy, etc. to segment the text into word sequences), removing stop words (high-frequency words such as "the", "and", "is" that have little effect on semantic judgment), etc. Through these operations, the input text is made more standardized and concise, which is convenient for the model to extract effective features for classification judgment.

[0086] Step 42, input the preprocessed English corpus data into a preset multi-label classification model, and output corresponding classification labels based on the feature-label mapping relationship learned by the preset multi-label classification model and each dimension of the corpus, to form a multi-dimensional label combination as a classification result. For example, for an English article, the label output by the model may be [intermediate, technology, formal written language], indicating that its language difficulty is intermediate, involves the field of technology, and the language style is formal written language.

[0087] In a possible implementation, before the preset multi-label classification model outputs a multi-dimensional label combination, the classification confidence of each label is obtained through the Softmax activation function, that is, the reliability of the classification result of the label is estimated. For example, for the label "technology", the output confidence is 0.8, which means that there is a high probability that the corpus belongs to the field of technology.

[0088] In an embodiment of the present application, the preset multi-label classification model can use a Transformer-based pre-trained model such as BERT. During the training process of the preset multi-label classification model, English corpus data is collected and annotated, and the collected English corpus data is used as the input of the preset multi-label classification model to train the preset multi-label classification model.

[0089] Step 204, extracting entities and relationships from the classified English corpus, and forming a knowledge graph by establishing semantic links and concept mapping.

[0090] In a possible implementation, entities and relations are extracted from classified English corpus, and a knowledge graph is formed by establishing semantic links and concept mapping, including:

[0091] Identify entities in the preliminarily classified English corpus through the preset NER tool, extract the relationship between entities through syntactic analysis and semantic role labeling; connect the extracted entities through semantic relationships to form semantic connections between entities;

[0092] The extracted entities, entity relationships, and semantic connections are stored in the form of a graph database to build the basic architecture of the knowledge graph.

[0093] Step 205, incrementally evaluate the collected English corpora of different categories through the knowledge graph to obtain evaluation results, and dynamically update the English corpora whose evaluation results meet the preset threshold into the corresponding classified English corpus data.

[0094] In a possible implementation, incremental evaluation is performed on the collected English corpora of different categories through the knowledge graph to obtain evaluation results, and the English corpora whose evaluation results meet the preset threshold are dynamically updated to the corresponding classified English corpus data, including:

[0095] The incremental evaluation of the collected English corpus of different categories through the knowledge graph includes: obtaining the information entropy change of the collected English corpus of different categories, performing incremental evaluation on the knowledge graph, and when the information entropy change of the collected English corpus of different categories meets the preset threshold, the collected English corpus is dynamically updated to the corresponding classified English corpus data, which is specifically manifested as follows:

[0096] The collected English corpus of different categories is segmented, the continuous text stream is converted into a word sequence, and the frequency of each word is counted.

[0097] Based on the calculation formula of information entropy, the information entropy of each type of newly collected English corpus before updating is obtained. Among them, H 更新前 represents the information entropy before updating, and w represents all the words in the classified corpus.

[0098] Merge the newly collected English corpus with the original English corpus to obtain a new English corpus data set to be evaluated, perform word segmentation on each type of the merged English corpus data set, re-count the frequency of occurrence of each word, and obtain the updated information entropy value according to the information entropy calculation formula Among them, H 更新后 represents the information entropy before updating, and k is a traversal of all words in the classification corpus.

[0099] Get the information entropy change value of different categories of English corpus ΔH=H 更新后 -H 更新前 When the English corpus ΔH under a certain category meets the preset threshold, the collected English corpus of the category is dynamically updated to the corresponding category English corpus data.

[0100] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0101] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.

[0102] Continue to refer Figure 5 The English learning corpus dynamic update system described in this embodiment includes:

[0103] A preliminary processing module 501 is used to distribute the English corpus data collection task to multiple nodes based on preset rules, and perform preliminary processing on the English corpus data collected by multiple nodes;

[0104] A standardized English corpus data acquisition module 502 is used to perform in-depth cleaning and standardization processing on the initially processed English corpus to acquire standardized English corpus data;

[0105] The hierarchical annotation module 503 is used to perform preliminary multi-dimensional classification on the standardized English corpus data based on a preset multi-label classification model, generate a series of labels, and obtain classification results;

[0106] The knowledge graph construction module 504 is used to extract entities and relationships from the classified English corpus and form a knowledge graph by establishing semantic links and concept mapping;

[0107] The dynamic update expansion module 505 is used to perform incremental evaluation on the collected English corpus of different categories through the knowledge graph, obtain the evaluation results, and dynamically update the English corpus whose evaluation results meet the preset threshold into the corresponding classified English corpus data.

[0108] To solve the above technical problems, the present application also provides a computer device. Figure 6 , Figure 6 This is a basic structural block diagram of the computer device in this embodiment.

[0109] The computer device 6 includes a memory 6a, a processor 6b, and a network interface 6c that are interconnected through a system bus. It should be noted that the figure only shows a computer device 6 with components 6a-6c, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (Application Specific Integrated Circuit, ASIC), programmable gate arrays (Field-Programmable Gate Array, FPGA), digital processors (Digital Signal Processor, DSP), embedded devices, etc.

[0110] The computer device may be a computing device such as a desktop computer, a notebook, a PDA, a cloud server, etc. The computer device may interact with a user through a keyboard, a mouse, a remote controller, a touch pad, or a voice control device.

[0111] The memory 6a includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 6a can be an internal storage unit of the computer device 6, such as a hard disk or memory of the computer device 6. In other embodiments, the memory 6a can also be an external storage device of the computer device 6, such as a plug-in hard disk equipped on the computer device 6, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (FlashCard), etc. Of course, the memory 6a can also include both the internal storage unit of the computer device 6 and its external storage device. In this embodiment, the memory 6a is usually used to store the operating system and various application software installed on the computer device 6, such as the program code of the dynamic update method of the English learning corpus, etc. In addition, the memory 6a can also be used to temporarily store various types of data that have been output or are to be output.

[0112] The processor 6b may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments. The processor 6b is generally used to control the overall operation of the computer device 6. In this embodiment, the processor 6b is used to run the program code stored in the memory 6a or process data, such as running the program code of the English learning corpus dynamic update method.

[0113] The network interface 6c may include a wireless network interface or a wired network interface. The network interface 6c is generally used to establish a communication connection between the computer device 6 and other electronic devices.

[0114] The present application also provides another implementation, namely, providing a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a program of a method for dynamically updating an English learning corpus, and the dynamic update of the English learning corpus can be executed by at least one processor so that the at least one processor executes the steps of the method for dynamically updating an English learning corpus as described above.

[0115] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0116] Obviously, the embodiments described above are only some embodiments of the present application, rather than all embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application is described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions recorded in the aforementioned specific implementation methods, or to perform equivalent replacement of some of the technical features therein. Any equivalent structure made using the contents of the specification and drawings of this application, directly or indirectly used in other related technical fields, is similarly within the scope of patent protection of this application.

Claims

1. A method for dynamically updating an English learning corpus, characterized in that: include: Distribute the English corpus data collection task to multiple nodes based on preset rules, and perform preliminary processing on the English corpus data collected by multiple nodes; Perform deep cleaning and standardization on the initially processed English corpus to obtain standardized English corpus data; Based on the preset multi-label classification model, the standardized English corpus data is preliminarily classified in multiple dimensions, a series of labels are generated, and the classification results are obtained; Extract entities and relations from classified English corpus and form knowledge graph by establishing semantic links and concept mapping; The collected English corpora of different categories are incrementally evaluated through the knowledge graph to obtain the evaluation results. The English corpora whose evaluation results meet the preset threshold are dynamically updated to the corresponding classified English corpus data.

2. The method for dynamically updating an English learning corpus according to claim 1, characterized in that: Based on preset rules, the English corpus data collection task is distributed to multiple nodes, and the English corpus data collected by multiple nodes is preliminarily processed, including: In the process of data collection at multiple nodes, English corpus data is extracted, preliminary data cleaning is performed on the English corpus data through preset regular expressions, and preliminary standardization is performed on the cleaned English corpus data; The initially standardized English corpus data is deduplicated to obtain the initially processed English corpus.

3. The method for dynamically updating an English learning corpus according to claim 2, characterized in that: Based on the update frequency of different nodes, set the corresponding collection time for different node data collection tasks.

4. The method for dynamically updating an English learning corpus according to claim 1, characterized in that: Before distributing the English corpus data collection task to multiple nodes based on preset rules and performing preliminary processing on the English corpus data collected by multiple nodes, it includes: Conduct statistical analysis on the quantity and usage frequency of corpora in various fields, obtain scarce fields, obtain the entropy values ​​of scarce fields, sort them from high to low according to the entropy values, and update the collection priority of English corpora.

5. The method for dynamically updating an English learning corpus according to claim 1, characterized in that: Based on the preset multi-label classification model, the standardized English corpus data is preliminarily classified in multiple dimensions, a series of labels are generated, and classification results are obtained, including: Preprocess the standardized English corpus data to obtain preprocessed data; The preprocessed English corpus data is input into the preset multi-label classification model. Based on the feature and label mapping relationship learned by the preset multi-label classification model and each dimension of the corpus, the corresponding classification labels are output respectively to form a multi-dimensional label combination as the classification result.

6. The method for dynamically updating an English learning corpus according to claim 5, characterized in that: Before the preset multi-label classification model outputs a multi-dimensional label combination, the classification confidence of each label is obtained through the Softmax activation function, that is, an estimate of the reliability of the classification result of the label.

7. The method for dynamically updating an English learning corpus according to claim 1, characterized in that: Extract entities and relations from classified English corpus and form a knowledge graph by establishing semantic links and concept mapping, including: Identify entities in the preliminarily classified English corpus through the preset NER tool, extract the relationship between entities through syntactic analysis and semantic role labeling; connect the extracted entities through semantic relationships to form semantic connections between entities; The extracted entities, entity relationships, and semantic connections are stored in the form of a graph database to build the basic architecture of the knowledge graph.

8. A system for dynamically updating an English learning corpus, used to implement the method for dynamically updating an English learning corpus of claims 1 to 7, characterized in that: include: A preliminary processing module, used to distribute the English corpus data collection task to multiple nodes based on preset rules, and perform preliminary processing on the English corpus data collected by multiple nodes; A standardized English corpus data acquisition module is used to perform in-depth cleaning and standardization on the initially processed English corpus to obtain standardized English corpus data; The hierarchical annotation module is used to perform preliminary multi-dimensional classification of the standardized English corpus data based on a preset multi-label classification model, generate a series of labels, and obtain classification results; The knowledge graph construction module is used to extract entities and relations from the classified English corpus and form a knowledge graph by establishing semantic links and concept mapping; The dynamic update expansion module is used to perform incremental evaluation on the collected English corpora of different categories through the knowledge graph, obtain the evaluation results, and dynamically update the English corpora whose evaluation results meet the preset threshold to the corresponding classified English corpus data.

9. An electronic device, characterized in that: include: one or more processors; Memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions, which, when executed by the device, enable the device to perform the steps of the method for dynamically updating an English learning corpus as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed on a computer, enables the computer to execute the steps of the method for dynamically updating an English learning corpus according to any one of claims 1 to 7.