Adaptive method and system for real-time dynamic updating of foreign language lexicon

Through network crawlers collect and filter foreign language vocabulary combination data and dynamically update the vocabulary database, the problem of fixed vocabulary combination in traditional vocabulary software is solved, and learning efficiency and depth of word memory are improved.

CN119961275AInactive Publication Date: 2025-05-09ZHENGZHOU UNIVERSITY OF AERONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510052459.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The associative vocabulary combination between each foreign language word in traditional foreign language vocabulary software is fixed, and it cannot be dynamically updated according to the actual situation of foreign language vocabulary, resulting in low learning efficiency and easy forgetting words.

Method used

The web crawler method is used to collect vocabulary combination data from multiple communication platforms in real time, filter the combined data not included in the vocabulary, filter and update according to the data frequency and welcomeness, and dynamically add it to the vocabulary.

Benefits of technology

Real-time dynamic update of foreign language vocabulary databases is realized, learning efficiency is improved, and word information is authentic and timely.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961275A_ABST
    Figure CN119961275A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of language learning, in particular to an adaptive method and system for real-time dynamic updating of a foreign language lexicon, and the method comprises the steps: collecting combined data of vocabulary combinations in a plurality of communication platforms through employing a web crawler method, screening the combined data which are not recorded in the lexicon in the combined data according to words and corresponding lexicon data, and storing the screened combined data in the lexicon; obtaining initial recording data; and according to the data frequency and the data popularity of the word or vocabulary related data in the initial recorded data, screening out the data of which the data frequency exceeds a preset frequency and / or the data popularity exceeds a preset popularity to obtain final recorded data, and adding the recorded data to a corresponding position of the word bank. When a user learns foreign languages through a terminal, latest foreign language learning information is provided for the user terminal. According to the method, the problem that association vocabularies among all foreign language words in vocabulary software are fixed in combination and cannot be dynamically updated according to actual conditions of the foreign language vocabularies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of language learning, and in particular to an adaptive method and system for real-time dynamic updating of a foreign language vocabulary. Background Art

[0002] At present, how to learn a foreign language well and improve the foreign language level is a common problem encountered by many foreign language learners. The main reason is that most students lack the method of mobilizing their brains, and they compare words and definitions and do a lot of mechanical memorization. However, this way of learning words is often inefficient, and learners do not remember the words very deeply and often forget them in a very short time.

[0003] In the prior art, most foreign language learners often memorize foreign language vocabulary by rote memorization on foreign language vocabulary APPs. Foreign language words memorized in this way are monotonous and easy to forget, and require long-term repeated review to be remembered for a long time; in addition, there are a large number of words on foreign language vocabulary software, and the associative vocabulary combinations between each foreign language word are fixed, and cannot be dynamically updated according to the actual situation of foreign language vocabulary; because vocabulary is always in dynamic development, a large number of new collocations and pronunciations appear on the Internet every day, and foreign language learners cannot learn new collocations and applications of vocabulary in time when learning the original meaning and collocation of vocabulary in traditional vocabulary software.

[0004] Therefore, the present invention provides an adaptive method and system for real-time dynamic updating of a foreign language vocabulary to solve the above problems. Summary of the invention

[0005] In view of the above situation, in order to overcome the defects of the prior art, the present invention provides an adaptive method and system for real-time dynamic updating of a foreign language vocabulary, so as to solve the problem that the associative vocabulary combinations between each foreign language word in traditional vocabulary software are fixed and cannot be dynamically updated according to the actual situation of the foreign language vocabulary.

[0006] In order to achieve the above object, the technical solution adopted by the present invention is:

[0007] In a first aspect, the present invention provides an adaptive method for real-time dynamic updating of a foreign language vocabulary, comprising: using a web crawler method to collect combination data of vocabulary combinations in multiple communication platforms in real time, and temporarily storing the combination data in a first database in chronological order, wherein the combination data includes at least two of sentence data, translation data, comment data, and pronunciation data;

[0008] Filtering the combined data that are not included in the vocabulary according to the words and the corresponding vocabulary data to obtain initial included data, and storing them in the second database;

[0009] According to the data frequency and data popularity of each data in the initial collected data, the data whose data frequency exceeds the preset frequency and / or the data popularity exceeds the preset popularity is screened out to obtain the collected data; and the sentence data, translation data and / or pronunciation data in the collected data are added to the corresponding positions of the vocabulary; when the user learns a foreign language through the terminal, the latest foreign language learning information is provided to the user terminal.

[0010] Preferably, the filtering of the combined data that are not included in the thesaurus according to the words and the corresponding thesaurus data to obtain the initial included data includes:

[0011] Search the filtered combination data corresponding to the word in the word search combination data, wherein the filtered combination data includes sentence data, translation data and / or pronunciation data; determine whether there is uncollected data in the filtered combination data according to the vocabulary data corresponding to the word, and if so, store the uncollected data corresponding to the word in a second database; and obtain initial comment data corresponding to the uncollected data, and store it in a corresponding position in the second database.

[0012] Preferably, the method of filtering out data having a data frequency exceeding a preset frequency and / or a data popularity exceeding a preset popularity according to the data frequency and the data popularity of each data in the initial collected data to obtain the collected data includes:

[0013] Count the number of repetitions of the initial collected data corresponding to the word within the preset time period to calculate the data frequency; determine the data popularity of the corresponding initial collected data according to the number of occurrences of the preset evaluation words in the preset popularity word library within the preset time period and the corresponding preset weight;

[0014] When the data frequency corresponding to the unrecorded data is greater than or equal to the first preset frequency, the unrecorded data is confirmed as the recorded data;

[0015] When the data popularity corresponding to the uncollected data is greater than or equal to the first preset popularity, the uncollected data is confirmed as collected data;

[0016] When the data frequency corresponding to the unrecorded data is greater than or equal to the second preset frequency and less than the first preset frequency, and the data popularity corresponding to the unrecorded data is greater than or equal to the second preset popularity and less than the first preset popularity, the unrecorded data is confirmed as included data.

[0017] Preferably, the calculation formula of the data frequency is:

[0018]

[0019] Among them, f is the data frequency corresponding to the unrecorded data, num(x) is the frequency statistical function corresponding to the unrecorded data, and T is the preset time length.

[0020] Preferably, the calculation formula for the data popularity is:

[0021]

[0022] Among them, w is the popularity of the data corresponding to the collected data, n is the total number of statistical times of evaluation data, and wel i (x) is the frequency statistical function of each preset evaluation word, θ i for wel i (x) The corresponding weight.

[0023] Preferably, a statistical method for calculating the data frequency of each data in the initially collected data is:

[0024] The change degree of the number of repetitions of the initial collected data corresponding to the word within the preset time interval within the preset time length is counted to obtain the data frequency; the data frequency calculation formula is:

[0025]

[0026] Among them, f1 is the data frequency corresponding to the collected data, T1 is the preset duration, t is the preset time interval, num j+t (x) is the statistical number of data not included at time j+t, num j (x) is the statistical number of data not included at time j.

[0027] Preferably, the statistical method of the frequency statistical function of the unrecorded data includes:

[0028] Perform feature extraction on the initial collected data to obtain multiple extracted information;

[0029] Input the extracted information into the frequency counting function to obtain the initial statistical times of the initial collected data;

[0030] The initial statistical times of the same initial collected data are accumulated to obtain the statistical times.

[0031] Preferably, a self-adaptive method for real-time dynamic updating of a foreign language vocabulary also includes:

[0032] After the included data is confirmed, it is reviewed; when the review is passed, the reviewed included data is added to the corresponding position of the vocabulary.

[0033] Preferably, a self-adaptive method for real-time dynamic updating of a foreign language vocabulary also includes:

[0034] Regularly detect the user activity of the communication platform; when the user activity exceeds a first preset activity, add the communication platform corresponding to the user activity to the web crawler list; when the user activity of the communication platform in the web crawler list is less than a second preset activity, remove the corresponding communication platform from the web crawler list.

[0035] In a second aspect, the present invention provides an adaptive system for real-time dynamic updating of a foreign language vocabulary, comprising: a data acquisition module, which uses a web crawler method to collect combination data of vocabulary combinations in multiple communication platforms in real time, and temporarily stores the combination data in a first database in chronological order, wherein the combination data includes at least two of sentence data, translation data, comment data, and pronunciation data;

[0036] A data screening module, screening the combined data that are not included in the thesaurus according to the words and the corresponding thesaurus data, so as to obtain initial included data, and store it in the second database;

[0037] The data collection module selects data with a frequency exceeding a preset frequency and / or a popularity exceeding a preset popularity according to the data frequency and the data popularity of each data in the initial collected data, so as to obtain the collected data; and adds the sentence data, translation data and / or pronunciation data in the collected data to the corresponding position of the vocabulary;

[0038] The data upload module provides the latest foreign language learning information to the user terminal when the user learns a foreign language through the terminal.

[0039] The beneficial effects of the present invention are:

[0040] 1. The present invention collects the combined data of vocabulary combinations in multiple communication platforms by adopting a web crawler method, and screens the combined data that are not included in the vocabulary according to the words and the corresponding vocabulary data to obtain the initial included data; then, according to the data frequency and data popularity of the words or vocabulary-related data in the initial included data, screens out the data whose data frequency exceeds the preset frequency and / or whose data popularity exceeds the preset popularity, obtains the final included data, and adds the included data to the corresponding position of the vocabulary; when the user learns a foreign language through the terminal, the latest foreign language learning information is provided to the user terminal. The present invention adopts the above method to solve the problem that the associative vocabulary combination between each foreign language word in the traditional vocabulary software is fixed and cannot be dynamically updated according to the actual situation of the foreign language vocabulary.

[0041] 2. When confirming the data frequency or data popularity, the present invention uses a frequency statistics function to perform feature extraction on the initial collected data to obtain multiple extraction information of the same word or the same vocabulary; then the extraction information is input into the frequency statistics function to obtain the initial statistical times of the initial collected data; then the initial statistical times of the same initial collected data are accumulated to obtain the statistical times; finally, the corresponding result is obtained according to the calculation formula of the data frequency or the calculation formula of the data popularity. The present invention uses the frequency statistics function to perform feature extraction on the uncollected data of the same word or the same vocabulary, and then performs frequency statistics, which can ensure the accuracy of the statistical results, avoid statistical omissions, and ensure the authenticity of the word information, vocabulary information or sentence information added to the vocabulary.

[0042] 3. When confirming the web crawler list, the present invention will regularly detect the user activity of the communication platform; the communication platform of the web crawler list will be regularly replaced according to the user activity, so that the data collected by the web crawler method is widely known. After analyzing the collected data and determining the included data, the authenticity of the word information, vocabulary information or sentence information added to the vocabulary is guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 A schematic flow chart of an adaptive method for real-time dynamic updating of a foreign language vocabulary according to the present invention;

[0044] Figure 2 The module diagram of an adaptive system for real-time dynamic updating of a foreign language vocabulary according to the present invention is shown in FIG. DETAILED DESCRIPTION

[0045] The following will refer to the attached Figure 1 To Attachment Figure 2 The embodiments of the present invention are described in detail. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.

[0046] The present invention provides an adaptive method for real-time dynamic updating of a foreign language vocabulary. Figure 1 As shown, the following steps are included:

[0047] Step S11: using a web crawler method to collect combination data of vocabulary combinations in multiple communication platforms in real time, and temporarily storing the combination data in a first database in chronological order.

[0048] Specifically, the combined data includes at least two of sentence data, translation data, comment data and pronunciation data, for example, sentence data and comment data; or sentence data and translation data; or sentence data and pronunciation data; or comment data and pronunciation data.

[0049] When collecting data in real time, the collected combined data are stored in the corresponding positions of the first database in chronological order, and when the data stored in the first database are processed and analyzed by steps S12 and S13 to obtain the included data, the relevant combined data collected in the first database are deleted to improve the storage capacity and data storage efficiency of the first database.

[0050] Step S12: Filter the data in the combined data that are not included in the vocabulary according to the words and the corresponding vocabulary data to obtain initial included data, and store them in the second database.

[0051] Through the screening process of step S12, the amount of data is reduced and key data is screened out, which helps to improve the processing speed of confirming the included data in step S13.

[0052] Step S13: based on the data frequency and data popularity of each data in the initial collected data, filter out the data whose data frequency exceeds the preset frequency and / or whose data popularity exceeds the preset popularity to obtain the collected data; and add the sentence data, translation data and / or pronunciation data in the collected data to the corresponding position of the vocabulary.

[0053] Preferably, the pronunciation data is in the form of, but not limited to, audio data and video data.

[0054] Specifically, when filtering qualified data based on the data frequency and data popularity of each data in the initial collected data, it means determining the data frequency and data popularity by counting the repetition of the relevant data within a set time based on the relevant data of each word or vocabulary in the initial collected data.

[0055] Step S14: When the user learns a foreign language through the terminal, the latest foreign language learning information is provided to the user terminal.

[0056] Specifically, the present invention collects the combined data of vocabulary combinations in multiple communication platforms by adopting a web crawler method, and screens the combined data that are not included in the vocabulary according to the words and the corresponding vocabulary data to obtain the initial included data; then, according to the data frequency and data popularity of the words or vocabulary-related data in the initial included data, screens out the data whose data frequency exceeds the preset frequency and / or whose data popularity exceeds the preset popularity, obtains the final included data, and adds the included data to the corresponding position of the vocabulary; when the user learns a foreign language through the terminal, the latest foreign language learning information is provided to the user terminal. The present invention adopts the above method to solve the problem that the associative vocabulary combination between each foreign language word in the traditional vocabulary software is fixed and cannot be dynamically updated according to the actual situation of the foreign language vocabulary.

[0057] In one embodiment of the present invention, the combination data that is not included in the thesaurus is screened according to the words and the corresponding thesaurus data to obtain the initial included data, including:

[0058] Search the filtered combination data corresponding to the word in the word search combination data, wherein the filtered combination data includes sentence data, translation data and / or pronunciation data; determine whether there is uncollected data in the filtered combination data according to the vocabulary data corresponding to the word, and if so, store the uncollected data corresponding to the word in a second database; and obtain initial comment data corresponding to the uncollected data, and store it in a corresponding position in the second database.

[0059] In this embodiment, a natural language processing model is used to analyze and identify sentence data and translation data, and then the identified data is screened to obtain the screening combination data corresponding to the word. The natural language processing model can be a GPT model, a model built based on a recurrent neural network, a word embedding model, etc. The method of data screening can be database screening, rule-based screening, or attribute-based screening. Through the setting method of this embodiment, the present invention can screen out the uncollected data corresponding to the word and store it in the second database, so that the uncollected data can be analyzed and processed later to enrich the data volume of the vocabulary.

[0060] In one embodiment of the present invention, based on the data frequency and data popularity in the initial collected data, data with a data frequency exceeding a preset frequency and / or a data popularity exceeding a preset popularity is screened out to obtain the collected data, including:

[0061] The number of repetitions of the initial collected data corresponding to the word within the preset time period is counted to calculate the data frequency; the data popularity of the corresponding initial collected data is determined according to the number of occurrences of the preset evaluation words in the preset popularity vocabulary within the preset time period and the corresponding preset weights.

[0062] When the data frequency corresponding to the unincluded data is greater than or equal to the first preset frequency, the unincluded data is confirmed as included data; when the data popularity corresponding to the unincluded data is greater than or equal to the first preset popularity, the unincluded data is confirmed as included data; when the data frequency corresponding to the unincluded data is greater than or equal to the second preset frequency and less than the first preset frequency, and the data popularity corresponding to the unincluded data is greater than or equal to the second preset popularity and less than the first preset popularity, the unincluded data is confirmed as included data.

[0063] Preferably, the first preset frequency is 5000 times / second, and the second preset frequency is 3000 times / second; the first preset popularity is 70%, and the second preset popularity is 40%; the preset frequency and preset popularity can be adjusted according to actual conditions to obtain data that meets user needs.

[0064] Furthermore, a calculation formula for the above data frequency is:

[0065]

[0066] Among them, f is the data frequency corresponding to the unrecorded data, num(x) is the frequency statistical function corresponding to the unrecorded data, and T is the preset time length.

[0067] Specifically, the number of times a certain item of initial collected data corresponding to a word repeatedly appears within a preset time period of a preset duration T is counted to calculate the data frequency of the initial collected data.

[0068] Furthermore, the calculation formula for data popularity is:

[0069]

[0070] Among them, w is the popularity of the data corresponding to the collected data, n is the total number of statistics of each preset evaluation word, and wel i (x) is the frequency statistical function of the preset evaluation word, θ i for wel i (x) The corresponding weight.

[0071] Preferably, the preset evaluation words include but are not limited to words or emoticons representing evaluation such as good, excellent, great, and poor; the weights corresponding to the preset evaluation words are fixed values, for example, excellent is 20%; the weight values ​​can be determined based on actual conditions so that the calculated data popularity fits the actual conditions.

[0072] Furthermore, another statistical method for calculating the data frequency of each data in the initial collected data is:

[0073] The change degree of the number of repetitions of the initial collected data corresponding to the word within the preset time interval within the preset time length is counted to obtain the data frequency; the data frequency calculation formula is:

[0074]

[0075] Among them, f1 is the data frequency corresponding to the collected data, T1 is the preset duration, t is the preset time interval, num j+t (x) is the statistical number of data not included at time j+t, num j (x) is the statistical number of data not included at time j.

[0076] Through the configuration of this embodiment, the present invention can select a suitable data frequency calculation formula to calculate the data frequency to obtain a data frequency that meets the actual situation.

[0077] In one embodiment of the present invention, the statistical method of the frequency statistical function of the unrecorded data includes the following steps:

[0078] Step S21: extracting features from the initial collected data to obtain a plurality of extracted information.

[0079] Step S22: input the extracted information into a frequency counting function to obtain an initial statistical frequency of the initial collected data.

[0080] Step S23: Accumulate the initial statistical times of the same initial collected data to obtain the statistical times.

[0081] In this embodiment, when confirming the data frequency or the data popularity, the present invention uses a frequency statistics function to perform feature extraction on the initial collected data to obtain multiple extraction information of the same word or the same vocabulary; then the extraction information is input into the frequency statistics function to obtain the initial statistical times of the initial collected data; then the initial statistical times of the same initial collected data are accumulated to obtain the statistical times; finally, the corresponding result is obtained according to the calculation formula of the data frequency or the calculation formula of the data popularity. The present invention uses the frequency statistics function to perform feature extraction on the uncollected data of the same word or the same vocabulary, and then performs frequency statistics, which can ensure the accuracy of the statistical results, avoid statistical omissions, and ensure the authenticity of the word information, vocabulary information or sentence information added to the vocabulary.

[0082] In one embodiment of the present invention, an adaptive method for real-time dynamic updating of a foreign language vocabulary also includes: after storing the data in a first database, preprocessing the combined data in the first database, wherein the preprocessing includes at least data standardization, outlier processing, and noise reduction processing.

[0083] In this embodiment, the combined data in the first database are preprocessed to unify the dimensions and data formats of each combined data and perform noise reduction processing on the pronunciation data, so as to facilitate subsequent analysis and screening of the combined data.

[0084] In one embodiment of the present invention, an adaptive method for real-time dynamic updating of a foreign language vocabulary also includes: after confirming the included data, reviewing the included data; after passing the review, adding the reviewed included data to the corresponding position of the vocabulary.

[0085] In this embodiment, a manual review mode is added. By manually reviewing the collected data, the translation quality, data popularity, and pronunciation quality of the collected data are ensured, so that the collected data added to the vocabulary meets the needs of users. When manually reviewing the collected data, relevant information will be automatically provided, including but not limited to videos, audios, number of comments, screenshots, etc. of the communication platform to facilitate the work of the reviewer. Through the setting method of this embodiment, the present invention can ensure the quality of the collected data added to the vocabulary and provide users with good learning services.

[0086] In one embodiment of the present invention, an adaptive method for real-time dynamic updating of a foreign language vocabulary also includes: regularly detecting the user activity of a communication platform; when the user activity exceeds a first preset activity, adding the communication platform corresponding to the user activity to a web crawler list; when the user activity of a communication platform in the web crawler list is less than a second preset activity, removing the corresponding communication platform from the web crawler list.

[0087] In this embodiment, when confirming the web crawler list, the present invention will regularly detect the user activity of the communication platform; the communication platform of the web crawler list is regularly replaced according to the user activity, so that the data collected by the web crawler method is widely known, and after the collected data is analyzed and the included data is determined, the authenticity of the word information, vocabulary information or sentence information added to the vocabulary is guaranteed.

[0088] The present invention also provides an adaptive system for real-time dynamic updating of foreign language word banks, as shown in the attached Figure 2 As shown, it mainly includes data acquisition module, data screening module, data collection module, data upload module, data storage module and central control module; the responsibilities of each module are as follows:

[0089] The data collection module uses a web crawler method to collect combination data of vocabulary combinations in multiple communication platforms in real time, and temporarily stores the combination data in a first database in chronological order, wherein the combination data includes at least two of sentence data, translation data, comment data and pronunciation data.

[0090] The data screening module screens the combined data that are not included in the vocabulary according to the words and the corresponding vocabulary data to obtain initial included data and store them in the second database.

[0091] The data collection module selects data with a frequency exceeding a preset frequency and / or a popularity exceeding a preset popularity according to the data frequency and the data popularity of each data in the initial collected data to obtain the collected data; and adds the sentence data, translation data and / or pronunciation data in the collected data to the corresponding position of the vocabulary.

[0092] The data upload module provides the latest foreign language learning information to the user terminal when the user learns a foreign language through the terminal.

[0093] The data storage module is used to store the data generated during the execution of the data acquisition module, the data screening module, and the data collection module.

[0094] The central control module is connected to each module in communication or electrical connection to control each module to perform corresponding operations through control instructions.

[0095] The present invention realizes the collection of combined data of vocabulary combinations in multiple communication platforms by the mutual cooperation between the above modules, and screens the combined data that are not included in the vocabulary according to the words and the corresponding vocabulary data to obtain the initial included data; and then screens the data frequency and data popularity of the words or vocabulary-related data in the initial included data to obtain the final included data, and adds the included data to the corresponding position of the vocabulary. When the user learns a foreign language through the terminal, the latest foreign language learning information is provided to the user terminal. The present invention adopts the above method to solve the problem that the associative vocabulary combination between each foreign language word in the traditional vocabulary software is fixed and cannot be dynamically updated according to the actual situation of the foreign language vocabulary.

[0096] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0097] It should be noted that in the description of the present invention, the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer" and the like indicating directions or positional relationships are based on the directions or positional relationships shown in the drawings, which are only for the convenience of description, and do not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.

[0098] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0099] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0100] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0101] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0102] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0103] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0104] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

Claims

1. An adaptive method for real-time dynamic updating of a foreign language vocabulary, characterized in that: include: Using a web crawler method to collect combination data of vocabulary combinations in multiple communication platforms in real time, and temporarily storing the combination data in a first database in chronological order, the combination data including at least two of sentence data, translation data, comment data and pronunciation data; Filtering the combined data that are not included in the vocabulary according to the words and the corresponding vocabulary data to obtain initial included data, and storing them in the second database; According to the data frequency and data popularity of each data in the initial collected data, the data with a data frequency exceeding a preset frequency and / or a data popularity exceeding a preset popularity is screened out to obtain the collected data; and adding the sentence data, translation data and / or pronunciation data in the collected data to the corresponding positions in the vocabulary; When a user learns a foreign language through a terminal, the latest foreign language learning information is provided to the user terminal.

2. The adaptive method according to claim 1, characterized in that: The step of screening the combined data that are not included in the vocabulary according to the words and the corresponding vocabulary data to obtain the initial included data includes: Search the filtered combination data corresponding to the word in the word search combination data, wherein the filtered combination data includes sentence data, translation data and / or pronunciation data; determine whether there is uncollected data in the filtered combination data according to the vocabulary data corresponding to the word, and if so, store the uncollected data corresponding to the word in a second database; and obtain initial comment data corresponding to the uncollected data, and store it in a corresponding position in the second database.

3. The adaptive method according to claim 2, characterized in that: The method of screening out data having a data frequency exceeding a preset frequency and / or a data popularity exceeding a preset popularity according to the data frequency and the data popularity of each data in the initial collected data to obtain the collected data includes: Count the number of repetitions of the initial collected data corresponding to the word within the preset time period to calculate the data frequency; determine the data popularity of the corresponding initial collected data according to the number of occurrences of the preset evaluation words in the preset popularity word library within the preset time period and the corresponding preset weight; When the data frequency corresponding to the unrecorded data is greater than or equal to the first preset frequency, the unrecorded data is confirmed as the recorded data; When the data popularity corresponding to the uncollected data is greater than or equal to the first preset popularity, the uncollected data is confirmed as collected data; When the data frequency corresponding to the unrecorded data is greater than or equal to the second preset frequency and less than the first preset frequency, and the data popularity corresponding to the unrecorded data is greater than or equal to the second preset popularity and less than the first preset popularity, the unrecorded data is confirmed as included data.

4. The adaptive method according to claim 3, characterized in that: The calculation formula of the data frequency is: Among them, f is the data frequency corresponding to the unrecorded data, num(x) is the frequency statistical function corresponding to the unrecorded data, and T is the preset time length.

5. The adaptive method according to claim 3, characterized in that: The calculation formula of the data popularity is: Among them, w is the popularity of the data corresponding to the collected data, n is the total number of statistical times of evaluation data, and wel i (x) is the frequency statistical function of each preset evaluation word, θ i for wel i (x) The corresponding weight.

6. The adaptive method according to claim 2, characterized in that: A statistical method for the data frequency of each data in the initial collected data is: The change degree of the number of repetitions of the initial collected data corresponding to the word within the preset time interval within the preset time length is counted to obtain the data frequency; the data frequency calculation formula is: Among them, f1 is the data frequency corresponding to the collected data, T1 is the preset duration, t is the preset time interval, num j+t (x) is the statistical number of data not included at time j+t, num j (x) is the statistical number of data not included at time j.

7. The adaptive method according to claim 4, characterized in that: The statistical method of the frequency statistical function of the uncollected data includes: Perform feature extraction on the initial collected data to obtain multiple extracted information; Input the extracted information into the frequency counting function to obtain the initial statistical times of the initial collected data; The initial statistical times of the same initial collected data are accumulated to obtain the statistical times.

8. The adaptive method according to claim 1, characterized in that: Also includes: After confirming the included data, review the included data; When the review is passed, the reviewed data will be added to the corresponding position of the vocabulary.

9. The adaptive method according to claim 1, characterized in that: Also includes: Regularly check the user activity of the communication platform; When the user activity exceeds a first preset activity, adding the communication platform corresponding to the user activity to the web crawler list; When the user activity of a communication platform in the web crawler list is less than a second preset activity, the corresponding communication platform is removed from the web crawler list.

10. An adaptive system for real-time dynamic updating of foreign language word banks, characterized in that: include: A data collection module, which uses a web crawler method to collect combination data of vocabulary combinations in multiple communication platforms in real time, and temporarily stores the combination data in a first database in chronological order, wherein the combination data includes at least two of sentence data, translation data, comment data, and pronunciation data; A data screening module, screening the combined data that are not included in the thesaurus according to the words and the corresponding thesaurus data, so as to obtain initial included data, and store it in the second database; The data collection module selects data with a frequency exceeding a preset frequency and / or a popularity exceeding a preset popularity according to the data frequency and the popularity of each data in the initial collected data, so as to obtain the collected data; and adding the sentence data, translation data and / or pronunciation data in the collected data to the corresponding positions in the vocabulary; The data upload module provides the latest foreign language learning information to the user terminal when the user learns a foreign language through the terminal.

Citation Information

Patent Citations

  • User behavior analysis-based dynamic word bank updating method

    CN111125299A

  • Foreign language association lexicon self-training method based on network information

    CN117370307A