A vocabulary processing method, apparatus, device, and storage medium
By analyzing the frequency and probability of target words in corpus data, the most frequently used words are identified and provided, which solves the problem of low vocabulary learning efficiency in existing technologies and achieves more efficient vocabulary learning results.
Patent Information
- Application Number
- CN202111020267.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-01
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2041-09-01
AI Technical Summary
Current technology cannot provide high-frequency word information for words in a vocabulary set, resulting in low efficiency for language learners in vocabulary learning.
By identifying the target vocabulary corpus data, statistically analyzing the frequency and probability of vocabulary usage in different word sets, and selecting the word set with the highest expected value as the highest frequency word set, we can provide it to language learners to improve their learning efficiency.
By identifying and providing the most frequent word lists, language learners can more accurately grasp the actual usage of words at specific learning stages, thereby improving learning efficiency.
Smart Images

Figure CN115730594B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to, but is not limited to, the field of vocabulary information processing technology, and in particular to a vocabulary processing method, apparatus, device, and storage medium. Background Technology
[0002] In the context of economic globalization, we inevitably encounter more situations requiring the use of foreign languages in our work, studies, and daily lives. Therefore, more and more people need or want to learn one or more languages besides their mother tongue. In particular, the number of learners of languages such as English, French, German, Spanish, and even Italian is increasing. For these language learners, vocabulary acquisition is the most fundamental and crucial step in mastering a language.
[0003] In language learning, learners first need to memorize vocabulary. In particular, language learning is a gradual process, which means dividing the learning process into multiple stages. Correspondingly, vocabulary needs to be categorized into different sets, from simple to difficult, based on factors such as complexity and frequency of occurrence. For example, for Chinese English learners, there are different level exams in university, such as CET-4, CET-6, TEM-4, and TEM-8, corresponding to different learning stages and, more specifically, different vocabulary sets.
[0004] Currently, there are two main approaches to vocabulary learning. In the first approach, when language learners encounter unfamiliar words, phrases, sentences, or paragraphs while reading articles, watching videos, listening to audio, or in daily life, they can input these words into online dictionaries or dictionary applications to look them up and learn. In this case, the system typically searches for and displays all possible definitions, even including pronunciation, part of speech, collocations, and example sentences.
[0005] In the second approach, language learners can engage in targeted vocabulary memorization. Previously, learners used textbook vocabulary lists or vocabulary books. With the widespread use of mobile electronic devices such as mobile phones, more and more applications have been developed for vocabulary learning. The most common method used by these applications is the "high-frequency vocabulary" approach. In this approach, for a given learning stage, the most frequently occurring words in that stage are compiled into a vocabulary set for language learners. Furthermore, some applications compile a vocabulary set based on the probability of a word appearing in various lexical categories within a corpus of data. A lexical category refers to the possible collocations of a word in actual use. However, the applications here simply use these probabilities to determine high-frequency vocabulary. The content displayed to language learners by these applications can also include pronunciation, part of speech, definition, fixed collocations, example sentences, etc. In some applications, to improve learning speed, only the definition is displayed to language learners.
[0006] Therefore, it is evident that the existing technology has a problem in that it cannot provide high-frequency word settings for words in a vocabulary set. Summary of the Invention
[0007] This disclosure provides a vocabulary processing method, apparatus, device, and storage medium to provide the most frequent words in a vocabulary set, enabling users to master vocabulary and thereby improve their learning efficiency.
[0008] According to a first aspect of the present disclosure, a vocabulary processing method is provided. The vocabulary processing method includes: determining a target vocabulary, wherein the target vocabulary includes at least one word; obtaining corpus data, wherein the corpus data includes at least a plurality of sentences, and at least one of the plurality of sentences contains the target vocabulary; based on the corpus data and the target vocabulary, determining lexical statistics data and vocabulary statistics data of a first word contained in the target vocabulary as a first lexical setting; determining expected information for the first word as a first lexical setting based on the lexical statistics data and the vocabulary statistics data; and determining a first characteristic of the first word based on the expected information, wherein the first characteristic is at least the lexical setting of the first word with the greatest expected value.
[0009] In one embodiment, determining the expected information of a first word as a first word setting based on word setting statistics and vocabulary statistics includes: determining a first expected value based on word setting statistics and vocabulary statistics, wherein the first expected value is the expected number of times the first word is used as a first word setting in the target vocabulary; determining a second expected value based on word setting statistics and vocabulary statistics, wherein the second expected value is the expected number of times the first word is used as a first word setting in the corpus data; and determining the expected number of times the first word is used as a first word setting based on the first expected value and the second expected value.
[0010] In one embodiment, the lexical setting statistics include the first number of times the first word is used as the first lexical setting in the target vocabulary and the second number of times the first word is used as the first lexical setting in the corpus data. The lexical statistics include the first number of times the first word appears in the target vocabulary and the second number of times the first word appears in the corpus data. Based on the lexical setting statistics and the lexical statistics, the expected information of the first word being used as the first lexical setting is determined, including: determining a first expected value based on the first number of times of use, the second number of times of use, the first number of times of appearance, and the second number of times of appearance; determining a second expected value based on the second number of times of use, the second number of times of use, the first number of times of appearance, and the second number of times of appearance; and determining the expected value of the number of times the first word is used as the first lexical setting based on the first expected value and the second expected value.
[0011] In one embodiment, determining a first characteristic of a first word based on expected information includes: determining whether the expected number of times the first word is used as a first word setting is greater than the expected number of times the first word is used as a second word setting, wherein the second word setting is different from the first word setting; and when it is determined that the expected number of times the first word is used as a first word setting is greater than the expected number of times the first word is used as a second word setting, setting the first word setting as the first characteristic of the first word.
[0012] In one embodiment, the above-described vocabulary processing method further includes: determining the probability value of a first word being used as a first word in the corpus data; and determining the average value of the first word being used as a first word based on the corpus data and the target vocabulary. Accordingly, determining the first characteristic of the first word according to the expected information includes: determining the first characteristic of the first word based on the expected information, the probability value of the word, and the average value of the word.
[0013] In one embodiment, determining the average value of the word setting when the first word is used as a first word setting based on corpus data and target vocabulary includes: determining the word setting statistics when the first word is used as a second word setting based on corpus data and target vocabulary; and determining the average value of the word setting when the first word is used as a first word setting based on the word setting statistics when the first word is used as a first word setting and the word setting statistics when the first word is used as a second word setting.
[0014] In one embodiment, determining a first characteristic of a first word based on expected information, word probability value, and word average value includes: determining whether the expected number of times the first word is used as a first word setting is greater than the expected number of times the first word is used as a second word setting; determining whether the word probability value of the first word is greater than the word average value; and when it is determined that the expected number of times the first word is used as a first word setting is greater than the expected number of times the first word is used as a second word setting and the word probability value of the first word is greater than the word average value, setting the first word setting as the first characteristic of the first word.
[0015] According to a second aspect of the present disclosure, a vocabulary processing apparatus is provided. The vocabulary processing apparatus includes a target vocabulary determination module, a corpus data acquisition module, a statistical data determination module, an expectation information determination module, and a feature determination module. The target vocabulary determination module is configured to determine a target vocabulary, wherein the target vocabulary includes at least one word. The corpus data acquisition module is configured to acquire corpus data, wherein the corpus data includes at least a plurality of statements, and at least one of the plurality of statements contains the target vocabulary. The statistical data determination module is configured to determine, based on the corpus data and the target vocabulary, a lexical setting statistical data and a statistical data of the first word contained in the target vocabulary as a first lexical setting. The expectation information determination module is configured to determine expectation information for the first word as a first lexical setting based on the lexical setting statistical data and the vocabulary statistical data. The feature determination module is configured to determine a first feature of the first word based on the expectation information, wherein the first feature is at least a lexical setting of the first word with the maximum expectation.
[0016] In one embodiment, the expectation information determination module includes a first expectation value determination submodule, a second expectation value determination submodule, and an expectation value determination submodule. The first expectation value determination submodule is configured to determine a first expectation value based on word attribute statistics and vocabulary statistics, wherein the first expectation value is the expected number of times the first word is used as a first word attribute in the target vocabulary. The second expectation value determination submodule is configured to determine a second expectation value based on word attribute statistics and vocabulary statistics, wherein the second expectation value is the expected number of times the first word is used as a first word attribute in the corpus data. The expectation information determination submodule is configured to determine the expected number of times the first word is used as a first word attribute based on the first expectation value and the second expectation value.
[0017] In one embodiment, the lexical statistics include a first usage count of the first word as a first lexical term in the target vocabulary and a second usage count of the first word as a first lexical term in the corpus data. The lexical statistics also include a first occurrence count of the first word in the target vocabulary and a second occurrence count of the first word in the corpus data. A first expected value determination submodule is configured to determine a first expected value based on the first usage count, the second usage count, the first occurrence count, and the second occurrence count. A second expected value determination submodule is configured to determine a second expected value based on the first usage count, the second usage count, the first occurrence count, and the second occurrence count. An expected information determination submodule is configured to determine the expected value of the number of times the first word is used as a first lexical term based on the first expected value and the second expected value.
[0018] In one embodiment, the feature determination module includes an expected value comparison submodule and a feature determination submodule. The expected value comparison submodule is configured to determine whether the expected number of times the first word is used as a first terminology is greater than the expected number of times the first word is used as a second terminology, wherein the second terminology is different from the first terminology. The feature determination submodule is configured to set the first terminology as a first feature of the first wordinology when it is determined that the expected number of times the first word is used as a first terminology is greater than the expected number of times the first word is used as a second terminology.
[0019] In one embodiment, the vocabulary processing apparatus further includes a probability determination module and an average value determination module. The probability determination module is configured to determine a probability value of the first vocabulary being used as a first lexical term in the corpus data. The average value determination module is configured to determine an average value of the first vocabulary being used as a first lexical term based on the corpus data and the target vocabulary. Accordingly, the feature determination module is configured to determine a first feature of the first vocabulary based on expected information, the probability value, and the average value.
[0020] In one embodiment, the average value determination module includes a statistical data determination submodule and a word setting average value determination submodule. The statistical data determination submodule is configured to determine word setting statistical data for the first word as a second word setting based on corpus data and target vocabulary. The word setting average value determination submodule is configured to determine the word setting average value for the first word as a first word setting based on the word setting statistical data for the first word setting as a first word setting and the word setting statistical data for the second word setting.
[0021] In one embodiment, the feature determination module is specifically configured to: determine whether the expected number of times the first word is used as a first word setting is greater than the expected number of times the first word is used as a second word setting; determine whether the word setting probability value of the first word is greater than the word setting average value; and when it is determined that the expected number of times the first word is used as a first word setting is greater than the expected number of times the first word is used as a second word setting and the word setting probability value of the first word is greater than the word setting average value, set the first word setting as the first feature of the first word.
[0022] According to a third aspect of the present disclosure, a vocabulary processing apparatus is provided. The vocabulary processing apparatus includes a processor and a memory for storing processor-executable instructions. The processor is configured to implement the vocabulary processing method described in the first aspect when executing the executable instructions.
[0023] According to a fourth aspect of the present disclosure, a non-volatile computer-readable storage medium is provided. The storage medium stores an executable program. When executed by a processor, the executable program implements the vocabulary processing method described in the first aspect.
[0024] The technical solutions provided in this disclosure may include the following beneficial effects.
[0025] In this disclosure, the expected information of the first word as the first word setting is determined based on the word setting statistics and the vocabulary statistics of the first word. Based on this, the first characteristic of the first word is determined, that is, at least the word setting with the greatest expectation. This provides the highest frequency word setting of the words in the vocabulary set, so that users can master the actual usage of the words at a specific learning stage, thereby improving the user's learning efficiency.
[0026] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0027] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0028] Figure 1 A flowchart of a vocabulary processing method provided in an exemplary embodiment is shown;
[0029] Figure 2 It shows Figure 1 A detailed flowchart of the vocabulary processing methods in the document;
[0030] Figure 3 A structural block diagram of a vocabulary processing method provided by another exemplary embodiment is shown;
[0031] Figure 4 It shows Figure 3 A detailed flowchart of the vocabulary processing methods in the document;
[0032] Figure 5 An application example of the vocabulary processing method according to the present invention is shown;
[0033] Figure 6 Another application example of the vocabulary processing method according to the present invention is shown;
[0034] Figure 7 This illustrates yet another application example of the vocabulary processing method according to the present invention;
[0035] Figure 8 A structural block diagram of a vocabulary processing apparatus provided in an exemplary embodiment is shown;
[0036] Figure 9 A structural block diagram of a vocabulary processing apparatus provided in another exemplary embodiment is shown;
[0037] Figure 10 A structural block diagram of a vocabulary processing device provided in an exemplary embodiment is shown. Detailed Implementation
[0038] The embodiments of this application are described below with reference to the accompanying drawings. In the following description, reference is made to the accompanying drawings, which form part of this application and illustrate specific aspects of the embodiments of this application or to which specific aspects of the embodiments of this application may be used. It should be understood that the embodiments of this application may be used in other aspects and may include structural or logical variations not depicted in the drawings. Therefore, the following detailed description should not be construed in a limiting sense, and the scope of this application is defined by the appended claims. For example, it should be understood that the disclosure of the described methods is equally applicable to corresponding devices or systems for performing the methods, and vice versa. For example, if one or more specific method steps are described, the corresponding device may include one or more units, such as functional units, to perform the described one or more method steps (e.g., one unit performs one or more steps, or multiple units, each performing one or more of multiple steps), even if such one or more units are not explicitly described or illustrated in the drawings. On the other hand, for example, if a specific apparatus is described based on one or more units such as functional units, the corresponding method may include a step to perform the functionality of one or more units (e.g., a step to perform the functionality of one or more units, or multiple steps, each of which performs the functionality of one or more units among a plurality of units), even if such one or more steps are not explicitly described or illustrated in the accompanying drawings. Furthermore, it should be understood that, unless otherwise expressly stated, features of the various exemplary embodiments and / or aspects described herein can be combined with each other.
[0039] Figure 1 A flowchart of a vocabulary processing method provided by an exemplary embodiment is shown. Figure 2 It shows Figure 1 A detailed flowchart of the vocabulary processing methods in [the document / system].
[0040] like Figure 1 As shown, the vocabulary processing method includes steps S100, S200, S300, S400 and S500.
[0041] Step S100: Identify target vocabulary.
[0042] Here, the target vocabulary can be one or more words entered by the user when searching for words, or it can be a specific word in the corresponding vocabulary set when using the application to memorize words.
[0043] Step S200: Obtain corpus data, wherein the corpus data includes at least multiple sentences, and at least one of the multiple sentences contains the target vocabulary.
[0044] For each learning stage, there is not only a vocabulary set corresponding to that learning stage, but also multiple sentences corresponding to that vocabulary set. In one embodiment, the corpus data may include these multiple sentences. In another embodiment, the corpus data may include a vocabulary set for a specific learning stage and the corresponding multiple sentences. In yet another embodiment, the corpus data may also include annotations for the words in the vocabulary set for each sentence. The corpus data typically exists in the form of a corpus and includes a large number of sentences. In one embodiment, the target word appears in at least one sentence in the corpus data.
[0045] Step S300: Based on the corpus data and target vocabulary, determine the vocabulary statistics of the first word contained in the target vocabulary as the first word setting and the vocabulary statistics of the first word.
[0046] The term "lexical setting" in this disclosure can be understood as a lexical configuration for a word, which can be used to represent the collocations used for that word, or combinations of parts of speech and collocations. For example, the parts of speech of a word can include, but are not limited to, the following: noun, pronoun, verb, preposition, adjective, adverb, article, conjunction, interjection, numeral, and article. As another example, the parts of speech of a word can be set using the parts of speech settings in NLTK (Natural Language Toolkit), which provides a richer selection of parts of speech. Collocations can include, but are not limited to, the following: noun + noun, adjective + noun, verb + noun, verb + infinitive, adverb + verb, noun + verb, and verb + clause.
[0047] Each word in the vocabulary set has one or more word settings. For example, for a specific word that has both noun and verb parts of speech, when it is a verb, the word setting can be "verb + noun", "verb + clause", "verb + preposition", etc., and when it is a noun, the word setting can be "preposition + noun", "noun + noun", etc.
[0048] The first vocabulary is a word contained within the target vocabulary. The first vocabulary may appear once or multiple times in the target vocabulary. In one embodiment, the vocabulary statistics include the first number of times the first vocabulary is used as the first vocabulary in the target vocabulary and the second number of times the first vocabulary is used as the first vocabulary in the corpus data. The vocabulary statistics also include the first number of times the first vocabulary appears in the target vocabulary and the second number of times the first vocabulary appears in the corpus data.
[0049] In one embodiment, the target vocabulary can be a single word, phrase, sentence, paragraph, or multiple unrelated words.
[0050] In one example, the relationship between the first number of uses and the first number of occurrences can be at least one of the following:
[0051] When the target vocabulary includes a single word, both the first usage count and the first occurrence count are 1;
[0052] When a word appears multiple times in the target vocabulary, if the word appears as the same word, then the first usage count and the first occurrence count are equal;
[0053] If the word appears in different word forms, the first usage count is less than the first occurrence count.
[0054] Step S400: Based on the lexical statistics and vocabulary statistics, determine the expected information for the first word to be used as the first lexical setting.
[0055] Here, the expected information for the use of the first word as the first word setting includes: the expected number of times the first word is used as the first word setting in the target vocabulary, i.e., the first expected value; and the expected number of times the first word is used as the first word setting in the corpus data, i.e., the second expected value.
[0056] Specifically, if we take the union of the target vocabulary and the corpus data as the set, then the probability of the first word appearing as the first word in this set is represented by the ratio between the number of times the first word appears as the first word (i.e., the sum of the first and second usages) and the total number of times the first word appears (i.e., the sum of the first and second occurrences). From another perspective, this probability value indicates the probability of each first word in the set using the first word. Therefore, the first expected value is the product of this probability value and the first usage count, and the second expected value is the product of this probability value and the second usage count.
[0057] For example, a specific first word appears a total of 100 times in the collection (2 times in the target vocabulary and 98 times in the corpus data), and its occurrences as the first word in the target vocabulary and the corpus data are 1 and 20 times respectively. Then, the probability of this first word appearing as the first word is 0.21 (or 21%), corresponding to a first expected value of 0.21 and a second expected value of 4.2. In other words, the expected number of times this first word appears as the first word in the target vocabulary is 0.21 and the expected number of times it appears in the corpus data is 4.2.
[0058] like Figure 2 As shown, in one embodiment, step S400, which is to determine the expected information for the first word to be used as the first word based on the word setting statistics and vocabulary statistics, specifically includes steps S410, S420 and S430.
[0059] Step S410: Determine the first expected value based on the word statistics and vocabulary statistics.
[0060] In one embodiment, step 410, namely determining the first expected value based on word statistics and vocabulary statistics, specifically includes: determining the first expected value based on the first number of uses, the second number of uses, the first number of occurrences, and the second number of occurrences.
[0061] In one example, the first expected value X1 can be calculated according to the following formula: In this formula, a is the first number of times it is used, b is the second number of times it is used, c is the first number of times it appears, and d is the second number of times it appears.
[0062] Step S420: Determine the second expected value based on the word statistics and vocabulary statistics.
[0063] In one example, the first expected value X2 can be calculated according to the following formula:
[0064] Step S430: Determine the expected number of times the first word is used as the first word based on the first expected value and the second expected value.
[0065] In one example, the expected number of times the first word is used as the first word, X, can be calculated using the following formula:
[0066] Step 500: Determine the first characteristic of the first word based on the expected information.
[0067] Here, the first characteristic is at least the word setting with the maximum expectation of the first word. Specifically, after determining the expected value of the number of times the first word is used as the first word setting in step S300, the word setting with the maximum expectation of the first word is determined as the first characteristic of the first word setting.
[0068] like Figure 2 As shown, in one embodiment, step S500, which is to determine the first characteristic of the first word based on the desired information, includes steps S510 and S520.
[0069] Step S510: Determine whether the expected number of times the first word is used as the first word is greater than the expected number of times the first word is used as the second word.
[0070] Here, the second lexical assumption can be any lexical assumption of the first lexical term other than the first lexical assumption. In other words, the expected value of the first lexical assumption of the first lexical term is compared with the expected values of other lexical assumptions of the first lexical term.
[0071] Step S520: When the expected number of times the first word is used as the first word setting is greater than the expected number of times the first word is used as the second word setting, the first word setting is determined as the first characteristic of the first word.
[0072] In one embodiment, the first characteristic can be a lexical feature of the first vocabulary; the first lexical feature determined as the first characteristic in step S520 can be a lexical feature of the first vocabulary; in this case, the first lexical feature has the maximum expected value, i.e., the highest number of uses. In another embodiment, the first characteristic can be multiple lexical features of the first vocabulary; in other words, the first lexical feature determined as the first characteristic can be N ≥ 2 lexical features of the first vocabulary; in this case, these N lexical features are the lexical features of the first vocabulary that rank in the top N in terms of the number of uses. In yet another embodiment, the first characteristic can be multiple lexical features of the first vocabulary, and the ranking information among these lexical features. This ranking information can indicate that the multiple lexical features are arranged in order of the most used to the least used.
[0073] In another embodiment, the first characteristic may also be the sorting information of all or part of the first vocabulary. In this case, when the vocabulary is sorted by frequency of use from most to least, the vocabulary with the highest sorting in the first characteristic is the first vocabulary.
[0074] Specifically, by comparing the first and second word settings, the word setting with the highest expected value among all word settings of the first vocabulary can be finally determined.
[0075] The vocabulary processing method in this embodiment determines the expected information of the first vocabulary as a first lexical setting based on the lexical setting statistics and the vocabulary statistics of the first vocabulary. Based on this expected information, the lexical setting with the maximum expected value of the first vocabulary is determined as the first feature. This provides the highest frequency lexical setting of the vocabulary in the vocabulary set, enabling users to master the actual usage of the vocabulary at a specific learning stage, thereby improving the user's learning efficiency.
[0076] Figure 3 A structural block diagram of a vocabulary processing method provided by another exemplary embodiment is shown. Figure 4 It shows Figure 3 A detailed flowchart of the vocabulary processing methods in [the document / system].
[0077] like Figure 3 As shown, in addition to the steps described above, the vocabulary processing method may also include step S600.
[0078] Step S600: Create or maintain corpus data.
[0079] In one embodiment, the corpus data may include a vocabulary set for a specific learning stage and multiple corresponding sentences. In another embodiment, the corpus data may further include annotations of the vocabulary set for each sentence. Therefore, creating corpus data may include at least one of the following: collecting multiple vocabulary sets for a specific learning stage, collecting multiple sentences containing the above vocabulary sets, and annotating the vocabulary sets in the above sentences.
[0080] In one embodiment, maintaining corpus data may include at least one of the following: adding or deleting words, adding or deleting sentences, annotating the lexical settings of added words in sentences, deleting the annotations of deleted words in sentences, and annotating the lexical settings of added words in sentences.
[0081] In addition, the vocabulary processing method may also include steps S700 and S800.
[0082] Step S700: Determine the probability value of the first word being used as the first word in the corpus data.
[0083] Specifically, the frequency of the first word as the first word in the corpus data is statistically analyzed, as is the total frequency of the first word in the corpus data. Based on the total frequency and the frequency of its use as the first word, the probability value of the first word as the first word in the corpus data is determined. In other words, the probability value is calculated using the following formula: As mentioned earlier, b is the second number of times it is used, and d is the second number of times it appears.
[0084] In one example, the probability value of a word can be calculated in real time. In another example, the probability value of a word can be pre-calculated and stored, and retrieved when needed.
[0085] In another example, given existing corpus data containing word probability values for each word, the word probability values can be obtained directly from the existing corpus data.
[0086] Step S800: Based on the corpus data and target vocabulary, determine the average value of the vocabulary used as the first word setting.
[0087] Here, the average word setting refers to the average value of the first word setting of the first word in the target vocabulary and corpus data among the various word settings of the first word.
[0088] like Figure 4 As shown, in one embodiment, step S800, which is to determine the average value of the first word as the first word setting based on the corpus data and the target vocabulary, includes steps S810 and S820.
[0089] Step S810: Based on the corpus data and target vocabulary, determine the vocabulary statistics of the first vocabulary used as the second vocabulary setting.
[0090] As mentioned above, the second lexical setting refers to any lexical setting of the first word other than the first lexical setting. In particular, here we statistically analyze all lexical settings of the first word other than the first lexical setting to obtain the corresponding lexical setting statistics.
[0091] Step S820: Determine the average value of the word used as the first word setting based on the word setting statistics of the first word setting used as the first word setting and the word setting statistics of the first word setting used as the second word setting.
[0092] In one example, the vocabulary statistics used as the first word setting and the vocabulary statistics used as the second word setting are added together to obtain the total vocabulary statistics of the first word setting used as all word settings, thereby further obtaining the average vocabulary value.
[0093] In one example, the word average value Av is calculated according to the following formula: Among them, a1, a2, ..., a n It represents the number of times the first word is used in the target vocabulary as one of n different word types, and b1, b2, ..., b... n This represents the number of times the first word is used as one of the n word types in the corpus data. Correspondingly, (a1+b1) represents the number of times the first word is used as the first word type, and (a2+b2)+...+(a...b1) represents the number of times the first word is used as the first word type. n +b n ) represents the number of times the first word is used as the second word, therefore (a1+b1)+(a2+b2)+...+(a n +b n () represents the number of times the first word is used as part of the entire vocabulary list.
[0094] Accordingly, step S400, which is to determine the first characteristic of the first word based on the expected information, includes: determining the first characteristic of the first word based on the expected information, the word probability value and the word average value.
[0095] like Figure 4 As shown, in one embodiment, step S500, which is to determine the first characteristic of the first word based on the expected information, the word probability value and the word average value, includes steps S510, S530 and S540.
[0096] Step S510: Determine whether the expected number of times the first word is used as the first word is greater than the expected number of times the first word is used as the second word. It should be noted that step S430 here is the same as step S410 in the previous text.
[0097] Step S530: Determine whether the probability value of the first word is greater than the average value of the word.
[0098] Specifically, the probability value of the first word being used as the first word setting is compared with the average value of the word setting when the first word is used as the first word setting.
[0099] Step S540: When the expected number of times the first word is used as the first word setting is greater than the expected number of times the first word is used as the second word setting, and the word setting probability value of the first word is greater than the word setting average value, the first word setting is determined as the first characteristic of the first word.
[0100] Specifically, in this case, it is required that not only is the expected number of times the first word is used as the first word setting greater than the expected number of times the first word is used as the second word setting, but also that the word setting probability value of the first word is greater than the word setting average value. The first word setting that satisfies both of these conditions is determined as the first characteristic.
[0101] By incorporating the probability value and average value of word settings into the expected information, the vocabulary processing method of this embodiment can provide the highest frequency word settings of words more accurately.
[0102] The following examples illustrate the practical application of the aforementioned methods.
[0103] Example 1:
[0104] When the target word is only "insist", the first word can only be "insist". Therefore, the method described above is used to determine the expected value, probability value, and average value of the frequency of use of "insist" for each word based on the target word and existing corpus data. Figure 5 The results for the expected value, probability value, and average value of the keyword for insert are given, represented by Keyness, P, and Avg, respectively. In the results list shown, the first column gives the keyword for insert, the second column gives the keyword probability value P for each keyword, the third column gives the average value Avg for each keyword, and the fourth column gives the expected value Keyness for each keyword.
[0105] from Figure 5 As can be seen from the data, among the various verb constructs of `insist`, the verbs_ccomp (i.e., verb + that clause) has the largest expected value of 7080.4900, and the corresponding probability value of 0.6250 is much greater than the average value of 0.0250. Therefore, the verbs_ccomp is determined as the first characteristic of `insist`.
[0106] Example 2:
[0107] In one scenario, the target word is obtained by receiving user input. For example, if the user inputs "brink" as the target word, then the first word will be "brink". Figure 6 As shown, the lexical construct in_pp (i.e., the object phrase) of the word brink has the largest expected value of 292.8700, and the corresponding lexical construct probability value of 0.9718 is greater than the average lexical construct value of 0.2916. Therefore, the lexical construct in_pp is determined as the first characteristic of brink.
[0108] Example 3:
[0109] In some cases, the target vocabulary is directly provided by the word memorization app. For example, if the word to be memorized is "plan," then the first word would be "plan." Figure 7As shown, the verb + infinitive tov (vtov) of the lexical plan has the largest expected value of 7531.5200, and the corresponding tov probability value of 0.3000 is greater than the average tov value of 0.0100. Therefore, the tov is determined as the first characteristic of the plan.
[0110] Figure 8 A structural block diagram of a vocabulary processing apparatus provided in an exemplary embodiment is shown. Figure 8 As shown, the vocabulary processing device includes a target vocabulary determination module 10, a corpus data acquisition module 20, a statistical data determination module 30, an expected information determination module 40, and a feature determination module 50. The target vocabulary determination module 10 is configured to determine target vocabulary, wherein the target vocabulary includes at least one word. The corpus data acquisition module 20 is configured to acquire corpus data, wherein the corpus data includes at least a plurality of sentences, and at least one of the plurality of sentences contains the target vocabulary. The statistical data determination module 30 is configured to determine, based on the corpus data and the target vocabulary, a lexical setting statistical data and a statistical data of the first word contained in the target vocabulary as a first lexical setting. The expected information determination module 40 is configured to determine expected information of the first word as a first lexical setting based on the lexical setting statistical data and the vocabulary statistical data. The feature determination module 50 is configured to determine a first feature of the first word based on the expected information, wherein the first feature is at least the lexical setting of the first word with the maximum expectation.
[0111] In one embodiment, the expectation information determination module 40 includes a first expectation value determination submodule 41, a second expectation value determination submodule 42, and an expectation value determination submodule 43. The first expectation value determination submodule 41 is configured to determine a first expectation value based on word attribute statistics and vocabulary statistics, wherein the first expectation value is the expected number of times a first word is used as a first word attribute in the target vocabulary. The second expectation value determination submodule 42 is configured to determine a second expectation value based on word attribute statistics and vocabulary statistics, wherein the second expectation value is the expected number of times a first word is used as a first word attribute in the corpus data. The expectation information determination submodule 43 is configured to determine the expected number of times the first word is used as a first word attribute based on the first expectation value and the second expectation value.
[0112] In one embodiment, the lexical statistics include the first number of times the first word is used as the first lexical term in the target vocabulary and the second number of times the first word is used as the first lexical term in the corpus data. The lexical statistics also include the first number of times the first word appears in the target vocabulary and the second number of times the first word appears in the corpus data. A first expected value determination submodule 41 is configured to determine a first expected value based on the first number of uses, the second number of uses, the first number of appearances, and the second number of appearances. A second expected value determination submodule 42 is configured to determine a second expected value based on the first number of uses, the second number of uses, the first number of appearances, and the second number of appearances. An expected information determination submodule 43 is configured to determine the expected value of the number of times the first word is used as the first lexical term based on the first expected value and the second expected value.
[0113] In one embodiment, the feature determination module 50 includes an expected value comparison submodule 51 and a feature determination submodule 52. The expected value comparison submodule 51 is configured to determine whether the expected value of the number of times the first word is used as a first term is greater than the expected value of the number of times the first word is used as a second term, wherein the second term is different from the first term. The feature determination submodule 52 is configured to determine that the first word is a first feature of the first word when the expected value of the number of times the first word is used as a first term is greater than the expected value of the number of times the first word is used as a second term.
[0114] It can be seen that, Figure 8 The vocabulary processing device shown is used to implement Figure 1 and 2 The vocabulary processing method shown is described above. Therefore, a detailed description of this vocabulary processing device can be found in the corresponding vocabulary processing method.
[0115] Figure 9 A structural block diagram of a vocabulary processing apparatus provided in another exemplary embodiment is shown.
[0116] like Figure 9 As shown, in one embodiment, in addition to the modules described above, the vocabulary processing apparatus further includes a corpus data creation / maintenance module 60. The corpus data creation / maintenance module 60 is configured to create / maintain corpus data.
[0117] like Figure 9 As shown, in one embodiment, the vocabulary processing apparatus further includes a probability value determination module 70 and an average value determination module 80. The probability value determination module 70 is configured to determine the probability value of the first vocabulary being used as a first lexical term in the corpus data. The average value determination module 80 is configured to determine the average value of the first vocabulary being used as a first lexical term based on the corpus data and the target vocabulary. Accordingly, the feature determination module 60 is configured to determine a first feature of the first vocabulary based on the desired information, the probability value, and the average value.
[0118] In one embodiment, the average value determination module 80 includes a statistical data determination submodule 81 and a word average value determination submodule 82. The statistical data determination submodule 81 is configured to determine word average data for a first word used as a second word setting based on corpus data and target vocabulary. The word average value determination submodule 82 is configured to determine the word average value for the first word used as a first word setting based on the word average data for the first word setting used as a first word setting and the word average data for the first word setting used as a second word setting.
[0119] In one embodiment, the feature determination module 50 includes an expected value comparison submodule 51, an average value comparison submodule 53, and a feature determination submodule 54. The expected value comparison submodule 51 is configured to determine whether the expected value of the number of times the first word is used as a first term is greater than the expected value of the number of times the first word is used as a second term. The average value comparison submodule 53 is configured to determine whether the term probability value of the first word is greater than the term average value. The feature determination submodule 54 is configured to determine that the first word is a first feature of the first word when it is determined that the expected value of the number of times the first word is used as a first term is greater than the expected value of the number of times the first word is used as a second term, and the term probability value of the first word is greater than the term average value.
[0120] It can be seen that, Figure 9 The vocabulary processing device shown is used to implement Figure 3 and 4 The vocabulary processing method shown is described above. Therefore, a detailed description of this vocabulary processing device can be found in the corresponding vocabulary processing method.
[0121] Figure 10 A structural block diagram of a vocabulary processing device 1000 provided in an exemplary embodiment is shown. For example... Figure 10 As shown, the vocabulary processing device may include: a processor 1100 and a memory 1200 for storing processor-executable instructions; wherein the processor 1100 is configured to implement the vocabulary processing method described in any embodiment of the present disclosure when executing the executable instructions.
[0122] Based on the same inventive concept, this disclosure also provides a computer-readable storage medium storing an executable program, wherein the executable program, when executed by a processor, implements the vocabulary processing method described in any embodiment of this disclosure.
[0123] Those skilled in the art will appreciate that the functionality described in conjunction with the various illustrative logic blocks, modules, and algorithmic steps disclosed herein can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality described by the various illustrative logic blocks, modules, and steps can be stored or transmitted as one or more instructions or codes on a computer-readable medium and executed by a hardware-based processing unit. The computer-readable medium may comprise a computer-readable storage medium, which corresponds to a tangible medium, such as a data storage medium, or a communication medium that includes any medium facilitating the transfer of a computer program from one place to another (e.g., according to a communication protocol). In this way, the computer-readable medium may substantially correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium, such as a signal or carrier wave. The data storage medium may be any available medium accessible by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this application. A computer program product may comprise a computer-readable medium.
[0124] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other media that can be used to store desired program code in the form of instructions or data structures and is accessible by a computer. Furthermore, any connection is properly referred to as computer-readable media. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of media. However, it should be understood that the computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other temporary media, but are specifically addressed to non-temporary tangible storage media. As used herein, disks and optical discs include compact optical discs (CDs), laser optical discs, optical discs, digital versatile optical discs (DVDs), and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. The combination of the above items should also be included in the scope of computer-readable media.
[0125] Instructions can be executed by one or more processors, such as digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the term "processor" as used herein can refer to any of the foregoing structures or any other structures suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described in the various illustrative logic blocks, modules, and steps described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Moreover, the techniques can be fully implemented within one or more circuit or logic elements.
[0126] The technology of this application can be implemented in a wide variety of devices or apparatuses, including wireless handheld devices, integrated circuits (ICs), or a set of ICs (e.g., chipsets). The various components, modules, or units described in this application are intended to emphasize functional aspects of the apparatuses used to perform the disclosed technology, but do not necessarily need to be implemented by different hardware units. In fact, as described above, the various units can be combined with suitable software and / or firmware within a codec hardware unit, or provided via interoperable hardware units (comprising one or more processors as described above).
[0127] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0128] The above description is merely an exemplary embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A vocabulary processing method, characterized in that, include: Identify target vocabulary, wherein the target vocabulary includes at least one word; Obtain corpus data, wherein the corpus data includes at least a plurality of statements, and at least one of the plurality of statements contains the target vocabulary; Based on the corpus data and the target vocabulary, determine the word setting statistics and vocabulary statistics of the first word contained in the target vocabulary as a first word setting, wherein the first word setting includes the collocation usage of the first word, or the combination of collocation usage and part of speech, the word setting statistics include the first number of times the first word is used as the first word setting in the target vocabulary and the second number of times the first word is used as the first word setting in the corpus data, and the vocabulary statistics include the first number of times the first word appears in the target vocabulary and the second number of times the first word appears in the corpus data; Based on the lexical statistics and the vocabulary statistics, a first expected value and a second expected value are determined. Based on the first expected value and the second expected value, an expected value is determined of the number of times the first word is used as the first lexical term, wherein the first expected value is the expected number of times the first word is used as the first lexical term in the target vocabulary, and the second expected value is the expected number of times the first word is used as the first lexical term in the corpus data; and Based on the expected value, a first characteristic of the first word is determined, wherein the first characteristic is at least the word characteristic of the first word with the maximum expected value.
2. The method according to claim 1, characterized in that, The step of determining the first expected value and the second expected value based on the word statistics and the vocabulary statistics includes: The first expected value is determined based on the first number of uses, the second number of uses, the first number of occurrences, and the second number of occurrences. The second expected value is determined based on the first number of uses, the second number of uses, the first number of occurrences, and the second number of occurrences.
3. The method according to claim 1, characterized in that, Determining the first characteristic of the first word based on the expected value includes: Determine whether the expected number of times the first word is used as the first word setting is greater than the expected number of times the first word is used as the second word setting, wherein the second word setting is different from the first word setting; and When it is determined that the expected number of times the first word is used as the first word setting is greater than the expected number of times the first word is used as the second word setting, the first word setting is set as the first characteristic of the first word.
4. The method according to claim 3, characterized in that, The method further includes: Determine the probability value of the first word being used as the first word in the corpus data; and Based on the corpus data and the target vocabulary, determine the average value of the vocabulary used as the first word setting; The determination of the first characteristic of the first word based on the expected value includes: The first characteristic of the first word is determined based on the expected value, the word probability value, and the word average value.
5. The method according to claim 4, characterized in that, Based on the corpus data and the target vocabulary, the average value of the vocabulary used as the first word setting is determined, including: Based on the corpus data and the target vocabulary, determine the vocabulary statistics of the first vocabulary used as the second vocabulary setting; and Based on the statistical data of the word setting used as the first word setting and the statistical data of the word setting used as the second word setting, the average word setting value of the first word setting used as the first word setting is determined.
6. The method according to claim 4, characterized in that, Based on the expected value, the word probability value, and the word average value, the first characteristic of the first word is determined, including: Determine whether the expected number of times the first word is used as the first word is greater than the expected number of times the first word is used as the second word; Determine whether the probability value of the first word is greater than the average value of the word; and When it is determined that the expected number of times the first word is used as the first word setting is greater than the expected number of times the first word is used as the second word setting, and the word setting probability value of the first word is determined to be greater than the word setting average value, the first word setting is set as the first characteristic of the first word.
7. A vocabulary processing device, characterized in that, include: The target vocabulary determination module is configured to determine target vocabulary, wherein the target vocabulary includes at least one word; The corpus data acquisition module is configured to acquire corpus data, wherein the corpus data includes at least a plurality of statements, and at least one of the plurality of statements contains the target vocabulary; The statistical data determination module is configured to determine, based on corpus data and the target vocabulary, a first word contained in the target vocabulary as a first word setting statistical data and a vocabulary statistical data of the first word. The first word setting includes the collocation usage of the first word, or a combination of collocation usage and part of speech. The word setting statistical data includes the first number of times the first word is used as the first word setting in the target vocabulary and the second number of times the first word is used as the first word setting in the corpus data. The vocabulary statistical data includes the first number of times the first word appears in the target vocabulary and the second number of times the first word appears in the corpus data. The expected information determination module is configured to determine a first expected value and a second expected value based on the word setting statistics and the vocabulary statistics, and to determine an expected value for the number of times the first word is used as the first word setting based on the first expected value and the second expected value, wherein the first expected value is the expected number of times the first word is used as the first word setting in the target vocabulary, and the second expected value is the expected number of times the first word is used as the first word setting in the corpus data; and The feature determination module is configured to determine a first feature of the first word based on the expected value, wherein the first feature is at least the word characteristic of the first word with the maximum expectation.
8. A vocabulary processing device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to implement the vocabulary processing method according to any one of claims 1 to 6 when executing the executable instructions.
9. A non-volatile computer-readable storage medium, characterized in that, The readable storage medium stores an executable program, wherein the executable program, when executed by a processor, implements the vocabulary processing method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Word segmentation device obtaining method and device and electronic equipment
CN112101016A
Medical named entity recognition model training method and medical named entity recognition method
CN113051905A