English text vocabulary sequence similarity detection method
By introducing a detection method for word order features in English text, the problem of ignoring word order features in the existing technology is solved, and more accurate word order similarity evaluation and semantic understanding are achieved.
Patent Information
- Application Number
- CN202510803821.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-19
AI Technical Summary
Existing methods for checking word order similarity in English texts ignore word order features, resulting in an inability to effectively evaluate the similarity of word orders and affecting semantic understanding.
A detection method based on lexical order features was designed. Through word segmentation, sentence segmentation and stemming, a sentence pair set was constructed to generate a joint vocabulary set. The lexical order similarity was evaluated using the lexical order vector and similarity calculation formula.
The effective integration of word order features improves the accuracy of detecting word order similarity in English texts and enhances the ability to understand sentence semantics.
Smart Images

Figure CN120670867A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a method for detecting similarity of word order in English texts based on the order characteristics of words in English texts. The present invention is only applicable to similarity analysis of word order in English texts. Background Art
[0002] Existing methods for checking the similarity of word order in English texts typically use the edit distance method, the longest common subsequence method, and the statistical model method. These methods evaluate the similarity of words by calculating the overlapping parts between words in English texts. Although this method can identify similar words in English texts, it generally ignores the word order characteristics in English texts, which are very important for understanding the semantics of sentences. To address the above problems, the present invention proposes a method for detecting the similarity of word order in English texts based on the word order characteristics in English texts. This method incorporates the word order characteristics in English texts and evaluates the similarity of word orders in English texts through a designed word order feature dual mechanism. Summary of the Invention
[0003] The present invention provides a method for detecting similarity in the order of English words. The method processing flow chart is as follows: Figure 1 As shown in , its processing flow is: first, the English text to be detected is segmented into words and sentences, and a vocabulary chain of the sentences is generated, and different forms of the same vocabulary in the vocabulary chain are de-stemmed; second, the sentences in the English text to be detected are matched sentence by sentence in order to construct a sentence pair set of the English text to be detected; third, each pair of sentences in the sentence pair set is processed by set operation to generate a joint vocabulary set of all non-repeated words in the sentence pair set; fourth, according to the occurrence of words in the joint vocabulary set in the sentences of the English text to be detected, and combined with the vocabulary order vector to construct a formula, the position number of the words in the joint vocabulary set in the sentences of the English text to be detected is generated. The position number generation method is: if the word appears in the sentence, then the word The position number of a word is the word position value in the sentence; if the word does not appear in the sentence, then the synonyms of the word are first calculated using the word synonym similarity calculation formula and the word synonym similarity judgment value calculation formula, and then the word position value of the obtained synonym in the sentence is used as the position number of the word; if the synonym of the word cannot be obtained using the word synonym similarity calculation formula and the word synonym similarity judgment value calculation formula, the position number of the word is 0; fifth, based on the constructed word order vector, the word order similarity calculation formula is first used to calculate the word order similarity, and then the word order similarity judgment result calculation formula is used to determine whether the word order is similar or dissimilar based on the obtained word order similarity.
[0004] The calculation formula of the present invention is defined as follows:
[0005] 1. Calculation formula for vocabulary synonym similarity
[0006]
[0007] In formula (1), word frequency i is the vocabulary in the English text to be detected i The ratio of the number of occurrences to the total number of words in the English text to be tested, word frequency j is the vocabulary in the English text to be detected j The ratio of the number of occurrences to the total number of words in the English text to be tested, co-occurrence frequency ij is the vocabulary in the English text to be detected i and vocabulary j The ratio of the number of co-occurrences to the total number of words in the English text to be tested. The value range of word synonym similarity is [0,1].
[0008] 2. Calculation formula for vocabulary synonymous similarity judgment value
[0009]
[0010] In formula (2), the synonymous similarity of words is calculated by formula (1), and the synonymous similarity judgment value is 1 or 0, where 1 means that the word i and vocabulary j is a synonym, 0 means vocabulary i and vocabulary j Not synonyms.
[0011] 3. Construction formula of vocabulary order vector
[0012] Vocabulary order vector = [v1, v2, v i ,…,v n ], i = 1, 2, ..., n, n is the total number of different words in the English text T to be detected,
[0013] in,
[0014] In formula (3), T is the English text to be detected, w i is the i-th word in the vocabulary set of the English text T to be detected, and the word u is calculated by formula (2) to determine whether it is the word w i Synonyms of v i It is the vocabulary w i or vocabulary u i The position number in the English text T to be detected, position (w i , T) is the vocabulary w iThe position number in the English text T to be detected, position (u, T) is the position number of the word u in the English text T to be detected.
[0015] 4. Calculation formula for word order similarity
[0016]
[0017] In formula (4), the vocabulary order vector i and word order vectors j It is calculated by formula (3), and the range of word order similarity is [0, 1].
[0018] 5. Calculation formula for word order similarity judgment results
[0019]
[0020] In formula (5), the word order similarity judgment result is calculated by formula (4), and the word order similarity judgment result is: similar or dissimilar.
[0021] Specific processing steps of the diagnostic method of the present invention
[0022] like Figure 1 As shown, the processing flow of the English text word order similarity detection method is as follows:
[0023] P101 starts;
[0024] P102 reads the English text to be tested;
[0025] P103 pre-processes the English text to be tested, including sentence segmentation, generating sentence word chains, and de-stemming different forms of the same word in the word chain;
[0026] P104 constructs a sentence pair set, matching the sentences in the English text to be detected sentence by sentence in order to construct a sentence pair set of the English text to be detected;
[0027] P105 generates a joint vocabulary set, performs set operation processing on each sentence in the sentence pair set, and generates a joint vocabulary set of all non-repeated words in the sentence pair;
[0028] P106 constructs an initial lexical order vector, generating a zero vector with the same dimension as the joint word set for each sentence in the sentence pair as the initial lexical order vector;
[0029] P107 fills the lexical order vector and analyzes the sentence pair to be tested. For words in the joint word set that exist in the current sentence, directly extract their position numbers using formula (3); for words in the joint word set that do not exist in the current sentence, find their synonyms based on formulas (1) and (2). Then, use formula (3) to extract the position number of the synonym and fill it into the corresponding position of the lexical order vector of the corresponding sentence.
[0030] P108 Lexical order similarity determination, based on the constructed lexical order vector, first use formula (4) to calculate the lexical order similarity, and then use formula (5) to determine whether the lexical order is similar or dissimilar based on the obtained lexical order similarity;
[0031] End of P109. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a processing flow chart of a method for detecting similarity of word order in English text. DETAILED DESCRIPTION
[0033] In the embodiment of the present invention, an English text titled "Technology in Education" is input. The English text is divided into a text to be detected and an original English text. The English text word order similarity detection method mainly includes the following steps:
[0034] (1) Read the English text to be tested. The results are as follows:
[0035] English text to be tested
[0036] In today’s world,technology has become an essential part ofeducation,playing a major role in improving the learning experience for bothteachers and students.In today’s education,technology plays a major role asan essential tool,significantly improving the learning experience for bothteachers and students.Digital tools like online platforms and smart devicesallow educators to create more interactive and personalized lessons.With thehelp of AI and digital learning systems,teachers can better understandstudent progress and adjust their methods accordingly.For example,virtualreality can bring difficult subjects tolife,making them easier andmoreinteresting for students.Interactive tools such as videos,podcasts,andeducational apps provide students with alternative ways to absorbinformation.One major benefit of using technology is that it gives studentsaccess to a wide range of resources anytime and anywhere.One major advantage of using technology is that it allows students to access a wide variety of educational resources at any time and from any place. However, not all students have equal access to these tools, and some may face privacy or distraction issues. Despite its benefits, there are still challenges like unequal access to devices and concerns about screen time and data safety. If used wisely, technology can truly transform how we teach and learn, making education more efficient and engaging. When implemented thoughtfully, technology can greatly enhance the quality of education and prepare students for the future.
[0037] (2) English text preprocessing. The results are as follows:
[0038] In today world, technology has become an essential part of education, play a major role in improve the learn experience for both teacher and student.
[0039] In today education, technology play a major role as an essential tool, significantly improve the learn experience for both teacher and student.
[0040] Digital tool like online platform and smart device allow educator tocreate more interactive and personalize lesson.
[0041] With the help of ai and digital learning systems,teachers can betterunderstand student progress and adjust their method accordingly.
[0042] For example,virtual reality can bring difficult subject to life,makethem easy and more interesting for student.
[0043] Interactive tool such as video,podcast,and educational app providestudent with alternative ways to absorb information.
[0044] One major benefit of use technology is that it give student access toa wide range of resource anytime and anywhere.
[0045] One major advantage of use technology is that it allow student toaccess a wide variety of educational resource at any time and from any place.
[0046] However, not all students have equal access to these tools, and some may face privacy or distraction issues.
[0047] Despite its benefits, there are still challenges like unequal access to devices and concerns about screen time and data safety.
[0048] If used wisely, technology can truly transform how we teach and learn, making education more efficient and engaging.
[0049] When implemented thoughtfully, technology can greatly enhance the quality of education and prepare students for the future.
[0050] (3) Construct a set of sentence pairs. The results are as follows:
[0051] Sentence pair 1 - ["In today's world, technology has become an essential part of education, playing a major role in improving the learning experience for both teachers and students.", "In today's education, technology plays a major role as an essential tool, significantly improving the learning experience for both teachers and students."] Sentence pair 2 - ["Digital tools like online platforms and smart devices allow educators to create more interactive and personalized lessons.", "With the help of AI and digital learning systems, teachers can better understand student progress and adjust their methods accordingly."]
[0052] Sentence pair 3 - ["For example, virtual reality can bring difficult subjects to life, making them easy and more interesting for students.", "Interactive tools such as videos, podcasts, and educational apps provide students with alternative ways to absorb information."]
[0053] Sentence pair 4 - ["One major benefit of using technology is that it gives students access to a wide range of resources anytime and anywhere.", "One major advantage of using technology is that it allows students to access a wide variety of educational resources at any time and from any place."]
[0054] Sentence pair 5 - ["However, not all students have equal access to these tools, and some may face privacy or distraction issues.", "Despite its benefits, there are still challenges like unequal access to devices and concerns about screen time and data safety."]
[0055] Sentence pair 6 - ["If used wisely, technology can truly transform how we teach and learn, making education more efficient and engaging.", "When implemented thoughtfully, technology can greatly enhance the quality of education and prepare students for the future."]
[0056] (4) Generate a combined vocabulary set. The results are as follows:
[0057] Combined Word Set 1 = [in, today, world, technology, has, become, an, essential, part, of, education, play, a, major, role, improve, the, learn, experience, for, both, teacher, and, student, as, tool, significantly]
[0058] Combined Word Set 2 = [digital, tool, like, online, platform, and, smart, device, allow, educator, to, create, more, interactive, personalize, lesson, with, the, help, of, ai, learning, systems, teachers, can, better, understand, student, progress, adjust, their, method, accordingly]
[0059] Combined Word Set 3 = [for, example, virtual, reality, can, bring, difficult, subject, to, life, make, them, easy, and, more, interesting, student, interactive, tool, such, as, video, podcast, educational, app, provide, with, alternative, ways, absorb, information]
[0060] Combined Word Set 4 = [one, major, benefit, of, use, technology, is, that, it, give, student, access, to, a, wide, range, of, resource, anytime, and, anywhere, advantage, allow, variety, educational, at, any, time, from, place]
[0061] Joint word set 5=[however,not,all,student,have,equal,access,to,these,tool,and,some,may,face,privacy,or,distraction,issue,despite,its,benefit,there,are,still,challenge,like,unequal,device,concern,about,screen,time,data,safety]
[0062] Combine word set 6 = [if,use,wisely,technology,can,truly,transform,how,we,teach,and,learn,make,education,more,efficient,engaging,when,implement,thoughtfully,greatly,enhance,the,quality,of,prepare,student,for,future](5) to construct the initial word order vector. The result is as follows:
[0063] Initial word order vector of sentence pair 1 = [0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0], [0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0]
[0064] Initial word order vector of sentence pair 2 = [0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0], [0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0]
[0065] Initial word order vector of sentence pair 3 = [0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0], [0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0]
[0066] Initial word order vector of sentence pair 4 = [0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0], [0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0]
[0067] Initial word order vector of sentence pair 5 = [0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0], [0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0]
[0068] Initial word order vector of sentence pair 6 = [0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0], [0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0]
[0069] (6) Fill in the vocabulary order vector. The result is as follows:
[0070] Word order vector of sentence pair 1 = [1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,18,19,20,21,22,23,24,25,26,0,14,0,0,0,0,0,0,0,0,0,0,0,0,0,0], [1,2,3,0,5,18,0,11,12,9,16,4,0,7,8,9,15,16,17,18,19,20,21,22,23,6,10,13,14,0,0,0,0,0,0 The word order vector of sentence pair 2 = [1,2,3,4,5,6,7,8,9,10,11,12,13,14,16,17,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0], [0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0]
[0071] The word order vector of sentence pair 3 = [2, 2, 3, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 17, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 14, 0, 0, 0, 6, 0, 0, 10, 1, 2, 3, 3, 4, 5, 7, 8, 9, 11, 12, 13, 15, 16, 0, 0, 0, 0, 0, 0, 0, 0, 0]
[0072] Word order vector of sentence pair 4 = [1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,18,19,20,21,0,0,0,0,0,0,0,0,16,0,0,0,0,0,0,0,0,0,0,0,0,0,0], [1,2,0,4,5,6,7,8,9,0,11,13,12,14,15,25,19,8,22,0,3,10,16,18,20,21,21,23,25,0,0,0,0,0,0,0,0,0,0,0,0,0]
[0073] The word order vector of sentence pair 5 = [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 0, 0, 0, 6, 1, 0, 0, 0, 0, 12, 0, 0, 0, 0, 0, 0, 0, 0, 0], [5, 0, 0, 0, 4, 9, 10, 3, 0, 12, 14, 0, 0, 0, 0, 0, 1, 2, 3, 4, 4, 5, 6, 7, 8, 11, 13, 14, 15, 15, 17, 18, 0, 0, 0, 0, 0, 0]
[0074] The word order vector of sentence pair 6 = [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18, 0, 0, 0, 0, 0, 0, 13, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 4, 5, 0, 0, 5, 0, 11, 0, 12, 10, 0, 0, 1, 2, 3, 6, 7, 8, 8, 9, 12, 13, 14, 16, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]
[0075] (7) Determination of vocabulary order similarity. The results are as follows:
[0076] Word order similarity values: 0.7886, 0.0000, 0.1714, 0.6479, 0.3231, 0.3770
[0077] Similarity judgment results: similar, dissimilar, dissimilar, similar, dissimilar, dissimilar.
Claims
1. A method for detecting similarity in word order in English texts, the diagnostic method comprising the following processing steps: first, segmenting the English text to be tested into words and sentences, generating a lexical chain of sentences therein, and performing stemming on different forms of the same word in the lexical chain; second, matching the sentences in the English text to be tested sentence by sentence in order to construct a set of sentence pairs in the English text to be tested; third, performing a set operation on each pair of sentences in the set of sentence pairs to generate a joint lexical set of all non-repeated words in the sentence pair set; Fourth, based on the occurrence of the words in the joint word set in the sentences of the English text to be tested, and combined with the word order vector to construct a formula, the position number of the words in the joint word set in the sentences of the English text to be tested is generated. The position number generation method is: if the word appears in the sentence, then the position number of the word is the word position value in the sentence; if the word does not appear in the sentence, then first use the word synonym similarity calculation formula and the word synonym similarity judgment value calculation formula to calculate the synonym of the word, and then use the word position value of the obtained synonym in the sentence as the position number of the word; If the synonyms of the word cannot be obtained using the word synonym similarity calculation formula and the word synonym similarity judgment value calculation formula, the position number of the word is 0; Fifth, based on the constructed word order vector, the word order similarity is first calculated using the word order similarity calculation formula, and then, based on the obtained word order similarity, the word order similarity judgment result calculation formula is used to determine whether the word order is similar or dissimilar.
2. The method for detecting similarity in word order of English text according to claim 1, wherein: The calculation formula of the English text word order similarity detection method is defined as follows: (1) Calculation formula for vocabulary synonym similarity In formula (1), word frequency i The word frequency is the ratio of the number of times word i appears in the English text to be tested to the total number of words in the English text to be tested. j is the vocabulary in the English text to be detected j The ratio of the number of occurrences to the total number of words in the English text to be tested, co-occurrence frequency ij is the vocabulary in the English text to be detected i and vocabulary j The ratio of the number of co-occurrences to the total number of words in the English text to be tested. The value range of word synonym similarity is [0,1]. (2) Calculation formula for vocabulary synonym similarity judgment value In formula (2), the synonymous similarity of words is calculated by formula (1), and the synonymous similarity judgment value is 1 or 0, where 1 means that the word i and vocabulary j is a synonym, 0 means vocabulary i and vocabulary j Not a synonym; (3) Formula for constructing vocabulary order vector Vocabulary order vector = [v1, v2, v i ,…,v n ], i = 1, 2, ..., n, n is the total number of different words in the English text T to be detected, In formula (3), T is the English text to be detected, w i is the i-th word in the vocabulary set of the English text T to be detected, and the word u is calculated by formula (2) to determine whether it is the word w i Synonyms of v i It is the vocabulary w i or vocabulary u i The position number in the English text T to be detected, position (w i , T) is the vocabulary w i The position number in the English text to be detected T, position (u, T) is the position number of the word u in the English text to be detected T; (4) Calculation formula for word order similarity In formula (4), the vocabulary order vector i and word order vectors j It is calculated by formula (3), and the range of word order similarity is [0, 1]; (5) Calculation formula for word order similarity judgment results In formula (5), the word order similarity judgment result is calculated by formula (4), and the word order similarity judgment result is: similar or dissimilar.
3. The method for detecting similarity in word order of English text according to claim 1, wherein: The processing steps of the English text word order similarity detection method are as follows: P101 starts; P102 reads the English text to be tested; P103 pre-processes the English text to be tested, including sentence segmentation, generating sentence word chains, and de-stemming different forms of the same word in the word chain; P104 constructs a sentence pair set, matching the sentences in the English text to be detected sentence by sentence in order to construct a sentence pair set of the English text to be detected; P105 generates a joint vocabulary set, performs set operation processing on each sentence in the sentence pair set, and generates a joint vocabulary set of all non-repeated words in the sentence pair; P106 constructs an initial lexical order vector, generating a zero vector with the same dimension as the joint word set for each sentence in the sentence pair as the initial lexical order vector; P107 fills the lexical order vector and analyzes the sentence pair to be tested. For words in the joint word set that exist in the current sentence, directly extract their position numbers using formula (3); for words in the joint word set that do not exist in the current sentence, find their synonyms based on formulas (1) and (2). Then, use formula (3) to extract the position number of the synonym and fill it into the corresponding position of the lexical order vector of the corresponding sentence. P108 Lexical order similarity determination, based on the constructed lexical order vector, first use formula (4) to calculate the lexical order similarity, and then use formula (5) to determine whether the lexical order is similar or dissimilar based on the obtained lexical order similarity; End of P109.