Text similarity calculation method based on language fragment feature items

By using verbs as feature items to build a semantic environment, the problem that traditional VSM methods cannot fully express text semantics is solved, and a higher accuracy of text similarity calculation is achieved, achieving a score accuracy of 85.24%.

CN120146024APending Publication Date: 2025-06-13WUXI ELECTRICAL & HIGHER VOCATIONAL SCHOOLS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510242245.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the traditional VSM method, using keywords as feature items cannot fully express text semantics, resulting in insufficient accuracy of text similarity calculation.

Method used

The semantic environment is constructed through the collection of verbs, reflecting the semantic connection between text contexts, using the TF-IDF method to calculate the weight of the feature term, and using the vector angle cosine method to calculate the text similarity.

Benefits of technology

The accuracy of text similarity calculation was improved. The experimental results showed that in the computer grading score of subjective questions in the test paper, the accuracy rate reached 85.24%, which was better than the traditional method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146024A_ABST
    Figure CN120146024A_ABST
Patent Text Reader

Abstract

The invention provides a text similarity calculation method based on phrase feature items, and aims to solve the problem that text semantic information cannot be completely reflected when keywords are adopted as feature items in a traditional vector space model (VSM) method and improve the text similarity calculation accuracy. The method comprises the following three steps: defining and constructing language slices, performing formalized representation on a text by using the language slices, and performing text similarity calculation. Specifically, the method comprises the following steps: combining two words in a text according to a grammar rule to form candidate language slices; calculating the correlation of the two words by using the point mutual information amount, and screening out language slices which accord with a threshold value as feature items; a TF-IDF method is adopted to calculate the weight of the feature item, and a vector included angle cosine method is adopted to calculate the text similarity. Experiments show that when the method is used for computer test paper judgment and scoring of test paper subjective questions, the accuracy rate can reach 85.24%, the method is remarkably superior to a traditional method adopting keywords as feature items, and an optimization scheme is provided for text similarity calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text classification in natural language processing, and in particular to a method for calculating text similarity based on language piece feature items. Background Art

[0002] Natural language processing is an important research direction in the fields of computer science and artificial intelligence. Text classification technology is a part of it, and its task is to automatically determine the category to which a text belongs according to the text content under a given classification system. As the core of text classification research, text similarity calculation uses mathematical methods to calculate the similarity degree among multiple texts. Computer grading of subjective questions in examination papers is a key application of text classification technology in the field of education, which is of great significance for students' self-test homework, large-scale online examinations, etc. Existing systems such as Project Essay Grade, Educational Testing Service, and E-rater all adopt the VSM (Vector Space Model) method. Research shows that the VSM method is one of the effective text representation methods. However, the traditional VSM method usually uses keywords as vector feature items. This way of representing with discrete terms lacks the semantic correlation between words, resulting in the inability of feature items to fully express the text semantics. Summary of the Invention

[0003] In view of the problems in the traditional VSM method, the present application provides a method for calculating text similarity based on language piece feature items, and innovatively proposes to use "language pieces" as feature items. A semantic environment is constructed through a set of language pieces to reflect the semantic connection and potential conceptual structure (such as the co-occurrence relationship between words, etc.) between the text context. The composition structure of a language piece is "main word + affiliated word", that is, two words that conform to the binary grammar rule in the text are combined together to form a candidate language piece; a domain corpus is selected, and the pointwise mutual information method is used to calculate the correlation between the two words in the candidate language piece, and the combinations that meet the threshold are screened out to form language pieces. Using language pieces as feature items, the TF-IDF (Term Frequency–Inverse Document Frequency) method is used to calculate the weight of the feature items, and the cosine method of vector included angle is used to calculate the text similarity. Experiments show that when this method is used for computer grading of subjective questions in examination papers, the accuracy rate can reach 85.24%, which is significantly better than the traditional method using keywords as feature items, providing an optimized solution for text similarity calculation.

[0004] The technical solution adopted by the present invention is as follows:

[0005] A text similarity calculation method based on language piece feature items, which includes the following steps:

[0006] Step S1: Define and construct language pieces;

[0007] Step S2: Formalize the text using language pieces;

[0008] Step S3: Calculate the text similarity.

[0009] The construction method of the language pieces is as follows: A language piece is composed of two words in the text. One of the words has a part of speech of noun, verb, adjective or adverb, which is called the main word; the other is a word in the text that can be orderly combined with the main word according to the binary grammar rule, which is called the affiliated word. The main word and the affiliated word form a language piece;

[0010] Take the language piece as a feature item of the text vector, calculate the feature item weight using the TF-IDF method, and calculate the similarity between texts using the vector cosine angle method.

[0011] The step S1 first preprocesses the text, and its operation steps are as follows: Optionally, use an existing word segmentation system (such as NLPIR-ICTCLAS of the Chinese Academy of Sciences) to perform operations such as word segmentation, part-of-speech tagging, removing special punctuation marks and useless modal particles, filtering high-frequency function words, removing stop words, merging synonyms and near-synonyms on the text to obtain the word set of the text.

[0012] Furthermore, the definition of the language piece in step S1:

[0013] Let Doc represent a text, which is an ordered sequence composed of a series of sentences, that is, Doc = (S 1 ; S 2 ; …; S p ), where p ∈ N (N is the set of natural numbers) is used to represent the number of sentences in Doc. For any k ∈ {1, 2, …, p}, use Sk to represent the kth sentence in Doc. After preprocessing Sk, the word set W of Sk can be obtained, and there is W = (w 1 ; w 2 ; …; w m ), where m ∈ N represents the number of words in Sk.

[0014] Use w a to represent a word in W with a part of speech of noun, verb, adjective or adverb, which is called the main word; use w b to represent another word in W that can be orderly combined with w a according to the binary grammar rule, which is called the affiliated word. The language piece is a binary ordered combination composed of w a and w b , and wa and w b meets the relevance requirement. In the language segment, w a and w b The combination order of is determined according to the bigram rules and can be w a in the front or w a in the back.

[0015] Thus, using t to represent a language segment, we have:

[0016] where 1 ≤ a ≤ m, 1 ≤ b ≤ m, and a, b ∈ N.

[0017] All such language segments form the language segment set T of sentence S k , that is, T = {t 1 ; t 2 ;...}. The text language segment set is the union of the sentence language segment sets.

[0018] The language segment contains the semantic fragment of the sentence, and the language segment set covers the general semantics of the text and follows the following two conventions: (1) For any two language segments t i and t j (i ≠ j) in the language segment set T, it satisfies t i ≠ t j , that is, the language segments in the language segment set are mutually different; (2) In the sentence language segment set T, there is no inherent order relationship such that for any two language segments t i and t j , it can be clearly said that t i is before or after t j , that is, the language segments in the language segment set have no order relationship.

[0019] Furthermore, the construction of the language segment in step S1 adopts a method combining rules and statistics. The specific steps are: Select two words from the word set to form a language segment. One of the words is a noun, verb, adjective, or adverb, which is called the main word; the other word can be combined with the main word in an orderly manner according to the bigram rules and is called the adjunct word. Combining the main word and the adjunct word gives the candidate language segment set. Next, based on the domain corpus, the point mutual information is used to calculate the relevance between the main word and the adjunct word. When the relevance between the two meets the preset threshold, the combination of these two words forms a language segment.

[0020] Furthermore, the bigram rules based on which the language segment is constructed can be formally expressed as: Let A 1 , A 2 respectively represent the part-of-speech of the main word and the adjunct word, and t represents the one formed by A 1 , A 2A phrase formed by combining the corresponding words in a specific order; then, the bigram grammar rule can be expressed as (A 1 , A 2 ) → t, where the part-of-speech A 1 , A 2 is taken from a predefined set of parts of speech (such as nouns, verbs, adjectives, adverbs, etc.). The bigram grammar rule emphasizes the mapping relationship between the combination of words based on their parts of speech and the generation of phrases. A set of multiple bigram grammar rules is called a bigram grammar rule set. The bigram grammar rule set comprehensively covers the rules for generating phrases by various combinations of parts of speech under a specific language analysis framework, providing a systematic specification for the construction of phrases.

[0021] Furthermore, in step S1, the Pointwise Mutual Information method is used to calculate the correlation between two words, and the combination of words that meets the pointwise mutual information threshold is used as a phrase. The calculation formula is:

[0022]

[0023] where C(w 1 , w 2 ) represents the number of times the words w 1 and w 2 co-occur in order in the corpus, C(w 1 ) represents the number of times the word w 1 appears, C(w 2 ) represents the number of times the word w 2 appears, and H represents the total number of words in the corpus.

[0024] Furthermore, in step S2, the formal representation of the text adopts the VSM method based on phrase feature terms, that is, a text vector is represented by a set of phrase feature terms and their weights, and the text similarity can be calculated by means of the mathematical relationship between vectors.

[0025] The text vector can be represented as:

[0026] D = {We 1 ; We 2 ; …; We i ; …; We n};

[0027] where D represents the text vector, We i (1 ≤ i ≤ n) represents the weight of the phrase feature term T i (1 ≤ i ≤ n), and both i and n are natural numbers (i, n ∈ N).

[0028] The steps for calculating the text similarity in step S3 are: The similarity degree of the text is represented by Sim(Di , D j ) indicates that the TF-IDF method is used to calculate the weights of the feature terms of the language segments, and the cosine of the vector angle method is used to calculate the text similarity. The formula is as follows:

[0029]

[0030] Among them, D i , D j are two texts for which the similarity is to be calculated, and d ik and d jk are the weights of the feature terms of D i and D j . Through the calculation of the formula, the cosine of their included angle is obtained. There is 0 ≤ Sim(D i , D j ) ≤ 1. When the calculated value is larger, the text similarity is larger; on the contrary, the text similarity is smaller. When Sim(D i , D j ) = 0, it indicates that the text D i and D j are completely dissimilar. When Sim(D i , D j ) = 1, it indicates that the text D i and D j are completely similar.

[0031] The beneficial effects of the present invention compared with the prior art are as follows:

[0032] The present invention creatively proposes to use "language segments" as feature terms for calculating text similarity in the vector space model (VSM). By constructing a semantic environment through the set of language segments, it more completely reflects the semantic connection and potential conceptual structure between the text context, and more accurately reflects the overall semantic features of the text. Experiments have proved that the accuracy rate of this method for computer grading of subjective questions in test papers can reach 85.24%, which is better than the keyword feature term method, and provides an optimized solution for text similarity calculation. Description of the Drawings

[0033] Figure 1 is a flowchart of the steps for calculating text similarity in the present invention;

[0034] Figure 2 is a schematic diagram of the language segment structure in the present invention;

[0035] Figure 3 is a schematic diagram of the process for calculating text similarity in the present invention. Detailed Embodiments

[0036] The following combines the drawings to illustrate the detailed embodiments of the present invention.

[0037] AsFigures 1 to 3 As shown in the figure, the present invention proposes a method for calculating text similarity based on language piece feature items, which mainly includes the following three steps:

[0038] Step S1: Define and construct language pieces;

[0039] Step S2: Formalize the text using language pieces;

[0040] Step S3: Calculate the text similarity.

[0041] In an embodiment of the present invention, the text preprocessing work uses the NLPIR-ICTCLAS word segmentation system.

[0042] Table 1 is the set of word part-of-speech marking symbols used in this application.

[0043]

[0044]

[0045] Table 1 Set of word part-of-speech marking symbols used in this application

[0046] In an embodiment of the present invention, the specific implementation manner of Step S1 is as follows:

[0047] For any sentence S in the text Doc k , after its preprocessing, the word set W of S can be obtained, and W = (w k ; w 1 ;...; w 2 ;...; w m ). Arbitrarily select a word in W whose part of speech is a noun, verb, adjective, or adverb as the main word; arbitrarily select another word in W, when it can be combined with the main word according to the binary grammar rule (A 1 , A 2 ) → t, then a candidate language piece is obtained. The set of all candidate language pieces is the candidate language piece set T, and there is T = {t 1 ; t 2 ;...}. The text candidate language piece set is the union of the sentence candidate language piece sets. The language pieces in the candidate language piece set are mutually different, and there is no order relationship among the language pieces.

[0048] This application constructs the binary grammar rule set shown in Table 2 centered on nouns, verbs, adjectives, and adverbs according to Chinese grammar rules, and these binary grammar rules follow certain lexical collocation norms. When determining the word combinations that make up the candidate set of language pieces, these rules are used according to the part of speech. Once a word combination that meets the rules is found, it is extracted to form a candidate language piece and recorded in the candidate language piece set.

[0049] Table 2 shows the bigram rule set for constructing language fragments, with a total of 11 types and 282 rules, which are formulated according to Chinese grammar.

[0050]

[0051]

[0052] Table 2 Bigram Rule Set

[0053] In this application, the bigram rule for word combination is expressed in the form of (A 1 , A 2 ) → t. For example: if A 1 = a (adjective) and A 2 = n (noun), then according to the "a + n" rule in the (adjective, noun) type, a candidate language fragment can be combined, such as (huge, space), where "huge" is the dependent word and "space" is the main word.

[0054] Since there are numerous word combinations in the text that satisfy the above bigram rules, after the processing of the above steps, the number of candidate language fragments in the candidate language fragment set is extremely large. However, some of these candidate language fragments are not reasonable. Combinations like "(RMB, eating)" obviously lack logical connection, so it is necessary to filter out the unreasonable candidate language fragments.

[0055] Given that there must be a strong correlation (Relevance) between words with definite semantics in the text, this application uses an index for measuring the mutual dependence relationship between words - Pointwise Mutual Information (PMI) to calculate the degree of correlation between two words, and then filters out the language fragments that meet the PMI threshold condition for subsequent processing. The formula for calculating the Pointwise Mutual Information is:

[0056]

[0057] where C(w 1 , w 2 ) represents the number of times the words w 1 and w 2 co-occur in order in the corpus, C(w 1 ) represents the number of times the word w 1 appears, C(w 2 ) represents the number of times the word w 2 appears, and H represents the total number of words in the corpus. By counting the number of times words appear, the number of times words co-occur, and the total number of words in the text, the Pointwise Mutual Information of word combinations can be calculated using the above formula.

[0058] The present invention uses the publicly available corpus of Fudan University and conducts research in the field of "Computer Technology" therein. One of its purposes is to first conduct in-depth research in the text language environment of a specific scope to fully verify the effectiveness of the present invention, and after the good effect is confirmed, then promote it; the other is to closely conform to the actual application scenario and be able to effectively improve the correctness and operation efficiency of the system.

[0059] Determination of the point mutual information threshold:

[0060] Based on the point mutual information calculation formula, the point mutual information of each word combination in the candidate phrase set can be calculated. The next key step is to determine a point mutual information threshold suitable for constructing phrases, and based on this, reasonable phrases are screened out. The point mutual information is an index to measure the correlation between two events. The larger its value, the stronger the correlation between the two events; the smaller the value, the weaker the correlation. On the basis of ensuring good statistical calculation results, this application comprehensively considers the factor of logical rationality and sets a suitable point mutual information threshold.

[0061] A phrase is a semantic association entity based on the combination of rules and statistics. After applying the bigram rule to the corpus in the field of "Computer Technology", a total of 233,756 candidate phrases are generated, and 79,721 reasonable phrases are screened out. Most of these reasonable phrases are distributed in the range where the point mutual information (PMI) is greater than or equal to 6. Therefore, the point mutual information threshold is determined to be PMI≥6.

[0062] In an embodiment of the present invention, the weight of the feature term adopts the relative frequency of the occurrence of the phrase, that is, the TF-IDF method.

[0063] Among them, TF = the number of occurrences of a certain phrase in the text / the total number of phrases in the text, IDF = log10 (the total number of texts / the number of texts in which a certain phrase appears), and the weight of a certain phrase = TF×IDF. For example, there are 1000 texts, among which 10 texts contain the phrase (memory, access), and the phrase (memory, access) appears 26 times in a text with 1625 phrases. Then the weight of the phrase (memory, access) in this text:

[0064]

[0065] In an embodiment of the present invention, the specific implementation method of text similarity calculation is: the similarity degree of the text is represented by Sim(D i , D j ), D i represents the source text, and D j represents the target text. The similarity between texts only depends on the explicit semantics of the texts being compared, and does not consider their implicit extensions, metaphorical meanings, etc. In the preferred embodiment of this application, cosine similarity is adopted, and the formula is as follows:

[0066]

[0067] Among them, D i and D j are two texts for which the similarity is to be calculated, and d ik and d jk are the feature item weights of D i and D j . Through the calculation of the formula, their cosine of the included angle is obtained, and 0 ≤ Sim(D i j , D j ) ≤ 1. When the calculated value is larger, the text similarity is larger; on the contrary, the text similarity is smaller. When Sim(D i , D j ) = 0, it means that the text D i and D j are completely dissimilar. When Sim(D i , D j ) = 1, it means that the text D i and D j are completely similar.

[0068] The following uses an example to illustrate the implementation process of the text similarity calculation method based on language segment feature items. The texts for which the similarity is to be calculated are Doc 1 and Doc 2 .

[0069] Doc 1 : "Because the speed at which the CPU accesses the cache is greater than the speed of accessing the memory, the processing speed is improved."

[0070] Doc 2 : "Setting up a cache improves the processing speed because the speed at which the CPU accesses the cache is greater than the speed of accessing the memory."

[0071] (1) Preprocess the text to obtain its word set W.

[0072] The word set of Doc 1 is W 1 , and W 1 = {because / p; cpu / nx; access / v; cache / nx; speed / n; greater / v; access / vn; memory / n; speed / n; improve / v; processing / vn; speed / n};

[0073] The word set of Doc 2 is W 2 , and W 2={Set / v; Cache / n; Improve / v; Process / vn; Speed / n; Because / c; CPU / nx; Access / v; Cache / nx; Speed / n; Greater than / v; Access / vn; Memory / n; Speed / n}.

[0074] (2) By applying the bigram rule and calculating the point mutual information, construct the set T of text segments of the text.

[0075] Construct the set T of text segments of Doc 1 of Doc 1 , and there is T 1 ={(Speed, Memory); (CPU, Speed); (CPU, Memory); (Cache, Speed); (Cache, Memory); (Speed, Greater than); (Speed, Improve); (CPU, Access); (CPU, Improve); (Cache, Improve); (Speed, Access); (Speed, Process); (Memory, Process); (CPU, Process); (Cache, Access); (Cache, Process); (Access, Speed); (Access, Memory); (Greater than, Memory); (Access, Cache)};

[0076] Construct the set T of text segments of Doc 2 of Doc 2 , and there is T 2 ={(Cache, Speed); (Cache, Memory); (Speed, Memory); (Cache, CPU); (Cache, Cache); (Speed, CPU); (Speed, Cache); (CPU, Memory); (Cache, Memory); (Cache, Improve); (Cache, Is); (Cache, Access); (Speed, Access); (Speed, Greater than); (CPU, Access); (Cache, Process); (Cache, Access); (Improve, Speed); (Improve, Memory); (Access, Memory); (Greater than, Memory); (Improve, CPU); (Improve, Cache); (Access, Cache); (Set, Cache); (Set, Memory); (Process, Memory); (Set, CPU); (Set, Cache); (Process, CPU)}.

[0077] (3) Based on the corpus and the formula, calculate the weight of feature terms and the cosine value of the vector angle to obtain Sim(Doc 1 , Doc 2 ) = 0.8532. The full score of this question is 20 points, and the system grading score = full score × Sim(Doc 1 , Doc 2 ) = 17.06 points.

[0078] In summary, as a representation form of text vector feature items, the present invention uses a method combining statistics and rules to generate language segments. The specific steps are as follows: First, clarify the structure of the language segment as "main word + attached word", and combine two words in the text that meet the bigram grammar rules to obtain a candidate language segment set. Second, based on the domain corpus, calculate the correlation between the two words in the candidate language segments using point mutual information, and filter out the combinations whose correlation meets the preset threshold range to form a language segment set. After completing the construction of the language segments, use the TF-IDF method to calculate the weights of each language segment, and then use the cosine method of vector angle to calculate the similarity between texts.

[0079] In summary, as an application of the text similarity calculation method, the system can automatically complete the grading of students' answers with reference to the standard answers. This process requires the system to reach or approach the level of teachers' grading in key indicators such as grading accuracy, thereby effectively improving the quality of education and teaching services.

[0080] To verify the feasibility of this similarity calculation method, it is necessary to establish a hardware module and system verification based on this method. Figure 3 The following is the data flow diagram of the text similarity calculation and verification process in the present invention:

[0081] Effectiveness verification and result analysis:

[0082] The evaluation method of the present invention is: taking the teacher's manual grading as the benchmark, comparing the system grading with the manual grading to obtain the accuracy Pr of the system grading. The accuracy calculation formula is as follows:

[0083]

[0084] Among them, the standard score is the full score of the question, the system grading is the score of the question evaluated by the grading system, and the manual grading is the score of the question evaluated by the teacher's manual marking.

[0085] Experimental cases: Select the subjective question part of the "Operating System" course test paper of a certain undergraduate university in 2022 as experimental cases. There are 3 questions in total, with a full score of 35 points. The standard answers are the answers provided with the test paper, and the user answers are the answers of the candidates, with a total of 1000 copies. As a comparison with the traditional VSM method, keywords were also used as feature items for system grading in the experiment, and the keywords were determined by the teacher.

[0086] By analyzing the system operation situation, the results shown in Table 3 are obtained:

[0087]

[0088] Table 3 Distribution of accuracy of system calculation results

[0089] Experimental results: When using the language fragment feature item method for system grading and scoring, the accuracy of 51% of the experimental cases is above 90%, and the accuracy of 15% of the experimental cases is between 80% and 90%. The average accuracy of the system is 85.24%, and satisfactory results have been achieved in text similarity calculation. At the same time, in the experiment of using the keyword feature item method for system grading and scoring, the average accuracy of the system is 71.62%. Obviously, the language fragment feature item method is superior to the keyword feature item method.

[0090] In summary, the present invention proposes a method based on the vector space model, uses language fragments as feature items, calculates the weights of feature items by the TF-IDF method, and calculates the text similarity by the vector cosine angle method. In the verification experiment, the scores given by the system are sampled and compared with the teacher's scores, and the accuracy rate is used as an index to evaluate the effectiveness of the method. The results show that the accuracy rate of the method of the present invention reaches 85.24%, showing satisfactory results, and fully proving the reliability and effectiveness of the method in practical applications.

[0091] The above description is an explanation of the present invention, not a limitation of the invention. For the scope defined by the present invention, please refer to the claims. Within the protection scope of the present invention, any form of modification can be made.

Claims

1. A text similarity calculation method based on language fragment feature items, characterized by: The method comprises the following steps: Step S1: define and construct a language segment; Step S2: formalizing the text using phrases; Step S3: Calculate text similarity; The method for constructing the language fragment is as follows: a language fragment is composed of two words in the text, one of which is a noun, a verb, an adjective or an adverb, and is called a stem word; The other is the words in the text that can be combined with the main words in an orderly manner according to the bigram rules, which are called attached words. The main words and attached words constitute a language fragment; The language fragments are used as the feature items of the text vector, the TF-IDF method is used to calculate the feature item weights, and the vector angle cosine method is used to calculate the similarity between texts.

2. The text similarity calculation method based on language fragment feature items according to claim 1, characterized in that: In the step S1, the text is first preprocessed, including word segmentation, part-of-speech frequency tagging, removal of special punctuation marks and useless modal particles, filtering of high-frequency function words, elimination of stop words, merging of synonyms and antonyms, and obtaining a word set.

3. The text similarity calculation method based on language fragment feature items according to claim 1, characterized in that: Definition of the language fragment in step S1: Let Doc represent a text, which is an ordered sequence of sentences, that is, Doc = (S1; S2; ...; S p ), where p∈N (N is a natural number set), is used to represent the number of sentences in Doc; for any k∈{1, 2, …, p}, S k Represents the kth sentence in Doc. k After pretreatment, S k The word set W has W = (w1; w2; ...; w m ), where m∈N, represents S k The number of words in Use w a Indicates a word in W that is a noun, verb, adjective or adverb, which is called the stem word; use w b Indicates that another in W can be combined with w a Words that are combined in an orderly manner according to the rules of bigrams are called adjuncts; fragments include w a and w b The binary ordered combination of a and w b The order of combinations is determined according to the bigram rules, including w a Before or w a in the back; Therefore, using t to represent a fragment, we can get: Where 1≤a≤m, 1≤b≤m, and a, b∈N; All such fragments constitute the sentence S k The text fragment set T is T = {t1; t2; ...}; the text fragment set is the union of the sentence fragment set.

4. The text similarity calculation method based on language fragment feature items according to claim 1, characterized in that: The method of constructing the fragment in step S1 is to combine the main words and the attached words that meet the bigram rules in the text word set to obtain a candidate fragment set; Based on the domain corpus, the point mutual information method is used to calculate the correlation between the main words and the attached words. The words that meet the threshold range are the language fragments, which are used as the feature items of the text vector to obtain the language fragment feature item set.

5. The text similarity calculation method based on language fragment feature items according to claim 1, characterized in that: The bigram rule based on which the language fragment is constructed in step S1 can be formally expressed as follows: let A1 and A2 represent the parts of speech of the main word and the subordinate word respectively, and t represents the language fragment composed of the words corresponding to A1 and A2 in a specific order; then, the bigram rule can be expressed as (A1, A2)→t, where the parts of speech A1 and A2 are taken from a predefined part-of-speech set (such as nouns, verbs, adjectives, adverbs, etc.); the bigram rule emphasizes the mapping relationship between the combination of words based on parts of speech and the generation of language fragments; a set of multiple bigram rules is called a bigram rule set; the bigram rule set comprehensively covers the rules for generating language fragments by combining various parts of speech under a specific language analysis framework, and provides a systematic specification for the construction of language fragments.

6. The text similarity calculation method based on language fragment feature items according to claim 1, characterized in that: In step S1, the correlation between two words is calculated by using the point mutual information method, and when the point mutual information threshold is met, it is regarded as a language fragment; The point mutual information calculation formula is: Among them, C(w1,w2) represents the number of times words w1 and w2 co-occur in order in the corpus, C(w1) represents the number of times word w1 appears, C(w2) represents the number of times word w2 appears, and H represents the total number of words in the corpus.

7. The text similarity calculation method based on language fragment feature items according to claim 1, characterized in that: In the VSM based on the language feature item in step S2, the text vector can be expressed as: D={We1;We2;…;We i ;…;We n }; Where D represents the text vector, We i (1≤i≤n) represents the feature item T of the segment i (1≤i≤n), and i and n are both natural numbers (i, n∈N).

8. The text similarity calculation method based on language fragment feature items according to claim 1, characterized in that: The steps of the text similarity calculation method in step S3 are: the similarity of the text is calculated by Sim(D i , D j ) indicates that the TF-IDF method is used to calculate the feature item weights, and the vector angle cosine method is used to calculate the text similarity. The formula is: Among them, D i , D j are two texts whose similarity is to be calculated, d ik and d jk Yes D i and D j The characteristic item weights of ; through the calculation of the formula, we get the cosine of their angle, 0≤Sim(D i , D j )≤1, when the calculated value is larger, the text similarity is larger, otherwise the text similarity is smaller; when Sim(D i , D j )=0 means text D i With D j Not similar at all, Sim(D i , D j )=1 indicates text D i With D j Completely similar.