Tibetan sentence similarity calculation method based on feature fusion

The Tibetan sentences are processed through feature fusion method, and the MPCNN and Bert model combined with Tibetan grammar rules are used to solve the problem of Tibetan sentence similarity calculation, improve the calculation accuracy and model representation ability, and improve the performance of related natural language processing tasks.

CN120471062APending Publication Date: 2025-08-12TIBET UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510573101.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-12

Smart Images

  • Figure CN120471062A_ABST
    Figure CN120471062A_ABST
Patent Text Reader

Abstract

The invention discloses a Tibetan sentence similarity calculation method based on feature fusion. The Tibetan sentence similarity calculation method comprises the following steps of S1, performing boundary recognition on Tibetan sentences; s2, preprocessing the Tibetan sentences recognized in the step S1, including distinguishing the sentences according to minimum semantic unit word granularity, and expressing the sentences by using trained word vectors to obtain a matrix form of the sentences; s3, constructing a sentence representation model to analyze the sentences to obtain sentence vectors; s4, calculating a similarity value based on the obtained sentence vector by using a sentence similarity calculation model; and S5, outputting the similarity. Compared with the prior art, the Tibetan sentence similarity calculation method based on the feature fusion has the advantage that the Tibetan sentence similarity calculation method based on the Tibetan application is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, and in particular to a Tibetan sentence similarity calculation method based on feature fusion. Background Art

[0002] With the rapid development of science and technology today, the Internet has become an indispensable part of people's daily lives. People are accustomed to using the Internet to obtain the information they need. However, since digital information is increasing exponentially year by year, a series of Tibetan information processing problems that need to be solved urgently have also arisen in the process of using computers and the Internet.

[0003] In the field of natural language processing, especially in Chinese information processing, the calculation of sentence similarity is a basic and core research topic. It has long been a hot topic and difficulty in people's research, and plays a very important role in various fields of natural language processing.

[0004] Sentence similarity generally refers to the degree of semantic similarity between texts and is widely used in various fields of natural language processing tasks. In the field of machine translation, it serves as an evaluation criterion for translation accuracy, helping to improve the performance of machine translation systems. In the field of search engines, it can be used to measure the similarity between the searched text and the retrieved text. In the field of intelligent question answering, it can be used to assess the semantic match between the input question and the output answer, effectively improving the efficiency of intelligent question answering systems and reducing the intervention rate of manual customer service. In the field of text plagiarism detection, similarity calculation can detect the degree of plagiarism between two texts, further optimizing the academic atmosphere. In the field of text clustering, similarity thresholds can be used as clustering criteria. In automatic summarization, sentence similarity can reflect the degree to which local information fits the topic, thereby improving the quality of text summarization systems and user experience.

[0005] Therefore, strengthening the systematic research on Tibetan sentence semantic similarity algorithms is the cornerstone for the further development of Tibetan information processing technology and plays a crucial role in subsequent information processing. It also has very important application value and academic research value. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to overcome the above technical defects and provide a Tibetan sentence similarity calculation method based on feature fusion based on Tibetan language application.

[0007] To solve the above technical problems, the present invention provides a technical solution: a method for calculating Tibetan sentence similarity based on feature fusion, comprising the following steps:

[0008] S1: perform boundary recognition on Tibetan sentences;

[0009] S2: Preprocess the Tibetan sentences recognized in S1, including dividing the sentences into word granularity based on the smallest semantic unit, and representing them with trained word vectors to obtain the matrix form of the sentences;

[0010] S3: Build a sentence representation model to parse the sentence and obtain the sentence vector;

[0011] S4: Use the sentence similarity calculation model to calculate the similarity value based on the obtained sentence vector;

[0012] S5: Output similarity.

[0013] Preferably, the boundary identification in S1 includes locating Tibetan language features as follows: TSG = {ΑGT, CM, PAT, VERB}, where:

[0014] TSG represents the Tibetan genitive sentence, AGT represents the agent, CM is the Tibetan genitive particle, PAT is the patient object, and VERB is the verb.

[0015] Preferably, the boundary identification in S1 includes establishing a statistical model based on feature parameters extracted from Tibetan sentences, and performing boundary identification on the Tibetan sentences based on the established statistical model;

[0016] The feature parameters include sentence boundary punctuation marks, non-ending words, ending words and special words.

[0017] Preferably, the preprocessing in S2 includes word segmentation, construction of semantic relationship network, Tibetan sentence dependency structure information representation and word vector representation.

[0018] Preferably, the word segmentation includes summarizing rules based on non-Tibetan character segmentation errors, Tibetan agglutinative word recognition errors, stop word segmentation errors, and unregistered word segmentation errors in the conditional random field word segmentation results, and reprocessing and outputting word segmentation results based on the rules;

[0019] The construction of the semantic relationship network includes integrating different forms of verbs, honorifics and rhetoric to construct a knowledge base;

[0020] The Tibetan sentence dependency structure information representation includes determining Tibetan word classes and dependency relationships;

[0021] The word vector representation is represented using the Bert model.

[0022] Preferably, the sentence representation model in S3 uses an MPCNN model to parse sentences based on Tibetan sentence dependency structure information.

[0023] Preferably, the sentence similarity calculation model includes similarity calculation of the full convolution layer and the single convolution layer, and the obtained similarity vectors are spliced into one-dimensional data and input into the fully connected layer to connect the softmax to obtain the probability value of the category to which the similarity value belongs.

[0024] The advantages of the present invention over the existing technology are: the technical method of the present invention targets the characteristics of Tibetan sentences, including boundary identification and preprocessing combined with Tibetan grammatical rules, which facilitates capturing the deep semantic associations of sentences, enhances the model's ability to represent Tibetan sentences, and increases the accuracy of analysis results. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is the flowchart of Tibetan sentence boundary recognition.

[0026] Figure 2 It is a flowchart for constructing dependency syntax.

[0027] Figure 3 is a dependency syntax tree.

[0028] Figure 4 It is the sentence model itself.

[0029] Figure 5 It is a sentence structure information model.

[0030] Figure 6 It is a schematic flow chart of the method of the present invention. DETAILED DESCRIPTION

[0031] The present invention will be described in further detail below with reference to the accompanying drawings.

[0032] This paper uses the MPCNN model to represent Tibetan sentences by incorporating dependency structure information. It also considers the Tibetan word semantic network to calculate the semantic similarity of Tibetan sentences. The Tibetan sentence semantic similarity calculation is divided into four steps: sentence boundary identification, sentence preprocessing (word segmentation, semantic relationship network, Tibetan sentence dependency structure information, and word vector representation), sentence representation model training, and sentence similarity calculation.

[0033] Tibetan sentence boundary recognition:

[0034] Since a Tibetan sentence is a string of characters, it is composed of some content words and function words with different meanings. The word order structure of Tibetan sentences belongs to the SOV type, that is, the word order structure of {subject + object + predicate}. The structure of Tibetan sentences is based on verbs, and uses various conjunctions to connect words to form sentences. For example, the main components of Tibetan genitive sentences are: TSG={AGT,CM,PAT,VERB}, where TSG represents Tibetan genitive sentences, and AGT represents the agent (noun, pronoun, noun phrase). CM is the Tibetan genitive particle, that is PAT represents the object and VERB represents the verb.

[0035] Tibetan sentence boundary recognition is a key technology in the early stages of calculating sentence similarity. The integrity of the sentence directly affects the final similarity calculation results.

[0036] The present invention adopts a statistical method to extract feature parameters from a large number of sentences, establishes a statistical model, and uses the statistical model to perform sentence boundary recognition as shown in Table 1.

[0037] Table 1 Tibetan sentence boundary feature parameters

[0038]

[0039] The difficulty in identifying sentence boundaries lies in complex sentences. Complex sentences are composed of two or more simple sentences that are related in meaning but do not constitute sentence components in structure. Simple sentences refer to sentences that are not fully expressed.

[0040] Usually these simple sentences are connected with conjunctions to express complete Tibetan complex sentences. The simple sentences in the complex sentences have relationships such as parallelism, sequence, progression, explanation and selection.

[0041] If the syllable to the left of the hammer sign is The syllable of exists in the Tibetan sentence boundary feature parameter table, indicating that the sentence is a complete sentence. On the contrary, it indicates that the sentence has not ended. The specific process is as follows Figure 1 shown.

[0042] Sentence preprocessing:

[0043] The sentences are divided into the smallest semantic unit word granularity, and then represented by the trained word vector to obtain the matrix form of the sentence

[0044] Word segmentation is the most basic technology in natural language processing, which represents sentences as the most basic word sequence. There are no obvious separators between words, which makes Tibetan word segmentation more difficult.

[0045] In the present invention, the conditional random field is first used to segment the corpus, and then the errors of the segmentation results are analyzed, and the rules are summarized, and the segmentation results are processed again using the rules;

[0046] Specifically, rules are summarized for problems such as non-Tibetan character segmentation errors, Tibetan agglutinative word recognition errors, stop word segmentation errors, and unregistered word segmentation errors in the word segmentation results based on conditional random fields. The rules are then used to reprocess the word segmentation results to obtain the final word segmentation results.

[0047] Building a semantic relationship network:

[0048] Since function words and auxiliary words in Tibetan text are useless for general text mining tasks, but they greatly affect the vector representation of words in the text, they need to be removed. At the same time, many verbs in Tibetan have tense changes (1:3), rhetoric (1:N) and honorifics (1:1), and it is necessary to integrate these different forms of words to build a knowledge base, as shown in Table 2-4 below.

[0049] Table 2 Tibetan verb knowledge base

[0050]

[0051] Table 3 Tibetan vocabulary knowledge base

[0052]

[0053]

[0054] Table 4 Tibetan honorifics

[0055]

[0056] Tibetan sentence dependency structure information representation:

[0057] Tibetan sentence structure feature information is the Tibetan syntactic tree library that has been annotated with dependency syntax. To perform dependency syntax annotating on it, we must first determine the Tibetan word classes and dependency relationships. The present invention adopts the dependency syntax tags in Zhaxijia's "Analysis of the Theory and Method of Tibetan Dependency Tree Library Construction". In this tag, ADV is the adverbial-predicate relationship, RAD is the post-attachment, ls is the action case, SBV is the subject-predicate relationship, RAD is the post-attachment, QUN is the quantity-attributive-predicate relationship, VOB is the object involved, ic is the clause core, cn is the linking conjunction, ADV is the adverbial-predicate relationship, ld is the same-aspect case, LNF is the object, RLD is the scope, ls is the action case, ATV is the agent, RLD is the scope, RLD is the scope, PVT is the object involved, ic is the clause core, cn is the linking conjunction, RLD is the same-aspect case of the predicate, ld is the same-aspect case, and HED is the core of the whole sentence. The specific dependency syntax construction process and dependency syntax tree are as follows: Figure 2-3 shown.

[0058] Word vector representation:

[0059] This paper adopts the Bert model. Because the model is based on the Transformer architecture, it has powerful feature extraction capabilities and can learn complex syntactic and semantic relationships between words and capture long-range dependencies, thereby better understanding the overall meaning of the sentence.

[0060] Its word vector is dynamically generated based on the context of the entire sentence, which perfectly solves the problems of Tibetan polysemy and tense deformation (such as verb ) and honorific conversion (such as ) problem, using self-attention to capture global dependencies between words in both directions. This is particularly well-suited for the long-distance grammatical structure of Tibetan SOV (subject-object-verb) word order. BERT also uses the WordPiece segmentation method to break words into smaller subword units, effectively handling out-of-view (OOV) words. Even if the model has never seen a word, it can infer its meaning through the combination of its subword units.

[0061] Sentence Representation Model:

[0062] The MPCNN model of the present invention uses convolutional filters with multiple granularity window sizes, followed by various types of pooling methods, to parse sentences from multiple perspectives (Multi-perspective) and extract more semantic and syntactic structures of sentences. This model is different from the sentence representation of other models and can extract more features.

[0063] The present invention uses the MPCNN model to add Tibetan sentence dependency structure information to further summarize the semantic information of the sentence. The specific sentence representation model is as follows Figure 4-5 shown.

[0064] Sentence similarity calculation model:

[0065] The sentence representation model generates sentence vectors, and the similarity between the two sentence vectors is calculated as the final similarity value of the two sentences. The MPCNN model, which incorporates sentence structure information, performs multiple convolutions and corresponding pooling methods on the sentences. To make sentence vector comparison and calculation more effective, the commonly used sentence similarity calculation methods—cosine similarity, Euclidean distance, and Manhattan distance—are combined to complete the calculation of sentence vectors. Euclidean distance measures the absolute distance between points in space and is directly related to the coordinates of each point. Cosine distance measures the angle between spatial vectors and reflects differences in direction rather than position. Manhattan distance is the sum of the north-south distance between two points and the east-west distance.

[0066] Sent1(x,y)={cos(x,y),Euclid(x,y),Manha(x,y)} (2.4)

[0067] Sent2(x,y)={cos(x,y),Euclid(x,y)} (2.5)

[0068] The sentence-level full convolutional layer uses formula (2.4) for similarity measurement. The sentence itself is evaluated using three distance calculation formulas, resulting in a 3×Filernum1 similarity, which accounts for the three pooling methods used. The sentence structure information is measured using formula (2.5) for similarity measurement in the single convolutional layer. The sentence structure information is evaluated using two distance calculation formulas, resulting in a 2×Filernum2 similarity, which accounts for the two pooling methods used.

[0069] After calculating the similarity between the full convolutional layer and the single convolutional layer, we obtain the corresponding similarity vectors. The two similarity values are concatenated into one-dimensional data and fed into the fully connected layer, where they are connected using the softmax function to obtain the probability value of the category to which the similarity value belongs.

[0070] The contents not described in detail in this specification belong to the prior art known to those skilled in the art.

[0071] The basic principles, main features and advantages of the present invention are shown and described above. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention, and such changes and modifications fall within the scope of the invention as claimed.

Claims

1. A Tibetan sentence similarity calculation method based on feature fusion, characterized by: The following steps are involved: S1: perform boundary recognition on Tibetan sentences; S2: Preprocess the Tibetan sentences recognized in S1, including dividing the sentences into word granularities based on the smallest semantic unit, and representing them with trained word vectors to obtain the matrix form of the sentences. S3: Build a sentence representation model to parse the sentence and obtain the sentence vector; S4: Use the sentence similarity calculation model to calculate the similarity value based on the obtained sentence vector; S5: Output similarity.

2. The method for calculating Tibetan sentence similarity based on feature fusion according to claim 1, characterized in that: The boundary identification in S1 includes locating Tibetan language features as follows: TSG = {ΑGT, CM, PAT, VERB}, where: TSG represents the Tibetan genitive sentence, AGT represents the agent, CM is the Tibetan genitive particle, PAT is the patient object, and VERB is the verb.

3. The method for calculating Tibetan sentence similarity based on feature fusion according to claim 2, characterized in that: The boundary identification in S1 includes establishing a statistical model based on feature parameters extracted from Tibetan sentences, and identifying the boundaries of Tibetan sentences based on the established statistical model; The feature parameters include sentence boundary punctuation marks, non-ending words, ending words and special words.

4. The method for calculating Tibetan sentence similarity based on feature fusion according to claim 1, wherein: The preprocessing in S2 includes word segmentation, building a semantic relationship network, representing Tibetan sentence dependency structure information, and representing word vectors.

5. The method for calculating Tibetan sentence similarity based on feature fusion according to claim 4, characterized in that: The word segmentation includes summarizing rules based on non-Tibetan character segmentation errors, Tibetan agglutinative word recognition errors, stop word segmentation errors, and unregistered word segmentation errors in the conditional random field word segmentation results, and reprocessing and outputting word segmentation results based on the rules; The construction of the semantic relationship network includes integrating different forms of verbs, honorifics and rhetoric to construct a knowledge base; The Tibetan sentence dependency structure information representation includes determining Tibetan word classes and dependency relationships; The word vector representation is represented using the Bert model.

6. The method for calculating Tibetan sentence similarity based on feature fusion according to claim 1 or 5, characterized in that: The sentence representation model in S3 uses the MPCNN model to parse sentences based on Tibetan sentence dependency structure information.

7. The method for calculating Tibetan sentence similarity based on feature fusion according to claim 6, characterized in that: The sentence similarity calculation model includes similarity calculation of the full convolution layer and the single convolution layer. The obtained similarity vectors are spliced into one-dimensional data and input into the fully connected layer to connect the softmax to obtain the probability value of the category to which the similarity value belongs.

Citation Information

Cited By

  • Low-resource language word segmentation model training and cross-model word list migration injection method

    CN122197877A

  • A method for low-resource language word segmentation model training and cross-model vocabulary table migration injection

    CN122197877B