Multidisciplinary academic paper language translation system based on large language model
By designing a multidisciplinary academic paper language translation system that includes preprocessing modules, corpus, large language models and quality evaluation modules, the shortcomings of the existing system in corpus collection and model training are solved, and more accurate and professional academic paper translation effects are achieved.
Patent Information
- Application Number
- CN202510084303.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing multidisciplinary academic paper language translation system based on large language models has shortcomings in corpus collection and model training, which leads to the inability to accurately understand and translate professional terms and expressions in specific contexts when translating multidisciplinary papers.
A multidisciplinary academic paper language translation system based on large language models is designed, which includes preprocessing modules, corpus, large language model, translation module, user interface and quality assessment module. The system collects and labels bilingual control data in multidisciplinary fields by cleaning and standardizing academic papers, and uses large language models for translation parameters training and optimization.
The system can generate more accurate and professional academic paper translation results, improve the professionalism and accuracy of translation, meet the translation needs of users of different disciplines, and continuously optimize the translation quality through the quality assessment module.
Smart Images

Figure CN120012789A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of language processing, and in particular to a multidisciplinary academic paper language translation system based on a large language model. Background Art
[0002] In today's globalized academic exchange environment, the demand for language translation of multidisciplinary academic papers is growing. Traditional academic paper translation methods mainly rely on manual translation. Although this method has certain guarantees in terms of accuracy, it has many limitations. On the one hand, manual translation is inefficient. Faced with a large amount of academic literature, the translation speed is difficult to meet the needs of scientific researchers to quickly obtain information. On the other hand, the quality of manual translation is limited by the translator's professional knowledge background and language ability. For some highly specialized multidisciplinary papers, it is extremely difficult to require translators to be proficient in professional knowledge of multiple disciplines and multiple languages at the same time. With the development of artificial intelligence technology, translation systems based on large language models have gradually emerged. However, the existing multidisciplinary academic paper language translation systems based on large language models still have obvious shortcomings. In terms of corpus, the corpus collection of many systems is not comprehensive and professional enough, and lacks extensive coverage of academic papers in multidisciplinary fields. In particular, there is less paper data for some niche or emerging disciplines. This makes it impossible for the model to accurately understand and translate professional terms and expressions in specific contexts when translating papers in these fields. In terms of model training and optimization, some systems do not make full use of the metadata of academic papers and cannot effectively guide the model to consider discipline-specific terms and contexts, resulting in a significant reduction in the professionalism and accuracy of the translation results. Summary of the invention
[0003] In order to solve the above technical problems, the present invention is implemented by the following technical solutions: a multidisciplinary academic paper language translation system based on a large language model, the system comprising: The preprocessing module is used to clean the text and standardize the format of the input academic papers; Corpus, which stores bilingual comparative data of multidisciplinary academic papers; A large language model, connected to the corpus, trained and optimized translation parameters based on the bilingual comparison data; A translation module, for inputting the preprocessed academic papers into the large language model to generate a translation output in a target language; A user interface that allows users to upload academic papers, select source and target languages, and display translation results; The quality assessment module is used to evaluate the quality of the translation results and provide feedback to the large language model to optimize future translations. The quality assessment module includes a machine assessment submodule and a manual assessment submodule. The machine assessment submodule is used to automatically evaluate the accuracy and fluency of the translation, and the manual assessment submodule is used to collect user feedback and expert evaluation to facilitate subsequent optimization of translation quality.
[0004] Preferably, the preprocessing module comprises: Tokenizer, which is used to segment academic paper text into words or phrases; A part-of-speech tagger is used to tag the words or phrases after word segmentation, and to identify and tag the part of speech of each unit; Entity recognizers for identifying proper nouns and terms in academic papers.
[0005] Preferably, the process of the word segmenter segmenting the academic paper text into words or phrases includes: using word boundary rules to segment the continuous text sequence in the standard text into words or phrases, firstly collecting word boundary rules in natural language, the word boundary rules are formed based on the grammar, vocabulary and punctuation of the language, and obtaining general word boundary rules, for different professional fields involved in multidisciplinary academic papers, sorting out the corresponding word boundary related rules in the field, obtaining usage rules, and at the same time, referring to the vocabulary and terminology dictionary of the professional field, extracting the word composition and boundary information reflected therein, and supplementing them into the rule system, so as to more accurately segment the academic paper text, and jointly forming professional field specific rules based on usage rules, word composition and boundary information, wherein the word boundary rules include general word boundary rules and professional field specific rules, integrating the general word boundary rules and professional field specific rules, constructing a rule base, and storing the rule base in the form of text files and database tables, so as to facilitate subsequent rule matching and query in the word segmentation process, and performing formatting of the input standard text, removing the redundant blank characters that may exist in the text, and replacing line breaks, making the text into a unified format, and finally performing the word segmentation process. The table characters are converted or standardized according to the unified format requirements to ensure that the format of the text is neat, so as to facilitate the subsequent accurate judgment of word boundaries according to the rules and obtain the standard text format. Based on the standard text format, the special characters in the text are identified. For some special characters that do not participate in the composition of words but may affect the judgment of word boundaries, the special characters include special symbols that are not English letters, numbers, and Chinese characters. They are processed to obtain the preprocessed text after the special characters are processed. The processing process includes: if it is a special character that is part of a word, its special attributes are retained and marked. If it is a special character that simply serves as a separator or decoration, it is temporarily removed or converted into a space according to specific needs so as not to affect the judgment of word boundaries. Based on the preprocessed text, starting from the beginning of the text, it is scanned character by character in order, reading one character each time, and recording the current character position and the character sequence that has been scanned at the same time. According to the current character position and the character sequence that has been scanned, it is matched with the general word boundary rules in the rule library. If a space or punctuation mark that meets the general word boundary rules is encountered, the character sequence scanned previously is determined as a word or phrase unit, extracted and recorded; For character sequences that meet the specific rules of the professional field, check whether the character sequence matches the specific rules of the professional field. If they match, it is confirmed as a complete professional term and the character sequence is extracted as a word unit. When the scanned character sequence does not meet the general word boundary rules, that is, the word boundary cannot be determined temporarily, continue to scan the characters, continuously expand the character sequence range, and try to match the rules again until a matching boundary rule is found to determine the word unit. After scanning the entire text, the word or phrase units extracted according to the word boundary rules are recorded in sequence to form a word or phrase list. This list is the result of word segmentation of the original continuous text sequence. Samples are randomly selected from the word or phrase list after word segmentation and manually checked to see if there are any incorrect word segmentations that do not meet language habits or professional field requirements. If an error is found during manual sampling, the location and content of the error are recorded, and the character sequence corresponding to the erroneous part is re-segmented to obtain an error correction sequence, and the character sequence is re-checked manually until the character sequence is correct.
[0006] Preferably, the process of storing bilingual comparison data of multidisciplinary academic papers in the corpus includes: first, collecting original academic papers and their corresponding high-quality translations in the fields of medicine, physics, computer science, economics, and literature from academic databases, academic journal websites, digital resources of professional books, and academic resource libraries of universities. These papers should be representative, including classic literature, cutting-edge research reports, and academic works of different types and levels, to ensure the richness and diversity of the corpus. At the same time, collecting metadata of academic papers, including subject classification information, author information, publication journal information, keywords, and abstracts. These metadata will be used in subsequent translation. The process provides important context and subject-specific information references for large language models. The collected original and translated texts are formatted in a unified manner, and unnecessary special characters, garbled characters, and redundant blanks are removed to ensure the standardization and consistency of the texts. This is convenient for subsequent text processing and analysis to obtain standard texts. Natural language processing tools are then used to check and correct spelling errors, grammatical errors, and irregular punctuation in the standard texts to improve the quality and accuracy of the corpus. The standard texts are segmented to divide continuous text sequences into words or phrases to better construct text features for subsequent indexing features. Use bilingual alignment technology to align the original text and the translation at the sentence level to ensure that each source language sentence can accurately correspond to the target language sentence, establish a one-to-one bilingual corpus pair, and annotate the aligned bilingual corpus pair. The annotation content includes part of speech, named entities, and grammatical structure. The annotation content provides richer language feature information for large language models, which helps to improve the accuracy and quality of translation. Use a part-of-speech tagger based on a statistical model to annotate the part of speech. For the recognition of named entities, it is used to identify the names of institutions, people, and places in academic papers, and annotate the types to maintain consistency and accuracy during the translation process; use By using dependency syntactic analysis, the grammatical relationship between words in a sentence is analyzed, the grammatical structure is annotated, and grammatical guidance is provided for translation. The cleaned, preprocessed, aligned and annotated bilingual corpus data is stored in a database. An appropriate data storage structure is adopted, and the inverted index technology is used to establish an index relationship between a word and the document or sentence containing the word. An indexing mechanism is constructed based on keywords, subject classifications, and sentence similarity to index the text data in the corpus, so that the bilingual comparison data and metadata related to the text to be translated can be quickly retrieved during the translation process, thereby improving the data retrieval efficiency and the system response speed. The corpus is updated regularly, and newly published academic papers and their translations are collected and incorporated into the corpus in a timely manner to maintain the timeliness of the corpus and its coverage of the latest academic achievements. At the same time, based on user feedback and problems found during system operation, erroneous data in the corpus is corrected and improved to continuously optimize the quality of the corpus.
[0007] Preferably, the process of using natural language processing tools to check and correct spelling errors, grammatical errors and irregular punctuation in the text includes: for spelling error checking and correction, firstly, using word boundary rules to segment the continuous text sequence in the standard text into words or phrases, obtaining a word or phrase list, and determining the basic composition of words in the text through lexical analysis, then collecting standard vocabulary lists in the fields of medicine, physics, computer science, economics, and literature, and constructing a spelling dictionary, which includes common professional terms and general vocabulary in academic papers. At the same time, using the Merriam-Webster dictionary and an extended dictionary containing network terms and emerging vocabulary as supplementary references to expand the vocabulary coverage; The Levenshtein distance algorithm is used to calculate the edit distance between each word in the text and the standard vocabulary in the spelling dictionary. The edit distance measures the minimum number of operations required to transform a word into another word by inserting, deleting, or replacing characters. If the edit distance is within a threshold range of less than or equal to 2, the word is considered to be misspelled. Based on the spelling error, for cases where the edit distance is close, the correct word with the smallest edit distance and that conforms to the context is searched from the spelling dictionary for replacement; For grammatical error checking and correction, based on the text annotated by the part-of-speech tagger, the part-of-speech category of each word in the text is determined. The part-of-speech categories include nouns, verbs, adjectives, and adverbs. Dependency syntactic analysis is performed. By using an analyzer based on the shift-reduce algorithm, a grammatical structure tree of the text sentence is constructed, the dependency relationship between each word is analyzed, and the grammatical components in the sentence and the modification and dominance relationship between the grammatical components are clarified. The grammatical components include subject, predicate, object, attributive, adverbial, and complement. Check whether the text sentences conform to the grammatical norms according to the pre-set grammatical rule library. The grammatical rule library includes general grammatical rules and grammatical habits. Among them, the grammatical rules include the basic sentence structure rules, tense collocation rules, and part of speech collocation rules in English. Grammatical habits include the grammatical format of citing literature in academic papers and the rules for using specific subject terms in sentences. Compare and analyze the sentence structure and grammatical rules after part of speech tagging and dependency syntax analysis, find out the places that do not conform to the grammatical rules, determine the grammatical errors, and based on the grammatical errors, correct them according to the grammatical rules on the one hand, and make adjustments based on the grammatical norms of the language and the contextual semantics on the other hand; For punctuation checking and correction, we first use the symbol recognition technology in natural language processing to accurately identify and classify the punctuation in the text. Punctuation includes period, comma, semicolon, colon, question mark, exclamation mark, quotation mark, brackets, and clarify the position and type of each punctuation in the text. Set punctuation rules. The punctuation rules include general punctuation standards and punctuation requirements in the field of academic papers. Punctuation standards include the use of periods to indicate the end of sentences and commas to separate parallel components and phrases in sentences. Punctuation requirements in the field of academic papers include the correct use of quotation marks and brackets when citing literature and the use of specific punctuation in formulas. Check the use of punctuation marks. According to the set punctuation mark usage rules, check the use of each punctuation mark in the text to see if there are any irregularities such as missing, redundant, or misused punctuation marks. Then, based on the irregularities found in the punctuation marks, correct them according to the punctuation mark usage rules.
[0008] Preferably, the process of determining the basic composition of words in a text by lexical analysis: first, checking the character encoding format of the standard text; if the character encoding format of the standard text does not meet the processing requirements, converting the character encoding format of the standard text into an adaptive encoding format, that is, converting the character encoding format of the text into an appropriate encoding format, removing irrelevant information in the text, irrelevant information including HTML tags, redundant blank characters, special control characters, and for academic papers, removing document reference marks and footnote numbers, thereby obtaining a preprocessed standard text; based on the preprocessed standard text, extracting language-related features from the text by counting the number of occurrences of each character and the frequency of adjacent character combinations; using a pre-trained large language model, matching the extracted language-related features with various language features in the large language model; the large language model classifies the text into the most similar language category by comparing the text features with typical features of a known language, calculating a similarity score; Use a pre-built rule dictionary, which contains word composition rules and part-of-speech rules. When scanning text, use the rule dictionary to determine word boundaries and segment the text into individual words. For morphologically rich English, use the n-gram model to count the occurrence probabilities of n consecutive words or characters in the text to determine word boundaries, where n is 2 or 3.
[0009] Preferably, the process of training a large language model includes: collecting text data covering different languages from public multilingual books, academic papers, news articles, and web page content, preliminarily organizing the collected text data, removing obvious error information, repeated content, and HTML tags, so that the text content is relatively pure and convenient for subsequent processing, using manual annotation to clearly mark the language category to which each piece of text data belongs, and classifying and organizing the text data according to the language type, constructing text data sets in different languages, ensuring that the text in each data set does belong to the corresponding annotated language, and the size of each language data set is as balanced as possible to avoid the data amount of a certain language being too small to affect the model training effect, and dividing the classified data set of each language into 70% to 80% The training set is used for model parameter learning, the validation set is used to evaluate the performance of the model during training and to assist in adjusting the hyperparameters of the model. The test set is used to objectively test the accuracy of the model in recognizing different languages after the model training is completed. The sum of the proportions of the training set, validation set and test set is 1. The frequency of occurrence of different characters and the frequency of occurrence of n-grams of characters in the text are counted to obtain character-level features. The frequency and distribution of common words in the text are extracted to obtain vocabulary-level features. Different languages have their own specific common words. The grammatical structure characteristics reflected in the text are analyzed to obtain grammatical structure features. Text features are constructed based on character-level features, vocabulary-level features and grammatical structure features. Then, principal component analysis (PCA) is used to standardize and normalize the extracted text features, remove some redundant features that do not contribute much to language recognition, simplify the data structure, improve the model training efficiency, reduce the risk of overfitting, unify the value range of different features, and avoid affecting the model training effect due to large differences in the magnitude of feature values; Based on the naive Bayes model, determine its prior probability calculation method and conditional probability estimation method, build a probability calculation model, and set the initial parameters of the probability distribution. The prior probability calculation method is determined according to the proportion of texts in each language category in the training set. The conditional probability estimation method uses maximum likelihood estimation. The feature vectors of the text are sequentially passed into the probability calculation model for probability calculation. Within the probability calculation model, according to the set value of the log-likelihood loss function, the parameters of the model are updated using stochastic gradient descent (SGD), so that the value of the log-likelihood loss function is continuously reduced, that is, the prediction results of the model are closer and closer to the actual situation. The log-likelihood loss function is used to measure the degree of difference between the model prediction results and the language category to which the text actually belongs. During the training process, after each predetermined training round or iteration, the performance of the probability calculation model is evaluated using the validation set to observe the accuracy, recall, and F1 value indicators. If it is found that the performance of the model on the validation set no longer improves or even decreases, the training is stopped in advance, the hyperparameters of the model are adjusted, and the amount of training data is increased to ensure the generalization ability of the model. After the model training is completed, the text data of the test set is input into the model to let the model predict the language category, and then the prediction results are compared with the real language annotations of the test set texts, and various performance indicators are calculated to understand the performance of the model on actual data that has not been seen, and to determine whether the model meets the expected language recognition ability requirements. Based on the model that meets the expected language recognition ability requirements, a large language model is obtained, where the performance indicators include accuracy, recall, and F1 value, where the accuracy is the ratio of the number of correctly predicted texts to the total number of test texts, the recall is the ratio of the number of correctly predicted texts in the language category to the actual number of texts in the language category, and the F1 value is the harmonic mean of the comprehensive consideration of accuracy and recall.
[0010] Preferably, the process of using the Levenshtein distance algorithm to calculate the edit distance between each word in the text and the standard vocabulary in the spelling dictionary is as follows: first, the target text content to be checked for spelling errors is obtained, and it is segmented to obtain a word list. These words are the objects for which the edit distance with the standard vocabulary in the spelling dictionary needs to be calculated. At the same time, a spelling dictionary is prepared, and each word in the spelling dictionary is used as a reference standard for subsequent comparison calculations, that is, a standard vocabulary. The spelling dictionary is a combination of an authoritative dictionary of a general language and a terminology dictionary in a specific professional field, ensuring that a rich standard vocabulary is covered. A two-dimensional array is created to store the calculation results of the edit distance. The rows of the two-dimensional array correspond to the words in the target text, and the columns correspond to the standard vocabulary in the spelling dictionary. Based on the method of initializing the values of the first row and the first column of the two-dimensional array according to the boundary conditions, each element in the two-dimensional array is initialized. Subsequently, from the word list after the target text is segmented, words are taken out one by one in order to calculate the edit distance. For each word taken out, each standard vocabulary in the spelling dictionary is traversed to calculate the edit distance between them. The target text word is set to word1, and the standard vocabulary is set to word2. The edit distance is calculated step by step by comparing the characters of word1 and word2. Suppose the two-dimensional array is dp[i][j], where i represents the current compared character position in word1, counting from 0 and initially 0, and j represents the current compared character position in word2, also counting from 0 and initially 0. When calculating dp[i][j], it is divided into three cases: insertion operation, deletion operation and replacement operation: For the insertion operation: if the first i characters of word1 are aligned with the first j-1 characters of word2, then inserting the jth character of word2 into word1 will align them. At this time, dp[i][j]=dp[i][j-1]+1, which means adding 1 to the previous state. The previous state refers to the edit distance when compared to the j-1th character of word2, indicating that an insertion operation has been performed. For the deletion operation: if the first i-1 characters of word1 are aligned with the first j characters of word2, then deleting the i-th character of word1 can align them. At this time, dp[i][j]=dp[i-1][j]+1, that is, in the previous state: the edit distance when comparing to the i-1-th character of word1, plus 1, represents a deletion operation; For the replacement operation: when the first i-1 characters of word1 are aligned with the first j-1 characters of word2, if the i-th character of word1 is the same as the j-th character of word2, then dp[i][j]=dp[i-1][j-1]; if they are different, then dp[i][j]=dp[i-1][j-1]+1, indicating that a replacement operation has been performed, and whether to add 1 is determined based on the previous state depending on whether the characters are the same.
[0011] Preferably, the process of taking out words one by one in order from the word list after the target text segmentation to calculate the edit distance also includes: by comparing the three cases of insertion operation, deletion operation and replacement operation, selecting the minimum value as the current value of dp[i][j], that is: dp[i][j]=min(dp[i][j-1]+1,dp[i-1][j]+1,dp[i-1][j-1]+(word1[i]!=word2[j])); where (word1[i]!=word2[j]) returns 1 if the characters are different and 0 if they are the same. Thus, starting from dp[0][0], the edit distance values of the corresponding positions in the entire two-dimensional array are calculated step by step, and updated continuously until all characters of word1 and word2 are compared. At this time, the corresponding dp[m][n] in the two-dimensional array is the edit distance between word1 and word2, and m and n are the last character positions of word1 and word2 respectively. Repeat the word traversal text and character-by-character comparison calculation until all words in the target text complete the edit distance calculation with each standard vocabulary in the spelling dictionary. The two-dimensional array also records the edit distance values between all words and standard vocabulary. The edit distance results of each target text word and each standard vocabulary in the spelling dictionary are extracted from the two-dimensional array to form a data structure convenient for subsequent analysis. The edit distance threshold is determined through experiments and experience, and the edit distance threshold is set to 2 or 3. For each target text word, its edit distance with all standard words in the spelling dictionary is checked. If the edit distance is less than or equal to the edit distance threshold, the target text word is identified as a possible misspelled word, and it is manually confirmed by combining the context; if the edit distance is greater than the edit distance threshold, the target text word is misspelled or is a very rare word that is not included in the dictionary, and it is manually confirmed by looking up materials.
[0012] Preferably, the large language model classifies the text into the most similar language category by comparing the text features with typical features of known languages and calculating the similarity score, including: According to the feature extraction methods: character-level features, vocabulary-level features, and grammatical structure features, feature extraction is performed on the target text to be recognized, including: counting the frequency of occurrence of each character in the target text, calculating the frequency of occurrence of bigrams and trigrams, identifying the occurrence of common words in the text, determining the proportion of stop words and high-frequency words in the text, and at the same time, analyzing the word order and part-of-speech collocation reflected in the text, and organizing the feature information into corresponding feature vectors; By analyzing, counting and processing a large amount of text data with well-labeled language categories, typical features of known languages are obtained. For the feature vector A of the text to be recognized and the typical feature vector B of the known language, the similarity is measured by calculating the cosine value of the angle between them. The calculation formula is: ; Where n is the dimension of the feature vector (such as 100 dimensions above), A i and B i are the values of vectors A and B in the i-th dimension respectively. The value range of cosine similarity is between [-1,1]. The closer to 1, the more similar the two vectors are. Based on the cosine similarity calculation formula, the feature vector of the text to be identified is calculated with the typical feature vectors of all known languages in turn, and the similarity scores calculated with each known language are compared to find the similarity score corresponding to the language category with the highest score, and the language category to which the text belongs is preliminarily determined. The threshold of the similarity score is set to 0.7. If all the calculated similarity scores are lower than this threshold, it means that the characteristics of the text are very different from the typical characteristics of the known languages and cannot be accurately classified. In this case, it is marked as "unknown language" or manual intervention is taken to judge; If the similarity scores of the text to be recognized and the two languages are very close and both are higher than the similarity score threshold, then other auxiliary information or further refined feature analysis is combined to determine the final language category.
[0013] The present invention provides a multidisciplinary academic paper language translation system based on a large language model, which has the following beneficial effects: This multidisciplinary academic paper language translation system based on a large language model can effectively handle spelling errors, grammatical errors and irregular punctuation during the translation process by using a variety of natural language processing technologies and model training methods, providing users with high-quality translations and saving users time and energy on text preprocessing and translation proofreading.
[0014] This multidisciplinary academic paper language translation system based on a large language model collects rich and representative bilingual comparison data and metadata of academic papers in multiple disciplines through a corpus. The bilingual alignment and annotation technology provides comprehensive and accurate language feature information for the large language model, enabling it to better understand the discipline-specific context and terminology during the translation process, thereby generating more accurate translations that meet academic requirements. Moreover, the corpus covers multiple disciplines and is constantly updated, which can keep up with the forefront of academic development, ensure the professionalism of the system's translation of professional terms and expressions in various disciplines, and meet the translation needs of users in different disciplines. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 A schematic diagram of a framework of a multidisciplinary academic paper language translation system based on a large language model according to the present invention; Figure 2 It is a schematic diagram of the framework of the preprocessing module in the present invention. DETAILED DESCRIPTION
[0016] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. The embodiments of the present invention are provided for the purpose of illustration and description, and are not intended to be exhaustive or to limit the present invention to the disclosed forms. Many modifications and variations will be apparent to those of ordinary skill in the art. The embodiments are selected and described in order to better illustrate the principles and practical applications of the present invention, and to enable those of ordinary skill in the art to understand the present invention and thereby design various embodiments with various modifications suitable for specific uses.
[0017] like Figure 1 and Figure 2 As shown, the present invention provides a technical solution: a multidisciplinary academic paper language translation system based on a large language model, the system comprising: The preprocessing module is used to clean the text and standardize the format of the input academic papers; Corpus, which stores bilingual comparative data of multidisciplinary academic papers; Large language models, connected to corpora, trained and optimized translation parameters based on bilingual comparison data; The translation module is used to input the preprocessed academic papers into a large language model to generate the translation output in the target language; A user interface that allows users to upload academic papers, select source and target languages, and display translation results; The quality assessment module is used to evaluate the quality of the translation results and provide feedback to the large language model to optimize future translations. The quality assessment module includes a machine assessment submodule and a manual assessment submodule. The machine assessment submodule is used to automatically evaluate the accuracy and fluency of the translation, and the manual assessment submodule is used to collect user feedback and expert evaluation to facilitate subsequent optimization of translation quality.
[0018] The preprocessing modules include: Tokenizer, which is used to segment academic paper text into words or phrases; A part-of-speech tagger is used to tag the words or phrases after word segmentation, and to identify and tag the part of speech of each unit; Entity recognizers for identifying proper nouns and terms in academic papers.
[0019] The process of the word segmenter segmenting the academic paper text into words or phrases includes: using word boundary rules to segment the continuous text sequence in the standard text into words or phrases. First, the word boundary rules in natural language are collected. The word boundary rules are formed based on the grammar, vocabulary and punctuation of the language. For example, in English, spaces are usually the natural boundaries between words. For example, "Helloworld" can easily distinguish the two words "Hello" and "world" through spaces; at the same time, punctuation includes periods, commas, and question marks. For example, "This is a book." can be divided into different word fragments based on periods to obtain general word boundary rules. For different professional fields involved in multidisciplinary academic papers, the corresponding word boundary-related rules in the field are sorted out to obtain usage rules. For example, in the medical field, medical terms are often fixed phrases. For example, "coronary atherosclerosis" is a complete professional term that cannot be split at will;For the sorting out of the field of computer science, specific keywords in programming languages (such as "if", "else", "for") and function names (such as "print" and "sum") are all regarded as independent word units, and their boundaries need to be determined according to their professional usage. These professional usages are the usage rules corresponding to the corresponding fields. At the same time, refer to the vocabulary and terminology dictionary of the professional field, extract the word composition and boundary information reflected in them, and add them to the rule system, so as to more accurately segment the academic paper text. Based on the usage rules, word composition and boundary information, professional field-specific rules are jointly formed. Among them, word boundary rules include general word boundary rules and professional field-specific rules. The general word boundary rules and professional field-specific rules are combined. Integrate domain-specific rules, build a rule base, and store the rule base in the form of text files and database tables to facilitate rule matching and query in the subsequent word segmentation process. Unify the format of the input standard text, remove unnecessary blanks (such as multiple consecutive spaces) that may exist in the text, convert or standardize line breaks and tabs according to unified format requirements, ensure that the text format is neat, and facilitate the subsequent accurate judgment of word boundaries based on rules. For example, if there are multiple consecutive spaces between words in the text, they are unified into a single space to avoid misjudgment of word boundaries due to format problems, and obtain a standard text format. Based on the standard text format, identify special characters in the text, and for some characters that do not participate in the composition of words but may affect Special characters that affect word boundary judgment, special characters include special symbols that are not English letters, numbers, and Chinese characters, are processed to obtain preprocessed text after special character processing. The processing process includes: if it is a special character that is part of a word (such as superscripts and subscripts in chemical element symbols, which are part of the overall word in professional fields), its special attributes are retained and marked; if it is a special character that simply serves as a separator or decoration (such as asterisks and wavy lines in some texts), it is temporarily removed or converted into a space according to specific needs without affecting the word boundary judgment. Based on the preprocessed text, start from the beginning of the text and scan each character in order, read one character each time, and record the current character position and the characters that have been scanned. For example, for the text "Hello, world!", first read "H" and record its position as the first character. As the scan progresses, keep tracking the changes of the characters. According to the current character position and the scanned character sequence, match them with the common word boundary rules in the rule base. If a space or punctuation mark that meets the common word boundary rules is encountered, the previously scanned character sequence is determined as a word or phrase unit, extracted and recorded. For example, when the comma in "Hello, world!" is scanned, the previous "Hello" is determined as a word and extracted according to the rule that the comma is used as the boundary; when a space is scanned, the "world" after the space is also determined as a word. For a character sequence that conforms to the specific rules of a professional field, check whether the character sequence matches the specific rules of the professional field. If it matches, confirm it as a complete professional term and extract the character sequence as a word unit. For example, when scanning "coronary atherosclerosis", by matching the rules in the medical field, it is confirmed as a complete professional term and the whole is extracted as a word unit without splitting. Among them, when the scanned character sequence does not conform to the general word boundary rules, that is, when the word boundary cannot be determined temporarily, continue to scan the characters backward, continuously expand the range of the character sequence, and try to match the rules again until a matching boundary rule is found to determine the word unit. For example, when dealing with some compound words or newly emerged lexical combinations, like the word "artificial intelligence", when scanning "ren" (person), the boundary may not be determined. Continue to scan "gong" (work), "zhi" (intelligence) until the whole "artificial intelligence" is scanned. By matching the relevant rules about new words and compound words in the rule base, it is determined as a word unit for extraction. After scanning the entire text, the word or phrase units extracted according to the word boundary rules are recorded in sequence to form a list of words or phrases. This list is the result of word segmentation for the original continuous text sequence. Randomly select samples from the list of words or phrases after word segmentation and check them manually to see if there are any incorrect word segmentation situations that do not conform to language habits or professional field requirements. For example, check whether there is a situation where a professional term that should not be split is split, or a situation where originally independent words are wrongly combined into a word unit; If an error is found during the manual random inspection, record the position where the error occurs and the content of the situation, and re-segment the character sequence corresponding to the error part to obtain a corrected sequence, and then manually re-check this character sequence until it is error-free. Through the above technical steps, it is possible to more accurately use the word boundary rules to split the continuous text sequence in the standard text into words or phrases, providing a good data basis for subsequent natural language processing related tasks.
[0020] Corpus, the process of storing bilingual comparison data of multidisciplinary academic papers includes: first, collecting academic papers and their corresponding high-quality translations covering the fields of medicine, physics, computer science, economics, and literature from academic databases, academic journal websites, professional book digital resources, and university academic resource libraries. These papers should be representative, including classic literature, cutting-edge research reports, and academic works of the same type and level, to ensure the richness and diversity of the corpus. At the same time, the metadata of academic papers is collected, including subject classification information, author information, publication journal information, keywords, and abstracts. These metadata will provide important context and subject-specific information references for large language models in the subsequent translation process. The collected original and translated texts are formatted in a unified manner, and unnecessary special characters, garbled characters, and redundant blank characters are removed to ensure the standardization and consistency of the text, so as to facilitate subsequent text processing and analysis and obtain standardized texts. Subsequently, natural language processing tools are used to check and correct spelling errors, grammatical errors, and irregular use of punctuation in the standardized texts to improve the quality and accuracy of the corpus. The standardized texts are segmented to divide continuous text sequences into words or phrases for better subsequent indexing, storage, and analysis operations. Use bilingual alignment technology to align the original text and the translation at the sentence level to ensure that each source language sentence can accurately correspond to the target language sentence, establish a one-to-one bilingual corpus pair, and annotate the aligned bilingual corpus pair. The annotation content includes part of speech, named entities, and grammatical structure. The annotation content provides richer language feature information for large language models, which helps to improve the accuracy and quality of translation. Use a part-of-speech tagger based on a statistical model to annotate the part of speech. For the recognition of named entities, it is used to identify the names of institutions, people, and places in academic papers, and annotate the types to maintain consistency and accuracy during the translation process; use dependency syntactic analysis to analyze the grammatical relationship between words in a sentence, complete the annotation of grammatical structures, and provide grammatical guidance for translation; The cleaned, preprocessed, aligned and annotated bilingual corpus data is stored in a database, using an appropriate data storage structure, such as a table in a relational database or a document-based non-relational database storage method, to efficiently manage and query the corpus data. The inverted indexing technology is used to establish an index relationship between a word and the document or sentence containing the word, and an indexing mechanism is constructed based on keywords, subject classifications, and sentence similarity to index the text data in the corpus, so that the bilingual comparison data and metadata related to the text to be translated can be quickly retrieved during the translation process, thereby improving the data retrieval efficiency and the system response speed; The corpus is updated regularly, and newly published academic papers and their translations are collected and incorporated into the corpus in a timely manner to maintain the timeliness of the corpus and its coverage of the latest academic achievements. At the same time, based on user feedback and problems found during system operation, the erroneous data in the corpus is corrected and improved to continuously optimize the quality of the corpus. Through the above specific technical steps, a high-quality, fully functional corpus can be constructed to provide a solid data foundation for the multidisciplinary academic paper language translation system based on large-scale language models, thereby improving the translation accuracy and professionalism of the system and better meeting the needs of users for multidisciplinary academic paper translation. In the actual implementation process, each step needs to be further optimized and adjusted according to specific system requirements, hardware resources and technical conditions to ensure the efficient operation and stable performance of the system. Among them, the metadata of academic papers is used to guide large-scale language models to consider discipline-specific terminology and context during the translation process.
[0021] The process of using natural language processing tools to check and correct spelling errors, grammatical errors, and irregular punctuation in text includes: For spelling error checking and correction, we first use word boundary rules to segment the continuous text sequence in the standard text into words or phrases to obtain a word or phrase list, and determine the basic composition of words in the text through lexical analysis. Then, we collect standard vocabulary lists in the fields of medicine, physics, computer science, economics, and literature to build a spelling dictionary. The spelling dictionary includes common professional terms and general vocabulary in academic papers. For example, in the medical field, professional vocabulary for the names of various diseases and drugs is included; in the computer science field, vocabulary for programming language-related terms and algorithm names is included. At the same time, the Merriam-Webster dictionary and the extended dictionary containing network terms and emerging vocabulary are used as supplementary references to expand the vocabulary coverage. The Levenshtein distance algorithm is used to calculate the edit distance between each word in the text and the standard vocabulary in the spelling dictionary. The edit distance measures the minimum number of operations required to transform a word into another word by inserting, deleting, or replacing characters. If the edit distance is within a threshold range of less than or equal to 2, the word is considered to be misspelled. Based on the spelling error, for cases where the edit distance is close, the correct word with the smallest edit distance and in line with the context is searched from the spelling dictionary for replacement. For example, if the edit distance between the word "recieve" (wrong spelling) and "receive" (correct spelling) is 1, and "receive" is judged to be grammatically and semantically reasonable based on the context, then the word is replaced and corrected; For grammatical error checking and correction, based on the text annotated by the part-of-speech tagger, the part-of-speech category of each word in the text is determined. The part-of-speech categories include nouns, verbs, adjectives, and adverbs. For example, for the sentence "Thestudentsstudieshard", after part-of-speech tagging, "The" will be marked as a determiner, "students" is a noun plural form, "studies" is a verb third-person singular form, and "hard" is an adverb. Through such annotations, we can preliminarily find clues to grammatical problems such as the inconsistency between the verb form and the subject number, and perform dependency syntax analysis. By using an analyzer based on the shift-reduce algorithm, a grammatical structure tree of the text sentence is constructed, the dependency relationship between each word is analyzed, and the grammatical components in the sentence and the modification and dominance relationship between the grammatical components are clarified. The grammatical components include subject, predicate, object, attributive, adverbial, and complement. For example, for the sentence: "Reading books are my hobby", through dependency syntactic analysis, we can find that "Reading books" as a whole is the subject, and its core word is "books", while the verb "are" is grammatically inconsistent with the subject (the correct one should be "is"), so that this grammatical error of inconsistent subject and predicate can be detected, and the text sentence is checked whether it conforms to the grammatical norms according to the pre-set grammatical rule library. The grammatical rule library includes general grammatical rules and grammatical habits, among which grammatical rules include basic sentence structure rules, tense collocation rules, and part of speech collocation rules in English. Grammatical habits include the grammatical format of citing literature in academic papers and the usage rules of specific subject terms in sentences. The sentence structure and grammatical habits after part of speech tagging and dependency syntactic analysis are compared and analyzed. Grammar rules are used to find places that do not conform to grammatical rules and identify grammatical errors. For example, if a structure such as "adjective + adverb + noun" appears that does not conform to the conventional order of adjectives modifying nouns (generally "adjective + noun"), it is determined that there may be a grammatical error. Based on grammatical errors, such as subject-verb agreement and tense errors, on the one hand, corrections are made according to grammatical rules. For example, after detecting that the subject is plural and the predicate verb uses a singular form, the predicate verb is changed to a plural form to make it conform to grammatical rules. On the other hand, adjustments are made with reference to the grammatical norms of the language and the semantics of the context. For example, in the case of a chaotic sentence structure, the order of sentence components is readjusted, and some necessary grammatical components (such as conjunctions and prepositions) are added or deleted to make the sentence conform to grammatical logic, while ensuring that the semantics of the corrected sentence remains unchanged or is more reasonable. For punctuation checking and correction, we first use the symbol recognition technology in natural language processing to accurately identify and classify the punctuation in the text. Punctuation includes period, comma, semicolon, colon, question mark, exclamation mark, quotation mark, brackets, and clarify the position and type of each punctuation in the text. For example, through text scanning and feature matching, we can identify the "." in the text as a period and the "," as a comma, so as to prepare for the subsequent analysis of whether their use is standardized. Set punctuation rules. Punctuation rules include general punctuation standards and punctuation requirements in the field of academic papers. Punctuation standards include the use of periods to indicate the end of sentences and commas to separate parallel components and phrases in sentences. Punctuation requirements in the field of academic papers include the correct use of quotation marks and brackets when citing literature, and the use of specific punctuation in formulas. For example, in academic papers, when quoting other people's opinions, double quotation marks should be used correctly to enclose the quoted content, and the source of the quotation should be indicated in the appropriate place. This is the embodiment of specific punctuation rules. Check the use of punctuation. According to the set punctuation rules, check the use of each punctuation in the text to see if there are any irregularities such as missing, redundant, or misused punctuation. For example, check whether there is a lack of necessary commas in long sentences, which makes it difficult to understand the semantics of the sentence; or whether there is an abuse of exclamation marks that does not meet the rigor of academic papers. For example, for the sentence "I like reading books I think it's si nteresting", through rule checking, it can be found that there is a lack of appropriate punctuation between the two sentences (commas or periods should be added to separate them), which is an irregular use of punctuation. Subsequently, according to the irregular punctuation found, corrections are made according to the punctuation usage rules. If punctuation is missing, the corresponding punctuation is added at the appropriate position; if punctuation is misused, such as a semicolon is used where a comma should be used, replacement adjustment is made; if it is redundant punctuation, it is deleted. For example, the sentence lacking punctuation mentioned above can be corrected to "I like reading books, I think it's interesting.", so that its punctuation usage conforms to the norm and is more conducive to the clear expression of semantics. Through the above series of specific technical steps, natural language processing tools are used to more comprehensively and meticulously check and correct spelling errors, grammatical errors and irregular use of punctuation in the text, improve the quality of the text, make it more in line with the language requirements of academic papers, and thus better serve the multidisciplinary academic paper language translation system based on large language models.
[0022] The process of determining the basic composition of words in a text through lexical analysis: first check the character encoding format of the standard text, which includes UTF-8 and ASCII. If the character encoding format of the standard text does not meet the processing requirements, convert the character encoding format of the standard text to an adaptive encoding format, that is, convert the character encoding format of the text to a suitable encoding format. For example, if the text is stored in an uncommon localized encoding format, convert it to an adaptive encoding format such as UTF-8 to ensure that subsequent processing can correctly identify characters and remove irrelevant information in the text. Irrelevant information includes HTML tags, redundant whitespace characters, and special control characters. If the text comes from a web page, HTML tags and redundant The blank characters refer to continuous spaces, tabs, and line feeds. Special control characters refer to ASCII control characters, Unicode control characters, and software-specific control characters. In ASCII encoding, most of the character codes 0-31 are control characters. For example, "\n" (line feed, ASCII code value 10), "\r" (carriage return, ASCII code value 13), and "\t" (tab, ASCII code value 9) are relatively common control characters. In the text preprocessing stage, these characters need to be processed. For example, multiple continuous line feeds or tabs can be normalized according to specific needs. There are also some control characters such as "\0" (null character, ASCII code value 0), such as If it appears in the middle of the text, it may cause some problems. For example, when processing text in some programming languages, it may be misunderstood as the end mark of the string, so it needs to be removed or converted appropriately. In Unicode encoding, there are also a series of control characters. For example, the range of "\u0001"-"\u001F" contains some control codes, such as "\u0008" (backspace), which may interfere with normal lexical analysis or text display during text processing. In addition, there are some special-purpose Unicode control characters, such as those used for text direction control (such as control characters related to Arabic and Hebrew written from right to left), text formatting (such as zero-width spaces, used to fine-tune text typesetting) If these characters are not necessary for the text itself, they may be removed as interference factors in the preprocessing stage. Among the software-specific control characters, some texts are extracted from specific software or systems, which contain control characters used by the software for internal processing. For example, some word processing software will insert some hidden characters in the document to mark the document structure, style or editing history. These characters are invisible to ordinary text readers, but may interfere with the lexical analysis process, so they need to be identified and processed. For academic papers, remove reference marks and footnote numbers. For example, replace multiple consecutive spaces with a single space, remove blank characters at the beginning and end of the text, and make the text format more regular.It is convenient for subsequent analysis to obtain the preprocessed standardized text. Based on the preprocessed standardized text, by counting the occurrence times of each character in the text and the frequencies of adjacent character combinations, language-related features are extracted therefrom. The language-related features include character distribution, character combination frequencies, and the frequencies of keywords in specific languages. Among them, the keywords in specific languages include the frequencies of "the", "a", and "is" in English. For example, the ratio of vowel letters to consonant letters in the text is counted. Different languages may have different vowel-consonant distribution rules; or the occurrence frequencies of common word prefixes and suffixes are calculated. For example, the usage frequencies of suffixes such as "-ing" and "-ed" in English vary greatly in different languages. A pre-trained large language model is used to match the extracted language-related features with various language features in the large language model. The large language model calculates a similarity score by comparing the text features with the typical features of known languages and classifies the text into the most similar language category. For example, if the features of the text have the highest similarity with the language model of English, the text language is determined to be English for subsequent use of the corresponding language's lexical analysis method. For languages with clear word boundary rules like Chinese, word segmentation is performed according to the word boundary rules. For example, by using the composition rules of single-character words, two-character words, and multi-character words, combined with punctuation marks (such as "的", "地", "得" as references for dividing boundaries) to divide words; A pre-constructed rule dictionary is used. Among them, the rule dictionary contains word composition rules and part-of-speech rules. When scanning the text, the word boundaries are judged according to the rule dictionary, and the text is segmented into individual words. For example, for the sentence "美丽的花朵" (beautiful flowers), according to the structure of "adjective + 的 + noun" in the rule dictionary, it is segmented into "美丽" (beautiful), "的" (de), and "花朵" (flowers). For morphologically rich English, an n-gram model is used to count the occurrence probabilities of consecutive n words or characters in the text to determine word boundaries. Among them, n can take 2 or 3. Taking the bigram model as an example, the probability of two adjacent words (or characters) appearing simultaneously is calculated. In the training stage, the probabilities of bigram combinations such as "thecat" and "adog" are counted through a large corpus; in the word segmentation stage, whether it is a single word is judged according to the probability size. If the probability of the combination "thecat" is lower than that of "the cat", it is segmented into two words, "the" and "cat".
[0023] The process of training a large language model includes: collecting text data covering different languages from publicly available multilingual books, academic papers, news articles, and web content. For example, collecting classic literary works in English, news reports in French, and academic papers in Chinese, and trying to ensure that there are sufficient and representative text samples for each language. Conducting preliminary collation on the collected text data, removing obvious error information, duplicate content, and HTML tags to make the text content relatively pure and facilitate subsequent processing. Manually annotating each piece of text data to clearly mark its language category. For example, marking a text from a French news website as "French" and a text of an ancient Chinese poem as "Chinese", and classifying and sorting the text data according to language types to construct text datasets in different languages, ensuring that the texts in each dataset truly belong to the corresponding marked language and that the scales of the language datasets are as balanced as possible to avoid the situation where the data volume of a certain language is too small and affects the model training effect. Dividing each language-classified dataset into a training set accounting for 70% - 80%, a validation set accounting for 10% - 15%, and a test set accounting for 10% - 20%. The training set is used for the parameter learning of the model, the validation set is used to evaluate the performance of the model during training and assist in adjusting the hyperparameters of the model, and the test set is used to objectively test the accuracy of the model in recognizing different languages after the model training is completed. And the sum of the proportions of the training set, validation set, and test set is 1. For example, for a total dataset containing 100,000 texts in different languages, about 70,000 can be divided as the training set, 10,000 - 15,000 as the validation set, and the rest as the test set, and it is necessary to ensure that the language distribution ratio in each set is roughly the same as that in the original dataset to avoid data bias. Counting the occurrence frequencies of different characters in the text, the occurrence frequencies of character n-grams (such as bigrams, trigrams) in the text to obtain character-level features. For example, in English, the bigram "th" "he" has a relatively high occurrence frequency, and different languages have their unique character combination rules. These frequency information can be used as features to distinguish languages. Extracting the occurrence frequencies and distribution of common words (such as stop words, high-frequency words) in the text to obtain word-level features. Different languages have their specific common words. For example, "the", "a", "is" in English, and "的", "是", "我" in Chinese. The occurrence frequencies of these words in texts of different languages are significantly different and can be used as important distinguishing features. Analyzing the grammatical structure characteristics reflected in the text to obtain grammatical structure features. The grammatical structure features include the word order characteristics in the sentence (such as Chinese is mostly in the subject-verb-object order, and the predicate in Japanese is often at the end of the sentence), and the collocation characteristics of parts of speech (the collocation methods of nouns, verbs, and adjectives are different in different languages), and converting them into a quantifiable feature form for the model to learn. Based on the character-level features, word-level features, and grammatical structure features, text features are constructed; Then, principal component analysis (PCA) is used to standardize and normalize the extracted text features, remove some redundant features that do not contribute much to language recognition, simplify the data structure, and improve the model training efficiency. It reduces the risk of overfitting, unifies the value range of different features, and avoids the influence of model training effect due to large differences in the magnitude of feature values. For example, for character frequency features and vocabulary frequency features, their values are mapped to the interval [0,1], so that the model can treat different features more fairly during training. For example, if some character n-gram features are highly correlated, they can be combined into a few key features to represent them through PCA. Based on the naive Bayes model, determine its prior probability calculation method and conditional probability estimation method, build a probability calculation model, and set the initial parameters of the probability distribution. The prior probability calculation method is determined according to the proportion of texts in each language category in the training set. The conditional probability estimation method adopts maximum likelihood estimation. The feature vectors of the text (including character frequency and vocabulary frequency) are sequentially passed into the probability calculation model for probability calculation. Inside the probability calculation model, according to the set value of the log-likelihood loss function, the parameters of the model are updated using stochastic gradient descent (SGD), so that the value of the log-likelihood loss function is continuously reduced, that is, the prediction results of the model are closer and closer to the actual situation. The log-likelihood loss function is used to measure the degree of difference between the model prediction results and the language category to which the text actually belongs. For example, in a deep learning model, the gradient is calculated by the back-propagation algorithm, and then the weight parameters of the network are adjusted in the opposite direction of the gradient according to the selected optimization algorithm. After multiple iterative training, the model gradually converges to a better parameter state. During the training process, after each predetermined training round or number of iterations, the performance of the probability calculation model is evaluated using the validation set to observe the accuracy, recall rate, and F1 value indicators. If it is found that the performance of the model on the validation set no longer improves or even decreases (i.e., overfitting occurs), the training is stopped early, the model's hyperparameters are adjusted, and the amount of training data is increased to ensure the generalization ability of the model. After the model training is completed, the text data of the test set is input into the model to let the model predict the language category, and then the prediction results are compared with the actual language annotations of the test set text to calculate the performance of each item. Indicators are used to understand the performance of the model on actual data that has not been seen, to determine whether the model meets the expected language recognition capability requirements, and to obtain a large language model based on the model that meets the expected language recognition capability requirements. Performance indicators include accuracy, recall, and F1 value. Accuracy is the ratio of the number of correctly predicted texts to the total number of test texts, recall is the ratio of the number of correctly predicted texts in a language category to the actual number of texts in the language category, and F1 is the harmonic mean that comprehensively considers accuracy and recall. These indicators are used to comprehensively evaluate the accuracy and precision of the model in recognizing different languages.
[0024] Using the Levenshtein distance algorithm, the process of calculating the edit distance between each word in the text and the standard vocabulary in the spelling dictionary is as follows: first, obtain the target text content to be checked for spelling errors, segment it, and obtain a word list. These words are the objects that need to calculate the edit distance with the standard vocabulary in the spelling dictionary. For example, for the text "Thesisasampletext.", after segmentation, the word list obtained is ["Thes", "is", "a", "sample", "text"]. At the same time, prepare a spelling dictionary, and use each word in the spelling dictionary as a reference standard for subsequent comparison calculations, that is, standard vocabulary. The spelling dictionary is a combination of an authoritative dictionary of a general language (such as the Merriam-Webster dictionary in English) and a terminology dictionary in a specific professional field to ensure that it covers a rich standard vocabulary. For example, the English dictionary contains many correctly spelled words such as "the", "is", and "sample". Create a two-dimensional array to store the calculation results of the edit distance. The rows of the two-dimensional array correspond to the words in the target text, and the columns correspond to the standard vocabulary in the spelling dictionary. For example, if the target The text has 5 words and the spelling dictionary has 1000 standard words, so a two-dimensional array with 5 rows and 1000 columns is created to record the edit distance value between each word and each standard word. Based on the way of initializing the values of the first row and the first column of the two-dimensional array according to the boundary conditions, each element in the two-dimensional array is initialized. For example, for the first row (corresponding to the first word of the target text), the initial value of the edit distance between it and the first standard word in the spelling dictionary can be set according to the length difference between the two words; similarly, the initial value of the edit distance between the first column (corresponding to the first standard word in the spelling dictionary) and each word in the target text is also initialized according to the corresponding rules and set to the word length difference, so as to facilitate the subsequent calculation and gradual update. Then, from the word list after the target text is segmented, the words are taken out one by one in order to calculate the edit distance. For example, the first word "Thes" is taken out first, and then the edit distance between this word and all the standard words in the spelling dictionary is calculated in turn. For each word taken out, each standard word in the spelling dictionary is traversed to fully calculate the edit distance between them. Based on the target text word currently being calculated and the standard vocabulary in the spelling dictionary, the target text word is set to word1 and the standard vocabulary is set to word2. The edit distance is calculated step by step by comparing the characters of word1 and word2. Let the two-dimensional array be dp[i][j], where i represents the current compared character position in word1, starting from 0 and initially 0, and j represents the current compared character position in word2, also starting from 0 and initially 0; when calculating dp[i][j], it is divided into three cases: insertion operation, deletion operation and replacement operation: For the insertion operation: if the first i characters of word1 are aligned with the first j-1 characters of word2, then inserting the j-th character of word2 into word1 will align them. At this time, dp[i][j]=dp[i][j-1]+1, that is, in the previous state: the edit distance when comparing to the j-1-th character of word2, plus 1, indicating that an insertion operation has been performed; For the deletion operation: if the first i-1 characters of word1 are aligned with the first j characters of word2, then deleting the i-th character of word1 can align them. At this time, dp[i][j]=dp[i-1][j]+1, that is, in the previous state: the edit distance when comparing to the i-1-th character of word1, plus 1, represents a deletion operation; For the replacement operation: when the first i-1 characters of word1 are aligned with the first j-1 characters of word2, if the i-th character of word1 is the same as the j-th character of word2, then dp[i][j]=dp[i-1][j-1]; if they are different, then dp[i][j]=dp[i-1][j-1]+1, indicating that a replacement operation has been performed, and whether to add 1 is determined based on the previous state depending on whether the characters are the same.
[0025] The process of taking out words one by one from the word list after the target text segmentation to calculate the edit distance also includes: by comparing the three cases of insertion operation, deletion operation and replacement operation, selecting the minimum value as the current value of dp[i][j], that is: dp[i][j]=min(dp[i][j-1]+1,dp[i-1][j]+1,dp[i-1][j-1]+(word1[i]!=word2[j])); the expression (word1[i]!=word2[j]) returns 1 if the characters are different and 0 if they are the same. Thus, starting from dp[0][0], the edit distance values of the corresponding positions in the entire two-dimensional array are calculated step by step, and updated continuously until all characters of word1 and word2 are compared. At this time, the corresponding dp[m][n] in the two-dimensional array is the edit distance between word1 and word2, and m and n are the last character positions of word1 and word2 respectively. Repeat the word traversal text and character-by-character comparison calculation until all single characters in the target text are compared. The edit distance between each word and each standard word in the spelling dictionary is calculated, and the two-dimensional array records the edit distance values between all words and standard words. The edit distance results between each target text word and each standard word in the spelling dictionary are extracted from the two-dimensional array to form a data structure that is convenient for subsequent analysis. For example, the edit distance value corresponding to each word can be organized into a dictionary form, with the key being the target text word and the value being a list. The list contains the edit distance values between the word and each standard word in the spelling dictionary, which is convenient for subsequent judgment of whether the word may have a spelling error based on the edit distance threshold. For example, for the word "Thes", the corresponding edit distance dictionary value may be {"Thes": [3,5,4,...]}, which indicates the edit distance situation with different standard words. The edit distance threshold is determined through experiments and experience, and the edit distance threshold is set to 2 or 3. For each target text word, its edit distance with all standard words in the spelling dictionary is checked. If the edit distance is less than or equal to the edit distance threshold, the target text word is identified as a possible misspelled word, and it is manually confirmed by combining the context; if the edit distance is greater than the edit distance threshold, the target text word is a misspelled word or a very rare word that is not included in the dictionary, and it is manually confirmed by consulting the data. Through the above specific technical steps, the Levenshtein distance algorithm can be used to effectively calculate the edit distance between each word in the text and the standard words in the spelling dictionary, providing important data basis for subsequent spelling error checking and correction.
[0026] Large language models compare text features with typical features of known languages, calculate similarity scores, and classify text into the most similar language category through the following process: According to the feature extraction methods: character-level features, vocabulary-level features, and grammatical structure features, feature extraction is performed on the target text to be recognized, including: counting the frequency of occurrence of each character in the target text, calculating the frequency of occurrence of two-gram character groups and three-gram character combinations; identifying the occurrence of common words in the text, and determining the proportion of stop words and high-frequency words in the text; at the same time, analyzing the word order and part-of-speech collocation reflected in the text, and organizing the feature information into corresponding feature vectors. Suppose we extract a feature vector containing 100 different feature dimensions, where each dimension represents a specific text feature (such as the frequency of a specific character, the proportion of a certain type of vocabulary). Then, after processing, the target text can be represented as a vector with 100 values for subsequent comparison with typical features of known languages. By analyzing, counting and processing a large amount of text data with well-labeled language categories, we can obtain the typical features of known languages. For each known language: English, Chinese, French, there is a corresponding set of typical feature vectors. For example, in the typical feature vector of the English language determined during the training process, the frequency of occurrence of the character "e" may be relatively high, the common words "the", "a", and "is" have a specific proportion of occurrence, and the sentence structure shows the common order of subject, predicate, and object. These features are combined to form a typical feature vector of the English language, and are also stored in the form of a fixed dimension (such as the 100 dimensions assumed above) to facilitate subsequent comparison operations. For the feature vector A of the text to be recognized and the typical feature vector B of the known language, the similarity is measured by calculating the cosine value of the angle between them. The calculation formula is: ; Where n is the dimension of the feature vector (such as 100 dimensions above), A i and B i are the values of vectors A and B in the i-th dimension respectively. The value range of cosine similarity is between [-1,1]. The closer to 1, the more similar the two vectors are; Based on the cosine similarity calculation formula, the feature vector of the text to be recognized is calculated in turn with the typical feature vectors of all known languages (such as English, Chinese, and German). Assuming that the model can recognize five languages, it is necessary to perform such calculation operations five times to obtain the cosine similarity scores of the target text and these five languages. For example, the cosine similarity score between the text to be recognized and the typical feature vector of English is calculated to be 0.8, and the cosine similarity score with the typical feature vector of Chinese is 0.3. The corresponding similarity scores with other languages are also obtained. These scores will serve as an important basis for judging the language category to which the text belongs. The calculated similarity scores with each known language are compared to find the similarity score corresponding to the language category with the highest score, and the language category to which the text belongs is preliminarily determined. For example, after the above calculation, it is found that the cosine similarity score between the text to be recognized and the typical feature vector of French is the highest, reaching 0.9, which is higher than the similarity scores with other languages. Then the language category to which the text belongs is preliminarily determined to be French. In order to avoid certain errors or noise interference due to feature similarity, If the similarity score of the text to be identified is very close to that of two languages (such as Spanish and Portuguese), and both are higher than the threshold of the similarity score, then other auxiliary information (such as the region of origin of the text, the background related to the text) or further refined feature analysis are used to determine the final language category to ensure the accuracy of text classification. Through the above specific technical steps, the large language model can calculate the similarity score based on the comparison between the text features and the typical features of the known languages, and accurately classify the text into the most similar language category, thereby providing a basis for subsequent text processing operations in the corresponding language.
[0027] It should be further explained that, in the specific implementation process, the corpus is used to collect rich and representative bilingual comparison data and metadata of academic papers in multiple disciplines. The bilingual alignment and annotation technology provides comprehensive and accurate language feature information for large-scale language models, enabling them to better understand the discipline-specific context and terminology during the translation process, thereby generating more accurate translations that meet academic requirements. Moreover, based on the fact that the corpus covers multiple disciplines and is constantly updated, it can keep up with the forefront of academic development, ensure the professionalism of the system's translation of professional terms and expressions in various disciplines, and meet the translation needs of users in different disciplines.
[0028] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A multidisciplinary academic paper language translation system based on a large language model, characterized by: The system includes: The preprocessing module is used to clean the text and standardize the format of the input academic papers; Corpus, which stores bilingual comparative data of multidisciplinary academic papers; A large language model, connected to the corpus, trained and optimized translation parameters based on the bilingual comparison data; A translation module, for inputting the preprocessed academic papers into the large language model to generate a translation output in a target language; A user interface that allows users to upload academic papers, select source and target languages, and display translation results; The quality assessment module is used to evaluate the quality of the translation results and provide feedback to the large language model to optimize future translations. The quality assessment module includes a machine evaluation submodule and a manual evaluation submodule. The machine evaluation submodule is used to automatically evaluate the accuracy and fluency of the translation, and the manual evaluation submodule is used to collect user feedback and expert evaluation.
2. According to claim 1, a multidisciplinary academic paper language translation system based on a large language model is characterized by: The preprocessing module comprises: Tokenizer, which is used to segment academic paper text into words or phrases; A part-of-speech tagger is used to tag the words or phrases after segmentation, and to identify and tag the part of speech of each unit; Entity recognizers for identifying proper nouns and terms in academic papers.
3. A multidisciplinary academic paper language translation system based on a large language model according to claim 2, characterized in that: The process of the word segmenter segmenting the academic paper text into words or phrases includes: using word boundary rules to segment the continuous text sequence in the standard text into words or phrases, firstly collecting word boundary rules in natural language to obtain general word boundary rules, sorting out the corresponding word boundary related rules in the field to obtain usage rules, and at the same time, referring to the vocabulary and terminology dictionary of the professional field, extracting the word composition and boundary information reflected therein, supplementing them to the rule system, and jointly forming professional field specific rules based on usage rules, word composition and boundary information, wherein the word boundary rules include general word boundary rules and professional field specific rules, integrating the general word boundary rules and professional field specific rules to construct a rule base, and storing the rule base in the form of text files and database tables, performing formatting of the input standard text, removing the redundant blank characters that may exist in the text, and converting line breaks and tabs into uniform formats. It is required to convert or standardize to obtain a standard text format. Based on the standard text format, special characters in the text are identified. Some special characters that do not participate in the composition of words but may affect the judgment of word boundaries are processed to obtain a preprocessed text after special character processing. The processing process includes: if it is a special character that is part of a word, its special attribute is retained and marked. If it is a special character that simply serves as a separator or decoration, it is selected to be temporarily removed or converted into a space. Based on the preprocessed text, starting from the beginning of the text, scan one character at a time in sequence, read one character at a time, and record the current character position and the scanned character sequence at the same time. According to the current character position and the scanned character sequence, match with the general word boundary rules in the rule library. If a space or punctuation mark that meets the general word boundary rules is encountered, the previously scanned character sequence is determined as a word or phrase unit, extracted and recorded; For character sequences that meet the specific rules of the professional field, check whether the character sequence matches the specific rules of the professional field. If they match, it is confirmed as a complete professional term and the character sequence is extracted as a word unit. When the scanned character sequence does not meet the general word boundary rules, continue to scan the characters, expand the character sequence range, and try to match the rules again until a matching boundary rule is found to determine the word unit. After scanning the entire text, the word or phrase units extracted according to the word boundary rules are recorded in sequence to form a word or phrase list. Samples are randomly selected from the word or phrase list after word segmentation and checked manually to see if there are any incorrect word segmentations that do not meet language habits or professional field requirements. If an error is found during manual sampling, the location and content of the error are recorded, and the character sequence corresponding to the erroneous part is re-segmented to obtain an error correction sequence, which is then manually rechecked until the character sequence is correct.
4. The multidisciplinary academic paper language translation system based on a large language model according to claim 3 is characterized by: The process of storing bilingual comparison data of multidisciplinary academic papers in the corpus includes: first, collecting original academic papers and their corresponding high-quality translations covering the fields of medicine, physics, computer science, economics, and literature from academic databases, academic journal websites, digital resources of professional books, and academic resource libraries of universities; at the same time, collecting metadata of academic papers, including subject classification information, author information, publication journal information, keywords, and abstracts; performing format unification processing on the collected original and translated texts, removing unnecessary special characters, garbled characters, and redundant blank characters to obtain standardized texts; then using natural language processing tools to check and correct spelling errors, grammatical errors, and irregular use of punctuation marks in the standardized texts; performing word segmentation processing on the standardized texts, and dividing continuous text sequences into words or phrases; Use bilingual alignment technology to align the original text and the translation at the sentence level to ensure that each source language sentence can accurately correspond to the target language sentence, establish a one-to-one bilingual corpus pair, and annotate the aligned bilingual corpus pair. The annotation content includes part of speech, named entity, and grammatical structure. Use a part-of-speech tagger based on a statistical model to annotate the part of speech. For named entity recognition, it is used to identify the names of institutions, people, and places in academic papers and annotate the types; use dependency syntactic analysis to analyze the grammatical relationship between words in the sentence, complete the annotation of the grammatical structure, and store the cleaned, preprocessed, aligned and annotated bilingual corpus data in the database. Use the inverted index technology to establish an index relationship between a word and the document or sentence containing the word, and build an index mechanism based on keywords, subject classification, and sentence similarity. The corpus is updated regularly, and newly published academic papers and their translations are collected and incorporated into the corpus in a timely manner. At the same time, based on user feedback and problems found during system operation, erroneous data in the corpus is corrected and improved to continuously optimize the quality of the corpus.
5. The multidisciplinary academic paper language translation system based on a large language model according to claim 4 is characterized by: The process of using natural language processing tools to check and correct spelling errors, grammatical errors, and irregular punctuation in text includes: For spelling error checking and correction, we first use word boundary rules to segment the continuous text sequence in the standard text into words or phrases to obtain a word or phrase list, and determine the basic word composition in the text through lexical analysis. Then, we collect standard vocabulary lists in the fields of medicine, physics, computer science, economics, and literature to build a spelling dictionary. The spelling dictionary includes common professional terms and general vocabulary in academic papers. At the same time, we use the Merriam-Webster dictionary and the extended dictionary containing network terms and emerging vocabulary as supplementary references. The Levenshtein distance algorithm is used to calculate the edit distance between each word in the text and the standard vocabulary in the spelling dictionary. If the edit distance is within a threshold range of less than or equal to 2, the word is considered to have a spelling error. Based on the spelling error, if the edit distance is close, the correct word with the smallest edit distance and in line with the context is searched from the spelling dictionary for replacement; For checking and correcting grammatical errors, based on the text annotated by the part-of-speech tagger, the part-of-speech category of each word in the text is determined, and dependency syntax analysis is performed. By using an analyzer based on the shift-reduce algorithm, a grammatical structure tree of the text sentence is constructed, the dependency relationship between each word is analyzed, and the grammatical components in the sentence and the modification and dominance relationship between the grammatical components are clarified. The grammatical components include subject, predicate, object, attributive, adverbial and complement. The text sentence is checked to see if it conforms to grammatical norms according to a pre-set grammatical rule library. The grammatical rule library includes common grammatical rules and grammatical habits. The sentence structure and grammatical rules after part-of-speech tagging and dependency syntax analysis are compared and analyzed, and the places that do not conform to the grammatical rules are found to determine grammatical errors. Based on the grammatical errors, on the one hand, corrections are made according to the grammatical rules, and on the other hand, adjustments are made with reference to the grammatical norms of the language and the contextual semantics. For punctuation checking and correction, we first use the symbol recognition technology in natural language processing to accurately identify and classify the punctuation in the text, clarify the position and type of each punctuation in the text, and set the punctuation usage rules. The punctuation usage rules include punctuation usage specifications and punctuation usage requirements in the field of academic papers; Check the use of punctuation marks. According to the set punctuation mark usage rules, check the use of each punctuation mark in the text to see if there are any irregularities such as missing, redundant, or misused punctuation marks. Then, based on the irregularities found in the punctuation marks, correct them according to the punctuation mark usage rules.
6. The multidisciplinary academic paper language translation system based on a large language model according to claim 5, characterized in that: The process of determining the basic composition of words in a text through lexical analysis: first check the character encoding format of the standard text. If the character encoding format of the standard text does not meet the processing requirements, convert the character encoding format of the standard text into an adaptive encoding format, remove irrelevant information in the text, including HTML tags, redundant blank characters, and special control characters. For academic papers, remove reference marks and footnote numbers to obtain preprocessed standard text. Based on the preprocessed standard text, extract language-related features from it by counting the number of occurrences of each character in the text and the frequency of adjacent character combinations. Use a pre-trained large language model to match the extracted language-related features with various language features in the large language model. The large language model compares text features with typical features of known languages, calculates similarity scores, and classifies the text into the most similar language category. Use a pre-built rule dictionary, where the rule dictionary contains word composition rules and part-of-speech rules. When scanning the text, the word boundaries are determined according to the rule dictionary, and the text is segmented into individual words. For English with rich morphology, use the n-gram model to count the occurrence probabilities of n consecutive words or characters in the text to determine the word boundaries, where n is 2 or 3.
7. The multidisciplinary academic paper language translation system based on a large language model according to claim 6, characterized in that: The process of training a large language model includes: collecting text data covering different languages from public multilingual books, academic papers, news articles, and web page content, preliminarily organizing the collected text data, removing obvious error information, repeated content, and HTML tags, using manual annotation to clearly mark the language category of each piece of text data, and classifying and organizing the text data according to the language type to construct text data sets in different languages. The classified data set for each language is divided into a training set accounting for 70% to 80%, a validation set accounting for 10% to 15%, and a test set accounting for 10% to 20%. The training set is used for model parameter learning, the validation set is used to evaluate the performance of the model during the training process and to assist in adjusting the model's hyperparameters. The test set is used to objectively test the model's accuracy in recognizing different languages after the model training is completed. The sum of the proportions of the training set, validation set, and test set is 1. The frequency of occurrence of different characters and the frequency of occurrence of n-grams of characters in the text are counted to obtain character-level features. The frequency and distribution of common words in the text are extracted to obtain vocabulary-level features. The grammatical structure characteristics reflected in the text are analyzed to obtain grammatical structure features. Text features are constructed based on character-level features, vocabulary-level features, and grammatical structure features. Then, principal component analysis (PCA) is used to standardize and normalize the extracted text features, remove some redundant features that do not contribute much to language recognition, simplify the data structure, improve the model training efficiency, reduce the risk of overfitting, and unify the value ranges of different features; based on the naive Bayes model, determine its prior probability calculation method and conditional probability estimation method, build a probability calculation model, set the initial parameters of the probability distribution, and pass the feature vectors of the text into the probability calculation model in sequence for probability calculation. Within the probability calculation model, according to the value of the set log-likelihood loss function, use stochastic gradient descent (SGD) to update the model parameters. During the training process, after each predetermined training round or number of iterations, use the validation set to evaluate the performance of the probability calculation model, and observe the accuracy, recall rate, and F1 value indicators. If it is found that the performance of the model on the validation set no longer improves or even decreases, stop the training early, adjust the model's hyperparameters, and increase the amount of training data to make adjustments; After the model training is completed, the text data of the test set is input into the model, and the model is asked to predict the language category. The prediction results are then compared with the actual language annotations of the test set text, and various performance indicators are calculated to understand the performance of the model on actual unseen data and to determine whether the model meets the expected language recognition capability requirements. Based on the model that meets the expected language recognition capability requirements, a large language model is obtained, where performance indicators include accuracy, recall rate, and F1 value.
8. The multidisciplinary academic paper language translation system based on a large language model according to claim 7 is characterized by: Using the Levenshtein distance algorithm, the process of calculating the edit distance between each word in the text and the standard vocabulary in the spelling dictionary is as follows: First, obtain the target text content to be checked for spelling errors, perform word segmentation on it, and obtain a word list. At the same time, prepare a spelling dictionary, and use each word in the spelling dictionary as a reference standard for subsequent comparison calculations, that is, standard vocabulary. Create a two-dimensional array to store the calculation results of the edit distance. The rows of the two-dimensional array correspond to the words in the target text, and the columns correspond to the standard vocabulary in the spelling dictionary. Based on the method of initializing the values of the first row and the first column of the two-dimensional array according to the boundary conditions, initialize each element in the two-dimensional array, and then take out the words one by one in order from the word list after the target text is segmented for edit distance calculation. For each word taken out, we need to traverse each standard word in the spelling dictionary and calculate the edit distance between them; set the target text word to word1 and the standard word to word2, and gradually calculate the edit distance by comparing the characters of word1 and word2. Let the two-dimensional array be dp[i][j], where i represents the current compared character position in word1, counting from 0, initially 0, and j represents the current compared character position in word2, also counting from 0, initially 0. When calculating dp[i][j], it is divided into three cases: insertion operation, deletion operation and replacement operation: For the insertion operation: if the first i characters of word1 are aligned with the first j-1 characters of word2, then inserting the jth character of word2 into word1 will align them. At this time, dp[i][j]=dp[i][j-1]+1, which means adding 1 to the previous state. The previous state refers to the edit distance when compared to the j-1th character of word2, indicating that an insertion operation has been performed. For the deletion operation: if the first i-1 characters of word1 are aligned with the first j characters of word2, then deleting the i-th character of word1 can align them. At this time, dp[i][j]=dp[i-1][j]+1, that is, in the previous state: the edit distance when comparing to the i-1-th character of word1, plus 1, represents a deletion operation; For the replacement operation: when the first i-1 characters of word1 are aligned with the first j-1 characters of word2, if the i-th character of word1 is the same as the j-th character of word2, then dp[i][j]=dp[i-1][j-1]; if they are different, then dp[i][j]=dp[i-1][j-1]+1, indicating that a replacement operation has been performed, and whether to add 1 is determined based on the previous state depending on whether the characters are the same.
9. The multidisciplinary academic paper language translation system based on a large language model according to claim 8, characterized in that: The process of taking out words one by one from the word list after the target text segmentation to calculate the edit distance also includes: by comparing the three cases of insertion operation, deletion operation and replacement operation, selecting the minimum value as the current value of dp[i][j], that is: dp[i][j]=min(dp[i][j-1]+1,dp[i-1][j]+1,dp[i-1][j-1]+(word1[i]!=word2[j])); where (word1[i]!=word2[j]) returns 1 if the characters are different and 0 if they are the same. Thus, starting from dp[0][0], the edit distance values of the corresponding positions in the entire two-dimensional array are calculated step by step, and updated continuously until all characters of word1 and word2 are compared. At this time, the corresponding dp[m][n] in the two-dimensional array is the edit distance between word1 and word2, and m and n are the last character positions of word1 and word2 respectively. Repeat the word traversal text and character-by-character comparison calculation until all words in the target text complete the edit distance calculation with each standard vocabulary in the spelling dictionary. The two-dimensional array also records the edit distance values between all words and standard vocabulary. The edit distance results of each target text word and each standard vocabulary in the spelling dictionary are extracted from the two-dimensional array; The edit distance threshold is set to 2 or 3. For each target text word, its edit distance with all standard words in the spelling dictionary is checked. If the edit distance is less than or equal to the edit distance threshold, the target text word is identified as a possible misspelled word, and is manually confirmed by combining the context. If the edit distance is greater than the edit distance threshold, the target text word is misspelled or is a very rare word that is not included in the dictionary, and is manually confirmed by looking up materials.
10. The multidisciplinary academic paper language translation system based on a large language model according to claim 9, characterized in that: Large language models compare text features with typical features of known languages, calculate similarity scores, and classify text into the most similar language category through the following process: According to the feature extraction method: character-level features, vocabulary-level features, and grammatical structure features, feature extraction is performed on the target text to be recognized, including: counting the frequency of occurrence of each character in the target text, calculating the frequency of occurrence of two-gram character groups and three-gram character combinations; identifying the occurrence of common words in the text, and determining the proportion of stop words and high-frequency words in the text; at the same time, analyzing the word order and part-of-speech collocation reflected in the text, and organizing the feature information into corresponding feature vectors, and obtaining typical features of known languages by analyzing, counting and processing a large amount of text data with language categories marked; for the feature vector A of the text to be recognized and the typical feature vector B of the known language, the similarity is measured by calculating the cosine value of the angle between them, and the calculation formula is: ; Where n is the dimension of the feature vector (such as 100 dimensions above), A i and B i are the values of vectors A and B in the i-th dimension respectively, and the value range of cosine similarity is between [-1,1].
Citation Information
Cited By
Multi-language intelligent analysis system for medical documents
CN121257560A