Natural language processing method

By extracting the vocabulary feature values ​​and sentence tone of natural pronunciation, and combining various technical means to generate a vocabulary list of target languages, the problem of insufficient analysis of existing translation tools when dealing with natural pronunciation is solved, and higher translation accuracy and scene adaptability are achieved.

CN120031050APending Publication Date: 2025-05-23SHANDONG ENERGY GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510051282.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When existing translation tools deal with natural speech, their analysis is not comprehensive enough, resulting in inaccurate translation and inability to express the user's true intentions.

Method used

By extracting the vocabulary feature values ​​and sentence tone of the sentence, combining the TF-IDF method, logistic regression model and singular value decomposition technology, a vocabulary list of the target language is generated, and the connection words are selected according to the grammatical format for translation.

Benefits of technology

It improves the accuracy and accuracy of natural speech translation and enhances the scene adaptability of translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031050A_ABST
    Figure CN120031050A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of natural language processing, and relates to a natural language processing method. A current translation tool easily has the conditions that the translation of a natural language is not accurate enough, and the real intention of a user cannot be expressed. According to the method, the translation vocabulary list is obtained by extracting the vocabulary feature value, and the target language vocabulary is translated by extracting the statement feature value and the statement tone. The invention aims to improve the translation quality of natural language processing by comprehensively considering the importance, emotion, statement characteristics and mood of vocabularies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing and relates to a natural language processing method. Background Art

[0002] Multilingual communication platforms are widely used in environments that require a large amount of text output, such as international conferences, online education, and global business exchanges. In these scenarios, participants come from different language backgrounds and require instant and accurate language translation to ensure smooth communication and accurate information delivery.

[0003] Although current translation tools can support basic translation needs, when processing natural speech with relatively high scene requirements, natural speech processing is limited by the incomplete analysis of the scene, and it is easy for the translation of natural language to be inaccurate and unable to express the user's true intentions. Summary of the invention

[0004] In order to solve the problems existing in the background technology, the present invention proposes a natural language processing method.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] A natural language processing method, comprising:

[0007] Extracting the vocabulary of the first sentence, and calculating the feature values ​​of all the vocabulary in the first sentence;

[0008] Acquire a vocabulary list of a target language corresponding to the vocabulary translation according to all vocabulary feature values ​​of the first sentence, wherein the feature value of each vocabulary in the vocabulary list is close to the vocabulary feature value of the first sentence;

[0009] Extracting a sentence feature value from the first sentence, and selecting a target language vocabulary from the vocabulary list based on the sentence feature value as a first factor;

[0010] Extracting sentence mood from the first sentence, and selecting target language vocabulary from the vocabulary list based on the sentence mood as a second factor;

[0011] After the target language vocabulary is selected, connectives are selected according to the grammatical format of the target language to connect the vocabulary to obtain a second sentence that has been translated.

[0012] Furthermore, the specific method of extracting the vocabulary of the first sentence and calculating the feature values ​​of all the vocabulary in the first sentence is:

[0013] The importance of the vocabulary to the first sentence is calculated by the TF-IDF method, and the importance feature value of the vocabulary is calculated by the TF-IDF method;

[0014] Construct an emotion dictionary, assign an emotion score to each word, and the value range of the emotion score is [-1, 1];

[0015] Use a logistic regression model to train the training vocabulary set, and the specific training method can be expressed by the following formula:

[0016]

[0017] where y = a is the emotion score a of each training word y, X is the feature vector, and x n β n is the mixed value of each training word and the model parameter;

[0018] Calculate the emotion feature value of each word through the trained logistic regression model.

[0019] Furthermore, the specific method for obtaining the vocabulary list of the corresponding word translation target language based on the entire vocabulary feature values of the first sentence is:

[0020] Establish a target language vocabulary database, and the attributes of each word in the target language vocabulary database include importance feature values and emotion feature values;

[0021] Retrieve according to the importance feature value and emotion feature value of each word;

[0022] First, perform a primary retrieval according to the emotion feature value, select the emotion feature value as the median, obtain the target language words within the first threshold, and generate an intermediate vocabulary list;

[0023] Extract the importance feature value of each word in the intermediate vocabulary list, and remove the words whose importance feature values deviate from the second threshold according to the second threshold to obtain the vocabulary list of the translation target language.

[0024] Furthermore, the specific method for extracting the sentence feature value from the first sentence is:

[0025] Construct the first sentence matrix A, and each element in the first sentence matrix corresponds to the frequency of each word in the first sentence appearing in the first sentence;

[0026] Perform singular value decomposition on the first sentence matrix, and the specific formula is:

[0027] A = U∑V T

[0028] where U is an M×k matrix, and each column represents the latent topic vector of a word; ∑ is a diagonal matrix containing singular values, indicating the importance of each latent topic; V Tis a k×N matrix where each row represents a latent topic vector for a document, where k represents the number of latent topics to be retained;

[0029] Keep the largest k singular values ​​and corresponding singular vectors, and compress the first statement matrix A. The specific formula is:

[0030]

[0031] Among them U k is an M×k matrix containing the topic representation of the vocabulary; k is a k×k diagonal matrix containing the k largest singular values; is a k×N matrix containing the topic representation of the document; according to ∑ k The size of the singular values ​​in , where larger singular values ​​correspond to more important topics;

[0032] Through U k Each column in represents the distribution of each word on each latent topic, Each column in represents the distribution of each document on each potential topic, according to U k and The vectors in are used to explain the specific meaning of each topic.

[0033] Furthermore, the specific method of extracting the sentence mood from the first sentence is:

[0034] Define the names of five modal particles represented by five modal features, and analyze the feature vectors of the five modal particles;

[0035] The feature vector of each word in the first sentence is extracted, and the feature vector of each word is compared with the feature vectors of the five modal word names respectively, and the feature vector of the modal word name closest to it is selected to determine the mood represented by the first sentence.

[0036] Furthermore, the specific method of selecting the target language vocabulary from the vocabulary list according to the first element and the second element is:

[0037] Extracting words with similar feature values ​​from a vocabulary list of the translation target language based on the feature value of the first sentence to generate a candidate vocabulary list;

[0038] Based on the tone of the first sentence, words that match the first sentence are extracted from the candidate word list, and a new word list is created according to the corresponding positions of the words in the first sentence using the extracted words.

[0039] Furthermore, the specific method of selecting conjunctions to connect words according to the grammatical format of the target language is:

[0040] Read the grammatical structure of the target language through a deep learning model and generate the grammatical rules of the target language;

[0041] Based on the vocabulary list of the current first sentence translation, a connecting word is generated through a deep learning model, and the order in the vocabulary list is adjusted.

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] The present invention analyzes the importance features and sentiment features of the sentences to be translated, and obtains a candidate vocabulary list that corresponds to the sentences to be translated. The difference between the words to be translated and the candidate vocabulary is greatly reduced, thereby improving the accuracy of natural speech translation.

[0044] The present invention also analyzes the sentence feature value and sentence tone of the sentence, and selects words from the candidate vocabulary list based on the sentence feature value and sentence tone, so that the words closest to the scene of the sentence to be translated can be selected from the candidate vocabulary list, thereby improving the accuracy and scene adaptability of the translation. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a flow chart of the method of the present invention; DETAILED DESCRIPTION

[0046] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0047] like Figure 1 As shown, the technical solution adopted by the present invention is as follows: a natural language processing method, comprising:

[0048] Extract the vocabulary of the first sentence and calculate the feature values ​​of all the vocabulary in the first sentence. The vocabulary includes words and phrases as well as slang, proverbs, etc. For example, Chinese idioms or poetry punctuation, proverbs, etc.

[0049] A vocabulary list of a target language corresponding to the vocabulary translation is obtained according to all vocabulary feature values ​​of the first sentence, wherein the feature value of each vocabulary in the vocabulary list is close to the vocabulary feature value of the first sentence.

[0050] A sentence feature value is extracted from the first sentence, and a target language vocabulary is selected from the vocabulary list based on the sentence feature value as a first element.

[0051] The sentence mood is extracted from the first sentence, and the target language vocabulary is selected from the vocabulary list according to the sentence mood as the second factor.

[0052] After the target language vocabulary is selected, connectives are selected according to the grammatical format of the target language to connect the vocabulary to obtain a second sentence that has been translated.

[0053] Before starting to translate natural language, you need to segment the natural language. Segmentation can be done using the jieba library in Python. Name the natural language as the first sentence for subsequent operations.

[0054] After extracting the vocabulary of the first sentence through the jiaba library, the feature values ​​of all the vocabulary in the first sentence are calculated. The feature values ​​include importance feature values ​​and sentiment feature values.

[0055] The importance of the vocabulary to the first sentence is calculated by the TF-IDF method, and the feature vector of the vocabulary is calculated by the TF-IDF method.

[0056] Calculate the frequency of each word appearing in the first sentence. The formula is:

[0057]

[0058] Where f(t,d) is the number of times word t appears in document d, ∑ k∈d f(k,d) is the total number of occurrences of all words in the document, and TF(t,d) is the frequency of word t appearing in document d. The higher the frequency, the more the current word can represent the characteristics of the current first sentence. On the contrary, the lower the frequency, the greater the deviation between the word and the characteristics of the current first sentence.

[0059] Count the number of times each word appears in all documents to calculate the inverse document frequency IDF. The specific formula is:

[0060]

[0061] Where N is the total number of words in the first sentence, |d∈D:t∈d| is the number of natural language data containing word t, and the inverse document frequency IDF(t,D) is the rarity of word t in the first sentence.

[0062] Calculate the TF-IDF value for each word in the first sentence to form a feature vector. The specific formula is:

[0063] TF-IDF(t,d,D)=TF(t,d)×IDF(t,D)

[0064] The feature vector of each word in the first sentence is extracted to obtain the importance feature value of the word in the first sentence.

[0065] Construct a sentiment dictionary and assign a sentiment score to each word. The sentiment score ranges from [-1, 1]. When the sentiment of the word is positive, the value is greater than 0, when the sentiment of the word is negative, the value is less than 0, and when the word is a neutral word, the value is 0.

[0066] The training vocabulary set is trained using a logistic regression model. The specific training method can be expressed by the following formula:

[0067]

[0068] Where y=a is the sentiment score a of each training word y, X is the feature vector, x n β n Mixed values ​​for each training vocabulary and model parameters.

[0069] The sentiment feature value of each word is calculated using the trained logistic regression model.

[0070] According to the importance feature values ​​and sentiment feature values ​​of all the words in the first sentence calculated in the above steps, a vocabulary list of the target language corresponding to the vocabulary translation is obtained.

[0071] A target language vocabulary database is established, and the attributes of each word in the target language vocabulary database include importance feature values ​​and sentiment feature values. The target language database should include all words in the target language and the importance feature values ​​and sentiment feature values ​​of all words.

[0072] Retrieval is performed based on the importance feature value and sentiment feature value of each word.

[0073] First, a search is performed based on the sentiment feature value, and the sentiment feature value is selected as the median to obtain the target language vocabulary within the first threshold and generate an intermediate vocabulary list. It is determined that all the words in the intermediate vocabulary list are consistent with the sentiment expressed by the words in the first sentence currently judged. The larger the first threshold is set, the more vague the sentiment expressed by the words in the intermediate vocabulary list is, and the smaller the first threshold is set, the more precise the sentiment expressed by the words in the intermediate vocabulary list is, and accordingly, the number of words will also decrease.

[0074] The importance feature value of each word in the intermediate word list is extracted, and the words whose importance feature value deviates from the second threshold value are removed according to the second threshold value to obtain the word list of the translation target language. Further screening is performed in the intermediate word list to ensure that all words in the word list of the translation target language meet the importance feature value and sentiment feature value of the current matching word in the first sentence.

[0075] After obtaining the vocabulary list, it is necessary to filter the vocabulary in the vocabulary list according to the specific features of the first sentence, so it is necessary to extract the sentence feature value of the first sentence.

[0076] A first sentence matrix A is constructed, wherein each element in the first sentence matrix corresponds to the frequency of occurrence of each word in the first sentence.

[0077] Perform singular value decomposition on the first statement matrix. The specific formula is:

[0078] A=U∑V T

[0079] Where U is an M×k matrix, each column represents a potential topic vector of a word. ∑ is a diagonal matrix containing singular values, indicating the importance of each potential topic, and its dimension is k×k. V T is a k×N matrix where each row represents the latent topic vector for a document, and k represents the number of latent topics to be retained.

[0080] Keep the largest k singular values ​​and corresponding singular vectors, and compress the first statement matrix A. The specific formula is:

[0081]

[0082] Among them U k It is an M×k matrix containing the topic representation of the vocabulary, and each column represents the distribution of a word in the latent topic space. k is a k×k diagonal matrix containing the k largest singular values. is a k×N matrix containing the topic representation of the document, and each row represents the distribution of a document in the latent topic space. k The size of the singular values ​​in , larger singular values ​​correspond to more important topics, and the singular values ​​indicate the importance of each topic. A larger singular value means that the topic occupies a more important position in the document.

[0083] Through U k Each column in represents the distribution of each word on each potential topic. By looking at the vector of a topic, we can determine which words best represent the topic. Each column in represents the distribution of each document on each potential topic, reflecting the topic composition of the document. k The singular value size in , larger singular values ​​correspond to more important topics, which can be calculated based on U k and The vectors in are used to explain the specific meaning of each topic.

[0084] After extracting the topic features of the first sentence, it is also necessary to extract the sentence tone from the first sentence.

[0085] Define the names of the five tone features, and analyze the feature vectors of the five tone names. In this embodiment, the five tones are defined as "anger", "happiness", "fear", "sadness", and "disgust". The user can freely define the five tones during execution.

[0086] The feature vector of each word in the first sentence is extracted, and the feature vector of each word is compared with the feature vectors of the five modal word names respectively, and the feature vector of the modal word name closest to it is selected to determine the modal represented by the first sentence.

[0087] Extract a sentence feature value from the first sentence, and select a target language vocabulary from the vocabulary list based on the sentence feature value as the first factor. Extract a sentence tone from the first sentence, and select a target language vocabulary from the vocabulary list based on the sentence tone as the second factor. Select a target language vocabulary from the vocabulary list based on the first factor and the second factor.

[0088] Based on the feature value of the first sentence, words with similar feature values ​​are extracted from a word list of the translation target language to generate a candidate word list.

[0089] Based on the tone of the first sentence, words that match the first sentence are extracted from the candidate word list, and a new word list is created according to the corresponding positions of the words in the first sentence using the extracted words.

[0090] After the target language vocabulary is selected, connectives are selected to connect the vocabulary according to the target language grammatical format, and the target language grammatical structure is read through the deep learning model to generate the grammatical rules of the target language.

[0091] Based on the vocabulary list of the current first sentence translation, a connecting word is generated through a deep learning model, and the position in the vocabulary list is adjusted to obtain a second sentence with a completed translation.

[0092] Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A natural language processing method, characterized in that: Included are: Extracting the vocabulary of the first sentence, and calculating the feature values ​​of all the vocabulary in the first sentence; Acquire a vocabulary list of a target language corresponding to the vocabulary translation according to all vocabulary feature values ​​of the first sentence, wherein the feature value of each vocabulary in the vocabulary list is close to the vocabulary feature value of the first sentence; Extracting a sentence feature value from the first sentence, and selecting a target language vocabulary from the vocabulary list based on the sentence feature value as a first factor; Extracting sentence mood from the first sentence, and selecting target language vocabulary from the vocabulary list based on the sentence mood as a second factor; After the target language vocabulary is selected, connectives are selected according to the grammatical format of the target language to connect the vocabulary to obtain a second sentence that has been translated.

2. A natural language processing method according to claim 1, characterized in that: The specific method of extracting the vocabulary of the first sentence and calculating the characteristic values ​​of all the vocabulary in the first sentence is: The importance of the vocabulary to the first sentence is calculated by the TF-IDF method, and the importance feature value of the vocabulary is calculated by the TF-IDF method; Construct a sentiment dictionary and assign a sentiment score to each word. The sentiment score ranges from [-1, 1]; The training vocabulary set is trained using a logistic regression model. The specific training method can be expressed by the following formula: Where y=a is the sentiment score a of each training word y, X is the feature vector, x n β n is a mixture of values ​​for each training vocabulary and model parameters; The sentiment feature value of each word is calculated using the trained logistic regression model.

3. A natural language processing method according to claim 2, characterized in that: The specific method of obtaining the vocabulary list of the corresponding vocabulary translation target language according to all vocabulary feature values ​​of the first sentence is: Establishing a target language vocabulary database, wherein the attributes of each word in the target language vocabulary database include importance feature values ​​and sentiment feature values; Search based on the importance and sentiment feature values ​​of each word; Firstly, a search is performed based on the sentiment feature value, the sentiment feature value is selected as the median, the target language vocabulary within the first threshold is obtained, and an intermediate vocabulary list is generated; The importance feature value of each word in the intermediate word list is extracted, and the words whose importance feature values ​​deviate from the second threshold are removed according to the second threshold, so as to obtain the word list of the translation target language.

4. A natural language processing method according to claim 1, characterized in that: The specific method of extracting the sentence feature value from the first sentence is: Construct a first sentence matrix A, where each element in the first sentence matrix corresponds to the frequency of each word in the first sentence appearing in the first sentence; Perform singular value decomposition on the first statement matrix. The specific formula is: A=U∑V T Where U is an M×k matrix, each column represents a potential topic vector of a word; ∑ is a diagonal matrix containing singular values, indicating the importance of each potential topic; V T is a k×N matrix where each row represents a latent topic vector for a document, where k represents the number of latent topics to be retained; Keep the largest k singular values ​​and corresponding singular vectors, and compress the first statement matrix A. The specific formula is: Among them U k is an M×k matrix containing the topic representation of the vocabulary; k is a k×k diagonal matrix containing the k largest singular values; is a k×N matrix containing the topic representation of the document; according to ∑ k The size of the singular values ​​in , where larger singular values ​​correspond to more important topics; Through U k Each column in represents the distribution of each word on each latent topic, Each column in represents the distribution of each document on each potential topic, according to U k and The vectors in are used to explain the specific meaning of each topic.

5. A natural language processing method according to claim 1, characterized in that: The specific method of extracting the sentence mood from the first sentence is: Define the names of five modal particles represented by five modal features, and analyze the feature vectors of the five modal particles; The feature vector of each word in the first sentence is extracted, and the feature vector of each word is compared with the feature vectors of the five modal word names respectively, and the feature vector of the modal word name closest to it is selected to determine the modal represented by the first sentence.

6. A natural language processing method according to claim 1, characterized in that: The specific method of selecting the target language vocabulary from the vocabulary list according to the first element and the second element is: Extracting words with similar feature values ​​from a vocabulary list of the translation target language based on the feature value of the first sentence to generate a candidate vocabulary list; Based on the tone of the first sentence, words that match the first sentence are extracted from the candidate word list, and a new word list is created according to the corresponding positions of the words in the first sentence using the extracted words.

7. A natural language processing method according to claim 6, characterized in that: The specific method of selecting conjunctions to connect words according to the grammatical format of the target language is: Read the grammatical structure of the target language through a deep learning model and generate the grammatical rules of the target language; Based on the vocabulary list of the current first sentence translation, a connecting word is generated through a deep learning model, and the position in the vocabulary list is adjusted.