Natural Language Processing Method, Apparatus and Electronic Device
By performing word segmentation processing on text data and using vector models to obtain semantic vectors, the problems of slow calculation and inaccurate analysis in existing natural language processing technologies are solved, and more efficient semantic expression and processing capabilities are achieved.
Patent Information
- Application Number
- CN202011479380.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-15
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-12-15
AI Technical Summary
In the existing natural language processing technology, processing is performed based on word segmentation, resulting in slow calculations and inaccurate analysis results in some scenarios.
A natural language processing method is proposed, by performing word segmentation processing on text data, obtaining text and/or vocabulary, and inputting text data and its corresponding domain attributes into the text vector model and the vocabulary vector model to obtain word vector and word vector. Then, based on these vectors and weights, sentence semantic vectors are determined for natural language processing.
It effectively improves the semantic expression ability of sentences, enhances the semantic expression ability of sentence-level natural language processing tasks, ensures the simplicity and efficiency of processing, and has a positive effect on downstream tasks.
Smart Images

Figure CN112528654B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer information processing, and is particularly applicable to the field of semantic recognition of machines. More specifically, it relates to a natural language processing method, device, electronic device and computer-readable medium. Background Art
[0002] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural languages. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Natural language processing does not generally study natural languages, but rather focuses on developing computer systems, especially software systems, that can effectively achieve natural language communication. Therefore, it is a part of computer science. In fact, natural language processing, that is, achieving natural language communication between humans and machines, or achieving natural language understanding and natural language generation, is very difficult. A Chinese text or a string of Chinese characters (including punctuation marks, etc.) may have multiple meanings. This is the main difficulty and obstacle in natural language understanding. Conversely, the same or similar meaning can also be represented by multiple Chinese texts or multiple strings of Chinese characters.
[0003] Modern NLP algorithms are based on machine learning, especially statistical machine learning. The machine learning paradigm is different from previous general attempts at language processing. The implementation of language processing tasks usually involves directly coding a large set of rules by hand. Generally, a machine learning model is trained based on a common corpus, the text data containing natural language is segmented, the result after segmentation is input into the trained machine learning model, and then semantic recognition is performed based on word vectors.
[0004] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The present invention aims to solve the dilemmas existing in the natural language processing of the prior art. Because in the natural language processing process of the prior art, it is all based on the method of word segmentation. However, in actual Chinese, single characters can also express many meanings. Moreover, the natural language processing models in the prior art are all trained based on broad corpora, aiming to obtain a natural language processing model applicable to all scenarios. The above two drawbacks make the natural language processing models in the prior art computationally slow and the analysis results given in some scenarios inaccurate.
[0006] To solve the above technical problems, one aspect of the present invention provides a natural language processing method, which includes: performing word segmentation on the text in the text data to obtain words and / or vocabulary; inputting the text data and its corresponding domain attribute into a word vector model to obtain word vectors; inputting the text data and its corresponding domain attribute into a vocabulary vector model to obtain vocabulary vectors; determining a first weight corresponding to the word and / or a second weight corresponding to the vocabulary based on the text data; determining a sentence semantic vector of the text data through the word vectors, the first weight and / or the vocabulary vectors, the second weight; and performing natural language processing on the real-time text data based on the sentence semantic vector.
[0007] According to a preferred embodiment of the present invention, it further includes: extracting sentence semantic vectors of a plurality of preset text data in a database; comparing the similarity between the text data and the plurality of preset text data based on the sentence semantic vectors; and determining target text data from the plurality of preset text data according to the similarity comparison result.
[0008] According to a preferred embodiment of the present invention, it further includes: training a deep neural network model based on a plurality of corpora with domain attributes to generate the word vector model; and training a shallow neural network model based on a plurality of corpora with domain attributes to generate the vocabulary vector model.
[0009] According to a preferred embodiment of the present invention, performing word segmentation on the text in the text data to obtain words and / or vocabulary includes: obtaining a word segmentation dictionary; performing word segmentation on the text data based on the word segmentation dictionary to generate a vocabulary network, where the vocabulary network is a directed acyclic graph; and determining the vocabulary based on the vocabulary network.
[0010] According to a preferred embodiment of the present invention, determining the vocabulary based on the vocabulary network includes: determining the maximum probability path in the vocabulary network based on a dynamic programming algorithm; and determining the vocabulary based on the maximum probability path.
[0011] According to a preferred embodiment of the present invention, after performing word segmentation on the text in the text data to obtain words and / or vocabulary, it further includes: determining the domain attribute of the text data based on the content of the text data; and / or determining the domain attribute of the text data based on the label of the text data.
[0012] According to a preferred embodiment of the present invention, inputting the text data and its corresponding domain attribute into a word vector model to obtain word vectors includes: inputting the text data and its corresponding domain attribute into a trained BERT model to generate word vectors.
[0013] According to a preferred embodiment of the present invention, inputting the text data and its corresponding domain attribute into a lexical vector model to obtain word vectors, including: inputting the text data and its corresponding domain attribute into a trained Word2vec model to generate word vectors.
[0014] According to a preferred embodiment of the present invention, determining a first weight corresponding to the character and / or a second weight corresponding to the vocabulary based on the text data, including: determining the first weight and / or the second weight based on the inverse document frequency corresponding to the character and / or the vocabulary in the text data.
[0015] According to a preferred embodiment of the present invention, determining a sentence semantic vector of the text data through the character vector, the first weight and / or the word vector, the second weight, including: splicing the character vector and / or the word vector according to the first weight and / or the second weight to generate the sentence semantic vector.
[0016] A second aspect of the present invention provides a natural language processing device, which includes: a word segmentation module for segmenting the characters in the text data to obtain characters and / or vocabulary; a character module for inputting the text data and its corresponding domain attribute into a character vector model to obtain character vectors; a vocabulary module for inputting the text data and its corresponding domain attribute into a lexical vector model to obtain word vectors; a weight module for determining a first weight corresponding to the character and / or a second weight corresponding to the vocabulary based on the text data; a vector module for determining a sentence semantic vector of the text data through the character vector, the first weight and / or the word vector, the second weight; and a semantic module for performing natural language processing on the real-time text data based on the sentence semantic vector.
[0017] A third aspect of the present invention provides an electronic device, including a processor and a memory, where the memory is used to store a computer executable program, and when the computer program is executed by the processor, the processor executes the method described above.
[0018] A fourth aspect of the present invention further provides a computer-readable medium storing a computer executable program, and when the computer executable program is executed, the method described above is implemented.
[0019] According to the natural language processing method, apparatus, electronic device, and computer-readable medium of the present disclosure, word segmentation is performed on the words in the text data to obtain words and / or vocabulary; the text data and its corresponding domain attribute are input into a word vector model to obtain word vectors; the text data and its corresponding domain attribute are input into a vocabulary vector model to obtain vocabulary vectors; based on the text data, a first weight corresponding to the word and / or a second weight corresponding to the vocabulary is determined; the sentence semantic vector of the text data is determined through the word vector, the first weight, and / or the vocabulary vector, the second weight; based on the sentence semantic vector, natural language processing is performed on the real-time text data, which can effectively improve the semantic expression ability of sentences, greatly enhance the semantic expression ability of sentence-level natural language processing tasks while ensuring their simplicity and efficiency, and achieve the purpose of having a positive and positive effect on downstream tasks.
[0020] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a system block diagram of the natural language processing method and apparatus according to an embodiment of the present invention.
[0022] Figure 2 It is a flowchart of the natural language processing method according to an embodiment of the present invention.
[0023] Figure 3 It is a flowchart of the natural language processing method according to an embodiment of the present invention.
[0024] Figure 4 It is a flowchart of the natural language processing method according to an embodiment of the present invention.
[0025] Figure 5 It is a block diagram of the natural language processing apparatus according to an embodiment of the present invention.
[0026] Figure 6 It is a schematic structural diagram of an electronic device according to an embodiment of the present invention;
[0027] Figure 7 It is a schematic diagram of a computer-readable recording medium according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] In the process of introducing the specific embodiments, the detailed description of the structure, performance, effect, or other features is to enable those skilled in the art to fully understand the embodiments. However, it does not exclude that those skilled in the art can implement the present invention with technical solutions that do not include the above structure, performance, effect, or other features under specific circumstances.
[0029] The flowcharts in the drawings are only exemplary flow demonstrations, and do not represent that all the contents, operations, and steps in the flowcharts must be included in the solution of the present invention, nor does it represent that they must be executed in the order shown in the figures. For example, some operations / steps in the flowchart can be decomposed, some operations / steps can be combined or partially combined, etc. Without departing from the gist of the present invention, the execution order shown in the flowchart can be changed according to the actual situation.
[0030] The boxes in the drawings Figure 1 generally represent functional entities, and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processing unit devices and / or microcontroller devices.
[0031] The same reference numerals in each drawing represent the same or similar elements, components, or parts. Therefore, the repeated description of the same or similar elements, components, or parts may be omitted hereinafter. It should also be understood that although the text may use attributives such as first, second, third, etc. to describe various devices, elements, components, or parts, these devices, elements, components, or parts should not be limited by these attributives. That is, these attributives are only used to distinguish one from another. For example, the first device can also be called the second device without departing from the essence of the technical solution of the present invention. In addition, the terms "and / or", "or / and" mean all combinations including any one or more of the listed items.
[0032] To solve the above technical problems, the present invention provides a natural language processing method, device, electronic device, and computer-readable medium, which perform word segmentation on the text in the text data to obtain words and / or vocabulary; input the text data and its corresponding domain attribute into a word vector model to obtain word vectors; input the text data and its corresponding domain attribute into a vocabulary vector model to obtain vocabulary vectors; determine a first weight corresponding to the word and / or a second weight corresponding to the vocabulary based on the text data; determine a sentence semantic vector of the text data through the word vectors, the first weight, and / or the vocabulary vectors, the second weight; and perform natural language processing on the real-time text data based on the sentence semantic vector, which can effectively improve the semantic expression ability of the sentence, greatly enhance the semantic expression ability of the sentence-level natural language processing task on the premise of ensuring its simplicity and efficiency, and achieve the purpose of having a positive and positive effect on downstream tasks.
[0033] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the following provides a further detailed description of the present invention in conjunction with specific embodiments and with reference to the accompanying drawings.
[0034] During the introduction of specific embodiments, the detailed descriptions of structures, performances, effects, or other features are provided to enable those skilled in the art to fully understand the embodiments. However, it does not exclude that those skilled in the art can implement the present invention with technical solutions that do not include the above-mentioned structures, performances, effects, or other features under specific circumstances.
[0035] The flowcharts in the accompanying drawings are merely exemplary flow demonstrations, and do not represent that all the contents, operations, and steps in the flowcharts must be included in the solutions of the present invention, nor does it represent that they must be executed in the order shown in the figures. For example, some operations / steps in the flowchart can be decomposed, some operations / steps can be combined or partially combined, etc. Without departing from the gist of the present invention, the execution order shown in the flowchart can be changed according to the actual situation.
[0036] The boxes in the accompanying drawings Figure 1 generally represent functional entities, and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processing unit devices and / or microcontroller devices.
[0037] The same reference numerals in the drawings represent the same or similar elements, components, or parts. Therefore, the repeated descriptions of the same or similar elements, components, or parts may be omitted hereinafter. It should also be understood that although the text may use ordinal adjectives such as first, second, third, etc. to describe various devices, elements, components, or parts, these devices, elements, components, or parts should not be limited by these ordinal adjectives. That is, these ordinal adjectives are only used to distinguish one from another. For example, the first device can also be called the second device without departing from the essence of the technical solution of the present invention. In addition, the terms "and / or", "or / and" mean all combinations including any one or more of the listed items.
[0038] Figure 1 is a system block diagram of a natural language processing method and apparatus shown according to an exemplary embodiment.
[0039] Such as Figure 1As shown, the system architecture 10 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0040] Users can use the terminal devices 101, 102, 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as auxiliary learning applications, web browser applications, instant messaging tools, email clients, social platform software, etc.
[0041] The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop portable computers, and desktop computers, etc.
[0042] In one embodiment, the terminal devices 101, 102, 103 may, for example, perform word segmentation on the text in the text data to obtain words and / or vocabulary; the terminal devices 101, 102, 103 may, for example, input the text data and its corresponding domain attributes into a word vector model to obtain word vectors; the terminal devices 101, 102, 103 may, for example, input the text data and its corresponding domain attributes into a vocabulary vector model to obtain vocabulary vectors; the terminal devices 101, 102, 103 may, for example, determine a first weight corresponding to the words and / or a second weight corresponding to the vocabulary based on the text data; the terminal devices 101, 102, 103 may, for example, determine a sentence semantic vector of the text data through the word vectors, the first weight, and / or the vocabulary vectors, the second weight; the terminal devices 101, 102, 103 may, for example, perform natural language processing on the real-time text data based on the sentence semantic vector. Among them, the word vector model and the vocabulary vector model may be located locally on the terminal devices 101, 102, 103 or at the server 105 side.
[0043] The server 105 may be a server providing various services, such as a background management server that supports video learning websites browsed by users using the terminal devices 101, 102, 103. The background management server may perform natural language processing on the received text data and feedback the processing results to the terminal devices 101, 102, 103.
[0044] In one embodiment, the server 105 may obtain text data from the terminal devices 101, 102, and 103, for example, and then perform word segmentation on the text in the text data to obtain words and / or vocabulary; the server 105 may, for example, input the text data and its corresponding domain attribute into a word vector model to obtain word vectors; the server 105 may, for example, input the text data and its corresponding domain attribute into a vocabulary vector model to obtain vocabulary vectors; the server 105 may, for example, determine a first weight corresponding to the word and / or a second weight corresponding to the vocabulary based on the text data; the server 105 may, for example, determine a sentence semantic vector of the text data through the word vectors, the first weight, and / or the vocabulary vectors, the second weight; the server 105 may, for example, perform natural language processing on the real-time text data based on the sentence semantic vector.
[0045] The server 105 may also, for example, extract the sentence semantic vectors of multiple preset text data in the database; the server 105 may also, for example, compare the similarity between the text data and the multiple preset text data based on the sentence semantic vectors; the server 105 may also, for example, determine target text data from the multiple preset text data according to the similarity comparison result.
[0046] The server 105 may also, for example, train a deep neural network model based on multiple corpora with domain attributes to generate the word vector model; the server 105 may also, for example, train a shallow neural network model based on multiple corpora with domain attributes to generate the vocabulary vector model.
[0047] The server 105 may be a physical server, or may also be composed of multiple servers, for example. A part of the server 105 may, for example, train machine learning models to generate a word vector model and a vocabulary vector model; and a part of the server 105 may also, for example, perform natural language processing on text data.
[0048] It should be noted that the natural language processing method provided by the embodiments of the present disclosure may be executed by the server 105 or the terminal devices 101, 102, and 103. Correspondingly, the natural language processing device may be disposed in the server 105 or the terminal devices 101, 102, and 103.
[0049] Figure 2 is a flowchart of a natural language processing method shown according to an exemplary embodiment. The natural language processing method 20 includes at least steps S202 to S212.
[0050] As Figure 2As shown, in S202, word segmentation processing is performed on the text in the text data to obtain words and / or vocabulary. Among them, the text data can be the text data of the user during the human-computer interaction process, or the text data converted from the user's voice data. The text data can include one or more sentences composed of natural language.
[0051] In the present disclosure, the word segmentation processing can be Chinese word segmentation processing. Word segmentation is the process of recombining a continuous sequence of characters into a sequence of words according to certain specifications. Existing word segmentation algorithms can be divided into three categories: string matching-based word segmentation methods, understanding-based word segmentation methods, and statistics-based word segmentation methods. According to whether it is combined with the part-of-speech tagging process, it can also be divided into pure word segmentation methods and integrated methods that combine word segmentation and tagging. In the present invention, one or more of the above methods can be used to perform word segmentation processing on the text data to generate multiple Chinese characters and vocabulary.
[0052] In one embodiment, it further includes: determining the domain attribute of the text data based on the content of the text data; and / or determining the domain attribute of the text data based on the label of the text data. The domain attribute of the text data can be obtained from the dialogue request of the human-computer dialogue, and can also be determined from the words after word segmentation of the text data. The present disclosure is not limited thereto.
[0053] In S204, the text data and its corresponding domain attribute are input into the word vector model to obtain word vectors. For example, the text data and its corresponding domain attribute can be input into the trained BERT model to generate word vectors.
[0054] In one embodiment, it further includes: training a deep neural network model based on multiple corpora with domain attributes to generate the word vector model; among them, the deep neural network model can be a deep neural network model of the BERT series, specifically including the BERT model, the ALBERT model, etc. BERT is a method of pre-training language representation. A general "language understanding" model is trained on a large amount of text corpora (Wikipedia), and then this model is used to perform the NLP tasks that you want to do. BERT is the first unsupervised, deep bidirectional system used in pre-training NLP. Unsupervised means that BERT only needs to be trained with pure text corpora because a large amount of text corpora can be publicly obtained on various language networks. The pre-trained representation can be context-independent or context-dependent, and moreover, the context-dependent representation can be unidirectional or bidirectional.
[0055] In an embodiment of the present invention, when performing BERT model training, when inputting corpus data, it is divided according to domain attributes. For example, in the "mathematics" domain, the "chemistry" domain, etc. Corpus data in different attribute domains are used to train different BERT models to generate word vector models for different domain attributes.
[0056] In S206, the text data and its corresponding domain attribute are input into the word vector model to obtain word vectors. For example, the text data and its corresponding domain attribute can be input into the trained Word2vec model to generate word vectors.
[0057] In one embodiment, it further includes: training a shallow neural network model based on multiple corpora with domain attributes to generate the word vector model. The shallow neural network model can be a Word2vec model. Word2vec is a group of related models used to generate word vectors. These models are shallow and double-layer neural networks, trained to reconstruct linguistic word texts. The network is represented by words and needs to guess the input words at adjacent positions. Under the bag-of-words model assumption in word2vec, the order of words is not important. After training, the word2vec model can be used to map each word to a vector, which can be used to represent the relationship between words. This vector is the hidden layer of the neural network.
[0058] In an embodiment of the present invention, when performing word2vec model training, when inputting corpus data, it is also divided according to domain attributes. For example, in the "mathematics" domain, the "chemistry" domain, etc. Corpus data in different attribute domains are used to train different word2vec models to generate word vector models for different domain attributes.
[0059] In S208, based on the text data, determine the first weight corresponding to the text and / or the second weight corresponding to the word. The first weight and / or the second weight can be determined based on the inverse document frequency corresponding to the text and / or the word in the text data.
[0060] Among them, the inverse document frequency (TF-IDF) is a statistical method used to evaluate the importance of a word or a term for a document set or a single document in a corpus. The importance of a word or a term increases proportionally with the number of times it appears in a document, but at the same time decreases inversely with the frequency of its appearance in the corpus. Various forms of TF-IDF weighting are used as the importance rating of words or terms in the present invention.
[0061] In S210, a sentence semantic vector of the text data is determined by the word vectors, the first weights, and / or the token vectors, the second weights. For example, the word vectors and / or the token vectors may be concatenated according to the first weights and / or the second weights to generate the sentence semantic vector.
[0062] In a specific embodiment, the text data is: "Today is cloudy."
[0063] 1) Word segmentation of "Today is cloudy": Today is cloudy;
[0064] 2) Obtain the token vectors of each word: Today: [0.1, 0.2, 0.3]; is: [0.4, 0.5, 0.6]; cloudy: [0.7, 0.8, 0.9];
[0065] 3) The second sentence vector corresponding to the token vector segmentation method of this sentence: (idf(Today) * [0.1, 0.2, 0.3] + idf(is) * [0.4, 0.5, 0.6] + idf(cloudy) * [0.7, 0.8, 0.9]) / 3, and the result is also a three-dimensional vector (it can also be a vector with more dimensions, and this application is not limited to this).
[0066] 4) Word segmentation of "Today is cloudy": Today is cloudy;
[0067] 5) Obtain the word vectors of each character: Jin: [0.1, 0.2, 0.3]; Tian: [0.12, 0.82, 0.92]; is: [0.4, 0.5, 0.6]; Yin: [0.7, 0.8, 0.9]; (Note: There is no restriction on the equality of the lengths of the word vectors in 2) above. Preferably, the difference in length between the character vectors and the word vectors should not be too large either.)
[0068] 6) The first sentence vector corresponding to the word vector segmentation method of this sentence: (idf(Jin) * [0.1, 0.2, 0.3] + idf(Tian) * [0.12, 0.82, 0.92] + idf(is) * [0.4, 0.5, 0.6] + idf(Yin) * [0.7, 0.8, 0.9] + idf(Tian) * [0.12, 0.82, 0.92]) / 5, and the result is a three-dimensional vector (it can also be a vector with more dimensions, and this application is not limited to this).
[0069] 7) Concatenate the sentence vectors of each granularity: (sentence vector 1, sentence vector 2...), and the final dimension is the sum of the lengths of each sentence vector.
[0070] In S212, natural language processing is performed on the real-time text data based on the sentence semantic vector.
[0071] In one embodiment, it further includes: extracting the sentence semantic vectors of multiple preset text data in a database; comparing the similarity between the text data and the multiple preset text data based on the sentence semantic vectors; and determining target text data from the multiple preset text data according to the similarity comparison result. For example, a user inputs a text data, which may be a math application problem. Searching in a question bank according to the data input by the user, the cosine distance between the sentence semantic vector of the text data and all the questions in the question bank can be calculated as the similarity between the two sentences, and then the question and the corresponding solution method most similar to the text data in the question bank are determined and the results are returned to the user.
[0072] According to the natural language processing method of the present disclosure, word segmentation is performed on the words in the text data to obtain words and / or vocabulary; the text data and its corresponding domain attribute are input into a word vector model to obtain word vectors; the text data and its corresponding domain attribute are input into a vocabulary vector model to obtain vocabulary vectors; the first weight corresponding to the word and / or the second weight corresponding to the vocabulary are determined based on the text data; the sentence semantic vector of the text data is determined through the word vectors, the first weight and / or the vocabulary vectors, the second weight; based on the sentence semantic vector, the natural language processing method for the real-time text data can effectively improve the semantic expression ability of the sentence, greatly enhance the semantic expression ability of the sentence-level natural language processing task on the premise of ensuring its simplicity and efficiency, and achieve the purpose of having a positive impact on downstream tasks.
[0073] It should be clearly understood that the present disclosure describes how to form and use specific examples, but the principles of the present disclosure are not limited to any details of these examples. On the contrary, based on the teachings of the content disclosed in the present disclosure, these principles can be applied to many other embodiments.
[0074] Figure 3 It is a flowchart of a natural language processing method shown according to another exemplary embodiment. Figure 3 The process 30 shown is for Figure 2 a detailed description of S202 "performing word segmentation on the words in the text data to obtain words and / or vocabulary" in the process shown.
[0075] As Figure 3 shown, in S302, a word segmentation dictionary is obtained.
[0076] In S304, word segmentation is performed on the text data based on the word segmentation dictionary to generate a lexical network, which is a directed acyclic graph. A directed acyclic graph refers to a directed graph without loops. If there is a non-directed acyclic graph, and starting from point A to B via C can return to A, forming a loop. Changing the direction of the edge from C to A to from A to C will turn it into a directed acyclic graph. The number of spanning trees of a directed acyclic graph is equal to the product of the in-degrees of the nodes with non-zero in-degrees.
[0077] In S306, the maximum probability path in the lexical network is determined based on the dynamic programming algorithm. The dynamic programming algorithm is usually used to solve problems with certain optimal properties. In such problems, there may be many feasible solutions. Each solution corresponds to a value, and we hope to find the solution with the optimal value. The dynamic programming algorithm is similar to the divide-and-conquer method. Its basic idea is also to decompose the problem to be solved into several sub-problems, solve the sub-problems first, and then obtain the solution of the original problem from the solutions of these sub-problems.
[0078] In the lexical network, given the state of a stage, a choice (action) from this state to a certain state in the next stage is called a decision. A sequence composed of decisions in each stage is called a strategy. For each actual multi-stage decision-making process, the range of available strategies is limited, and this range is called the set of allowable strategies. The strategy that achieves the optimal effect in the set of allowable strategies is called the optimal strategy. In the present invention, the optimal strategy is located at the maximum probability path of all segmented words.
[0079] More specifically, in the present invention, all word segmentation paths are first searched through the lexical network. Then the word segmentation path is the path with the maximum probability, and the probability of each path = the product of the probabilities of all words on this path.
[0080] In S308, the vocabulary is determined based on the maximum probability path.
[0081] Figure 4 It is a flowchart of a natural language processing method shown according to another exemplary embodiment. Figure 4 The shown process 40 is a detailed description of the whole process of the natural language processing method of the present invention.
[0082] As Figure 4 shown, in S402, corpus data is collected to generate multiple corpus data sets. Public domain corpus is collected.
[0083] In S404, it is determined whether the data in each corpus data set has been processed.
[0084] In S406, cleaning and screening are performed to obtain sentences. The corpus data is screened and cleaned to obtain text sentences.
[0085] In S408, word / phrase segmentation processing is performed, and their IDFs are respectively counted. The text sentences are segmented at the character and word granularities respectively to obtain the segmented text sentences. For the segmented characters and words respectively, their idfs are counted.
[0086] In S410, for the sentences after word / phrase segmentation respectively, word / phrase vectors are trained. The neural network is used to train the sentences segmented by characters and words respectively to obtain the corresponding word / phrase vectors and save them.
[0087] In S412, the model is saved.
[0088] In S414, the word / phrase vectors are obtained, weighted by IDF and concatenated to obtain sentence vectors. For a new sentence, the corresponding word / phrase vectors are obtained after segmenting words and characters, weighted by tf-idf respectively and then concatenated, so as to generate the sentence semantic vector, which is used as the representation of this sentence.
[0089] Those skilled in the art can understand that all or part of the steps of implementing the above embodiments are implemented as a computer program executed by a CPU. When this computer program is executed by the CPU, the above functions defined by the above method provided by the present disclosure are executed. The said program can be stored in a computer-readable storage medium, and this storage medium can be a read-only memory, a magnetic disk or an optical disc, etc.
[0090] In addition, it should be noted that the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.
[0091] The following is an embodiment of the device of the present disclosure, which can be used to execute the embodiment of the method of the present disclosure. For details not disclosed in the embodiment of the device of the present disclosure, please refer to the embodiment of the method of the present disclosure.
[0092] Figure 5 It is a block diagram of a natural language processing device shown according to an exemplary embodiment. As Figure 5 shown, the natural language processing device 50 includes: a word segmentation module 502, a character module 504, a vocabulary module 506, a weight module 508, a vector module 510, and a semantic module 512.
[0093] The word segmentation module 502 is used to perform word segmentation processing on the characters in the text data to obtain characters and / or vocabulary;
[0094] The character module 504 is used to input the text data and its corresponding domain attributes into a character vector model to obtain character vectors;
[0095] The vocabulary module 506 is used to input the text data and its corresponding domain attributes into a vocabulary vector model to obtain word vectors;
[0096] The weight module 508 is used to determine a first weight corresponding to the text and / or a second weight corresponding to the vocabulary based on the text data;
[0097] The vector module 510 is used to determine a sentence semantic vector of the text data through the character vector, the first weight, and / or the word vector, the second weight;
[0098] The semantic module 512 is used to perform natural language processing on the real-time text data based on the sentence semantic vector.
[0099] According to the natural language processing device of the present disclosure, word segmentation processing is performed on the text in the text data to obtain words and / or vocabulary; the text data and its corresponding domain attributes are input into a character vector model to obtain character vectors; the text data and its corresponding domain attributes are input into a vocabulary vector model to obtain word vectors; a first weight corresponding to the text and / or a second weight corresponding to the vocabulary is determined based on the text data; a sentence semantic vector of the text data is determined through the character vector, the first weight, and / or the word vector, the second weight; and natural language processing is performed on the real-time text data based on the sentence semantic vector. In this way, the semantic expression ability of sentences can be effectively improved, and for sentence-level natural language processing tasks, on the premise of ensuring its simplicity and efficiency, its semantic expression ability is greatly enhanced, achieving the purpose of having a positive and positive effect on downstream tasks.
[0100] Figure 6 FIG. is a schematic structural diagram of an electronic device according to an embodiment of the present invention. The electronic device includes a processor and a memory. The memory is used to store computer-executable programs. When the computer program is executed by the processor, the processor executes a vehicle intelligent assisted propulsion method based on rotation angle monitoring.
[0101] As Figure 6 shown, the electronic device is presented in the form of a general-purpose computing device. The processor can be one or multiple and work cooperatively. The present invention does not exclude distributed processing, that is, the processors can be dispersed in different physical devices. The electronic device of the present invention is not limited to a single entity and can also be the sum of multiple physical devices.
[0102] The memory stores computer-executable programs, usually machine-readable codes. The computer-readable programs can be executed by the processor so that the electronic device can execute the method of the present invention or at least some steps of the method.
[0103] The memory includes volatile memory, such as random access memory units (RAM) and / or cache memory units, and may also include non-volatile memory, such as read-only memory units (ROM).
[0104] Optionally, in this embodiment, the electronic device further includes an I / O interface for data exchange between the electronic device and external devices. The I / O interface may represent one or more of several bus structures, including a memory unit bus or a memory unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.
[0105] It should be understood that Figure 6 The displayed electronic device is merely an example of the present invention. The electronic device of the present invention may also include elements or components not shown in the above examples. For example, some electronic devices also include a display unit such as a display screen, and some electronic devices also include human-computer interaction elements, such as buttons, keyboards, etc. As long as the electronic device can execute the computer-readable program in the memory to implement at least part of the steps of the method of the present invention or the method, it can be considered as the electronic device covered by the present invention.
[0106] Figure 7 is a schematic diagram of a computer-readable recording medium according to an embodiment of the present invention. As Figure 7 shown, the computer-readable recording medium stores a computer-executable program, and when the computer-executable program is executed, the vehicle intelligent assisted propulsion method based on rotation angle monitoring described above of the present invention is implemented. The computer-readable storage medium may include data signals propagated in a baseband or as part of a carrier wave, which carry readable program codes. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium may also be any readable medium other than the readable storage medium, and the readable medium may send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program codes contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.
[0107] The program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0108] From the above description of the embodiments, those skilled in the art can easily understand that the present invention can be implemented by hardware capable of executing a specific computer program, such as the system of the present invention, and the electronic processing unit, server, client, mobile phone, control unit, processor, etc. included in the system. The present invention can also be implemented by a computer software that executes the method of the present invention, such as the control software executed by the microprocessor and electronic control unit at the locomotive end, the client, the server end, etc. However, it should be noted that the computer software for executing the method of the present invention is not limited to being executed in one or a specific number of hardware entities. It can also be implemented in a distributed manner by unspecified specific hardware. For example, some method steps executed by a computer program can be executed at the locomotive end, and another part can be executed in a mobile terminal or a smart helmet, etc. For computer software, the software product can be stored in a computer-readable storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.), or it can be distributed and stored on the network, as long as it can enable an electronic device to execute the method according to the present invention.
[0109] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the present invention is not inherently related to any specific computer, virtual device, or electronic device, and various general-purpose devices can also implement the present invention. The above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A natural language processing method, characterized in that, Including: Obtain a word segmentation dictionary; Segment the real-time text data based on the word segmentation dictionary to generate a vocabulary network, where the vocabulary network is a directed acyclic graph; Determine the maximum probability path in the vocabulary network based on the dynamic programming algorithm, including: first search out all word segmentation paths through the vocabulary network, and then the word segmentation path with the highest probability is the one, and the probability of each path = the product of the probabilities of all words in this path; Determine the words and vocabulary based on the maximum probability path; determine the domain attribute of the text data based on the content of the text data; and / or determine the domain attribute of the text data based on the label of the text data; Input the text data and its corresponding domain attribute into a word vector model to obtain word vectors; Input the text data and its corresponding domain attribute into a vocabulary vector model to obtain vocabulary vectors; Determine the first weight corresponding to the words and the second weight corresponding to the vocabulary based on the text data; Generate a first sentence vector according to the first weight and the word vectors; Generate a second sentence vector according to the second weight and the vocabulary vectors; Concatenate the first sentence vector and the second sentence vector to generate a sentence semantic vector; Perform natural language processing on the real-time text data based on the sentence semantic vector.
2. The natural language processing method according to claim 1, characterized in that It further includes: Extract the sentence semantic vectors of multiple preset text data in the database; Compare the similarity between the text data and the multiple preset text data based on the sentence semantic vectors; Determine the target text data from the multiple preset text data according to the similarity comparison result.
3. The natural language processing method according to claim 1, characterized in that, It further includes: Train a deep neural network model based on multiple corpora with domain attributes to generate the word vector model; Train a shallow neural network model based on multiple corpora with domain attributes to generate the vocabulary vector model.
4. The natural language processing method according to claim 1, wherein Input the text data and its corresponding domain attribute into a word vector model to obtain word vectors, including: Input the text data and its corresponding domain attribute into the trained BERT model to generate word vectors; Optionally, input the text data and its corresponding domain attribute into a vocabulary vector model to obtain vocabulary vectors, including: Input the text data and its corresponding domain attribute into the trained Word2vec model to generate vocabulary vectors.
5. The natural language processing method according to claim 1, characterized in that Determine the first weight corresponding to the words and the second weight corresponding to the vocabulary based on the text data, including: Determine the first weight and the second weight based on the inverse document frequency corresponding to the words and the vocabulary in the text data.
6. A natural language processing device, characterized in that, Adopt the method according to any one of claims 1-5, including: A word segmentation module, configured to obtain a word segmentation dictionary; segment text data based on the word segmentation dictionary to generate a vocabulary network, where the vocabulary network is a directed acyclic graph; determine the maximum probability path in the vocabulary network based on the dynamic programming algorithm; determine words and vocabulary based on the maximum probability path; determine the domain attribute of the text data based on the content of the text data; and / or determine the domain attribute of the text data based on the label of the text data; A text module, configured to input the text data and its corresponding domain attribute into a text vector model to obtain word vectors; A vocabulary module, configured to input the text data and its corresponding domain attribute into a vocabulary vector model to obtain word vectors; A weight module, configured to determine a first weight corresponding to the text and a second weight corresponding to the vocabulary based on the text data; A vector module, configured to generate a first sentence vector according to the first weight and the word vectors; generate a second sentence vector according to the second weight and the word vectors; splice the first sentence vector and the second sentence vector to generate a sentence semantic vector; A semantic module, configured to perform natural language processing on the real-time text data based on the sentence semantic vector.
7. An electronic device, comprising a processor and a memory, wherein the memory is configured to store a computer executable program, and is characterized in that: When the computer executable program is executed by the processor, the processor executes the method according to any one of claims 1-5.
8. A computer-readable medium storing a computer-executable program, characterized in that, When the computer executable program is executed, the method according to any one of claims 1-5 is implemented.
Citation Information
Patent Citations
Method and device for processing session information and terminal equipment
CN110209774A
Semantic processing method and related device
CN110457689A
Chinese text matching method and system
CN111914067A
Text information processing method and terminal
CN112052331A