Language processing method and system constructed based on the assumption of word meaning distribution
By inputting the sentences into syntax model and semantic analysis model, combining the above-mentioned meaning to predict candidate phrases and match word components, the problem of difficult-to-understand ambiguousness of Chinese vocabulary in the prior art is solved, and more accurate and efficient language understanding is achieved.
Patent Information
- Application Number
- CN202111461699.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-02
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-02
AI Technical Summary
Existing language understanding machines have difficulty accurately understanding the ambiguity of Chinese vocabulary, especially in the absence of context.
By entering the sentence into the syntax model for preliminary sentence breaking, the word components are obtained, and they are input into the semantic analysis model one by one, and predict candidate phrases based on the above meaning, match them with the word components to give them meaning, and finally the meaning of the sentence is derived.
It realizes the accurate understanding and analysis of Chinese vocabulary ambiguity without relying on too many contexts, and improves the accuracy and efficiency of language comprehension.
Smart Images

Figure CN114254177B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network multimedia, and in particular, to a method and system for language processing constructed based on the assumption of word meaning distribution. Background Art
[0002] With the rapid development of the network, there is a need for automated machines that can quickly and accurately understand the meaning of language. However, existing language understanding machines are difficult to understand accurately, especially in the case of Chinese words with multiple meanings, and machines are even more difficult to handle. It is necessary to develop machines that can understand the meaning of words in combination with the context.
[0003] Therefore, there is an urgent need for a targeted method and system for language processing constructed based on the assumption of word meaning distribution. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and system for language processing constructed based on the assumption of word meaning distribution. By inputting a sentence into a syntactic model for preliminary sentence segmentation to obtain a first word component, inputting the first word component into a semantic analysis model one by one to obtain a second word component, obtaining the context of the current sentence, predicting the next candidate phrase based on the meaning of the context, matching it with the second word component, and assigning the meaning of the second word component according to the matching result, thereby obtaining the meaning of the sentence.
[0005] In a first aspect, the present application provides a method for language processing constructed based on the assumption of word meaning distribution, the method comprising:
[0006] Obtain network data streams, extract sentences therefrom, input the sentences into a syntactic model for preliminary sentence segmentation to obtain a first word component, the syntactic model sets extraction windows with different widths according to each word type, and uses the extraction window as the basis for sentence segmentation, and the words within the window width form the first word component;
[0007] Input the first word component into a semantic analysis model one by one. If it can still be recognized as a short sentence, it is determined that the preliminary sentence segmentation of the first word component is not successful, and the first word component needs to be input into the syntactic model again for sentence segmentation to obtain a second word component; if it cannot be recognized as a short sentence and is recognized as a phrase, it is determined that the preliminary sentence segmentation of the first word component is successful, and the first word component is directly marked as the second word component; the phrase is composed of several words and does not have a syntactic structure;
[0008] Set the context width to N, where N is a positive integer, obtain the context of the current sentence according to the context width, input the context into the semantic analysis model, analyze the meaning of the context and predict the next candidate phrase of the context, match the candidate phrase with the second word component, and assign the meaning of the second word component according to the matching result;
[0009] Among them, the matching refers to comparing each word in the candidate phrase with the words in the second word component one by one, calculating the number of identical words, and when the number is greater than a preset threshold, it is determined that the candidate phrase matches the second word component;
[0010] Recombine the second word component to form a new sentence and obtain the meaning of this new sentence.
[0011] Combined with the first aspect, in the first possible implementation manner of the first aspect, the setting of extraction windows with different widths according to each word type includes updating the word type and establishing a corresponding relationship between the new word type and the extraction window width.
[0012] Combined with the first aspect, in the second possible implementation manner of the first aspect, the semantic analysis model performs semantic analysis according to the requirements of sentence grammar.
[0013] Combined with the first aspect, in the third possible implementation manner of the first aspect, the cores of both the semantic analysis model and the syntactic model use neural network models.
[0014] In a second aspect, the present application provides a language processing system constructed based on the assumption of word meaning distribution. The system includes a processor and a memory:
[0015] The memory is used to store program code and transmit the program code to the processor;
[0016] The processor is used to execute the method described in any one of the four possibilities of the first aspect according to the instructions in the program code.
[0017] In a third aspect, the present application provides a computer-readable storage medium, which is used to store program code, and the program code is used to execute the method described in any one of the four possibilities of the first aspect.
[0018] The present invention provides a method and system for language processing constructed based on the assumption of word meaning distribution. By inputting a sentence into a syntactic model for preliminary sentence segmentation to obtain a first word component, inputting the first word component into a semantic analysis model one by one to obtain a second word component, obtaining the context of the current sentence, predicting the next candidate phrase according to the meaning of the context, matching it with the second word component, and endowing the meaning of the second word component according to the matching result, and then obtaining the meaning of the sentence. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 This is a flowchart of the method of the present invention. Detailed implementation manners
[0021] The preferred embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making the protection scope of the present invention more clearly defined.
[0022] Figure 1 This is a flowchart of the language processing method constructed based on the word sense distribution hypothesis provided by this application, including:
[0023] Obtain network data streams, extract sentences therefrom, input the sentences into a syntactic model for preliminary sentence segmentation to obtain first word components. The syntactic model sets extraction windows of different widths according to each word type, and uses the extraction window as the basis for sentence segmentation. The words within the window width form the first word components;
[0024] Input the first word components into the semantic analysis model one by one. If it can still be recognized as a short sentence, it is determined that the preliminary sentence segmentation of this first word component is not successful, and this first word component needs to be input into the syntactic model again for sentence segmentation to obtain second word components; if it cannot be recognized as a short sentence and is recognized as a phrase, it is determined that the preliminary sentence segmentation of this first word component is successful, and the first word component is directly marked as the second word component; the phrase is composed of several words and does not have a syntactic structure;
[0025] Set the width of the previous context as N, where N is a positive integer. Obtain the previous context of the current sentence according to the width of the previous context, input the previous context into the semantic analysis model, analyze the meaning of the previous context and predict the candidate phrases following the previous context, match the candidate phrases with the second word components, and assign the meaning of the second word components according to the matching result;
[0026] Among them, the matching refers to comparing the words in the candidate phrases with the words in the second word components one by one, calculating the number of identical words. When the number is greater than a preset threshold, it is determined that the candidate phrase matches the second word component;
[0027] Recombine the second word components to form a new sentence and obtain the meaning of this new sentence.
[0028] In some preferred embodiments, setting extraction windows of different widths according to each word type includes updating the word types and establishing a corresponding relationship between the new word types and the extraction window widths.
[0029] In some preferred embodiments, the semantic analysis model performs semantic analysis according to the requirements of sentence grammar.
[0030] In some preferred embodiments, neural network models are used for the kernels of both the semantic analysis model and the syntactic model.
[0031] This application provides a language processing system constructed based on the assumption of word meaning distribution. The system includes a processor and a memory:
[0032] The memory is used to store program codes and transmit the program codes to the processor;
[0033] The processor is used to execute the method described in any one of all the embodiments of the first aspect according to the instructions in the program codes.
[0034] This application provides a computer-readable storage medium. The computer-readable storage medium is used to store program codes, and the program codes are used to execute the method described in any one of all the embodiments of the first aspect.
[0035] In specific implementation, the present invention further provides a computer storage medium. The computer storage medium may store a program. When the program is executed, it may include some or all of the steps in various embodiments of the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (abbreviation: ROM) or a random access memory (abbreviation: RAM), etc.
[0036] Those skilled in the art can clearly understand that the technologies in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0037] For the same or similar parts among the various embodiments of this specification, reference can be made to each other. In particular, for the embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method embodiments.
[0038] The above-described embodiments of the present invention do not constitute a limitation on the protection scope of the present invention.
Claims
1. A language processing method constructed based on the assumption of word meaning distribution, characterized in that, the method includes: Obtain network data streams, extract sentences therefrom, input the sentences into a syntactic model for preliminary sentence segmentation to obtain a first word component. The syntactic model sets extraction windows of different widths according to each word type, and uses the extraction window as the basis for sentence segmentation. The words within the window width form the first word component; Input the first word component into the semantic analysis model one by one. If it can still be recognized as a short sentence, it is determined that the preliminary sentence segmentation of the first word component is not successful, and the first word component needs to be input into the syntactic model again for sentence segmentation to obtain a second word component. If it cannot be recognized as a short sentence and is recognized as a phrase, it is determined that the preliminary sentence segmentation of the first word component is successful, and the first word component is directly labeled as the second word component. The phrase consists of several words and does not have a syntactic structure; Set the width of the previous context as N, where N is a positive integer. Obtain the previous context of the current sentence according to the width of the previous context, input the previous context into the semantic analysis model, analyze the meaning of the previous context and predict the candidate phrases following the previous context, match the candidate phrases with the second word component, and assign the meaning of the second word component according to the matching result; wherein, the matching refers to comparing the words in the candidate phrase with the words in the second word component one by one, calculating the number of identical words, and when the number is greater than a preset threshold, it is determined that the candidate phrase matches the second word component; Recombine the second word component to form a new sentence and obtain the meaning of the new sentence.
2. The method according to claim 1, characterized in that: The setting of extraction windows of different widths according to each word type includes updating the word type and establishing a corresponding relationship between the new word type and the extraction window width.
3. The method according to any one of claims 1-2, characterized in that: The semantic analysis model performs semantic analysis according to the requirements of sentence grammar.
4. The method according to any one of claims 1-3, characterized in that: The kernels of the semantic analysis model and the syntactic model both use neural network models.
5. A language processing system constructed based on the assumption of word meaning distribution, characterized in that, the system includes a processor and a memory: The memory is used to store program codes and transmit the program codes to the processor; The processor is used to execute the method according to any one of claims 1-4 according to the instructions in the program codes.
6. A computer-readable storage medium, characterized in that, the computer-readable storage medium is used to store program codes, and the program codes are used to execute the method according to any one of claims 1-4.
Citation Information
Patent Citations
A text generation method and device based on artificial intelligence
CN109670185A
Information processing method based on natural language recognition, related equipment and storage medium
CN110334347A