Chinese sentence structure analysis method, device and storage medium

CN115293135BActive Publication Date: 2026-10-09林原
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210973274.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-15
Publication Date
2026-10-09
Estimated Expiration
2042-08-15

AI Technical Summary

Technical Problem

比如,中文语法要求某个词语一定是个名词,这在深度学习中就无法做到

Benefits of technology

[0033]The Chinese analyzers in the various embodiments of this application are innovative and possess controllable Chinese sentence structure analysis capabilities. When analyzing polysemous and ambiguous phenomena, the parse tree and all options for polysemous words and/or ambiguous phenomena adhere to Chinese grammar. Then, the Chinese analyzer translates these options into a foreign language, with differences translated as [MASK]. This method transforms the analysis of ambiguity and polysemous words into a fill-in-the-blank exercise, making the Chinese analyzer a controllable artificial intelligence method rather than a black-box algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115293135B_ABST
    Figure CN115293135B_ABST
Patent Text Reader

Abstract

The application relates to a Chinese sentence structure analysis method, device and storage medium. The Chinese sentence structure analysis method comprises the following steps: analyzing the structure of a Chinese sentence by a preset Chinese parser to generate a first parsing tree; determining an analysis object existing in the first parsing tree; recommending a single Chinese meaning of the analysis object by a foreign language model in a preset Chinese analyzer; and reanalyzing the structure of the Chinese sentence according to the single Chinese meaning of the analysis object by the Chinese parser to generate a second parsing tree. The Chinese analyzer has controllable Chinese sentence structure analysis function. In comparison, the prior art is completely uncontrollable. The application can improve the accuracy of the result, does not need to establish a huge corpus, solves a major technical problem for Chinese sentence structure analysis, and can improve the Chinese sentence structure analysis efficiency of human-computer interaction and reduce the artificial cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing, and in particular to a method, device and storage medium for analyzing Chinese sentence structure. Background Technology

[0002] Sentence structure analysis plays a crucial role in Natural Language Processing (NLP), serving as the foundation for machine translation, reading comprehension, information extraction, and question-answering systems. However, Chinese sentences contain numerous polysemous words and ambiguities. While humans can easily resolve these issues, developing a set of rules that computers can execute is extremely difficult.

[0003] Currently, computer analysis of sentence structure mainly employs rule-based methods, probabilistic statistics, and deep learning. Among these, rule-based methods attempt to establish a set of rules that can resolve polysemy and ambiguity and run on a computer. However, rule-based methods are ineffective for natural languages, especially languages ​​like Chinese which exhibit numerous polysemy and ambiguities; to date, no universally accepted and feasible set of rules has been found.

[0004] Furthermore, probability statistics and deep learning also face significant technical challenges. First, both require building corpora. This necessitates labeling a vast amount of Chinese sentences, accurately identifying the meaning of each polysemous word and highlighting any ambiguities within the sentences. Then, statistical or machine learning methods are applied to the corpus to build language models, which then resolve polysemy and ambiguity. However, the labeling process must be done manually, which is inefficient, difficult to automate, and requires substantial manpower, making it unacceptable for many companies. Second, probability statistics and deep learning, especially deep learning, are black-box algorithms—uncontrollable. For example, Chinese grammar requires a word to be a noun, which is impossible in deep learning. Current language models have hundreds of millions of parameters; it's impractical to adjust these parameters for the part of speech of a single word.

[0005] In conclusion, there is an urgent need to invent a new artificial intelligence method that, while adhering to Chinese grammar (i.e., the user can control the analysis process), without requiring the construction of a corpus or the labeling of large amounts of data, can analyze polysemy and ambiguity in Chinese sentences, providing accurate prompts to help users make final decisions. Furthermore, it could improve the efficiency of human-computer interactive Chinese sentence structure analysis systems and modify the final results. Summary of the Invention

[0006] To solve the above-mentioned technical problems, or at least partially solve them, this application provides a Chinese sentence structure analysis method, device, and storage medium.

[0007] Firstly, this application provides a method for analyzing the structure of Chinese sentences, the method comprising:

[0008] A first parse tree is generated by parsing the structure of a Chinese sentence according to Chinese grammar using a preset Chinese parser; and the parsing objects existing in the first parse tree are identified; the parsing objects are Chinese polysemous words and / or Chinese ambiguous structures.

[0009] The foreign language model in the preset Chinese analyzer recommends a single Chinese meaning for the parsed object; the foreign language is a non-Chinese language.

[0010] The Chinese parser re-parses the structure of the Chinese sentence according to the single Chinese meaning of the parsed object, generating a second parse tree.

[0011] Optionally, recommending a single Chinese meaning for the parsed object using a foreign language model in a preset Chinese analyzer includes:

[0012] The Chinese analyzer translates the Chinese sentence into a translation composed of a foreign language, thereby avoiding the need to build a corpus and label the data in the corpus.

[0013] Optionally, the ambiguous structure includes part-of-speech ambiguity and / or voice ambiguity.

[0014] Optionally, after generating a second parse tree by re-parseing the structure of the Chinese sentence according to the single Chinese meaning of the parsed object using the Chinese parser, the process includes:

[0015] The second parsing tree and / or the parsing object are displayed in a preset human-machine interface to prompt the user to confirm the single Chinese meaning of the second parsing tree and / or the parsing object;

[0016] Upon receiving an unconfirmed instruction, the human-machine interface receives a modification operation on the single Chinese meaning of the second parse tree and / or the parsed object.

[0017] Optionally, recommending a single Chinese meaning for the parsed object using a foreign language model in a preset Chinese analyzer includes:

[0018] The Chinese analyzer translates the Chinese sentence into a translation composed of the target language, and replaces the translated words corresponding to the parsed object with preset tags in the translation.

[0019] The markers are inferred as target words using the foreign language model.

[0020] Compare the meanings of the target word and the parsed object;

[0021] Based on the comparison results, a single Chinese meaning for the parsed object is recommended.

[0022] Optionally, the comparison of the meanings of the target word and the parsed object; based on the comparison result, recommending a single Chinese meaning of the parsed object, including:

[0023] The first parse tree is used to determine the multiple Chinese meanings of the parsed object;

[0024] The target word is compared with the multiple Chinese meanings respectively;

[0025] When the comparison result is the same as or similar to one of the Chinese meanings, that Chinese meaning is recommended as the single Chinese meaning of the parsed object.

[0026] No recommendation will be made if the comparison results show that the meanings are all different or not very similar.

[0027] Optionally, the Chinese analyzer includes multiple foreign language models, and the step of recommending a single Chinese meaning of the parsed object through the foreign language models in the preset Chinese analyzer further includes:

[0028] When the Chinese analyzer fails to recommend a single Chinese meaning for the parsed object, the next foreign language language model is invoked to re-recommend the single Chinese meaning of the parsed object according to the preset calling order of the multiple foreign language models.

[0029] Secondly, this application provides a Chinese sentence structure analysis device, which includes: a memory, a processor, and a computer program stored in the memory and capable of running on the processor;

[0030] When the computer program is executed by the processor, it implements the steps of the Chinese sentence structure analysis method described in any of the above descriptions.

[0031] Thirdly, this application provides a computer-readable storage medium storing a Chinese sentence structure analysis program, which, when executed by a processor, implements the steps of the Chinese sentence structure analysis method described in any of the above claims.

[0032] The technical solutions provided in this application have the following advantages compared with the prior art:

[0033] The Chinese analyzers in the various embodiments of this application are innovative and possess controllable Chinese sentence structure analysis capabilities. When analyzing polysemous and ambiguous phenomena, the parse tree and all options for polysemous words and / or ambiguous phenomena adhere to Chinese grammar. Then, the Chinese analyzer translates these options into a foreign language, with differences translated as [MASK]. This method transforms the analysis of ambiguity and polysemous words into a fill-in-the-blank exercise, making the Chinese analyzer a controllable artificial intelligence method rather than a black-box algorithm.

[0034] After the Chinese analyzers in each implementation of this application translate the Chinese sentences into translations composed of foreign languages, they can use open-source large-scale artificial intelligence language models, which improves the accuracy of the results without the need to build a costly corpus.

[0035] The embodiments of this application effectively solve two major technical problems associated with rule-based methods, probabilistic statistics, and deep learning: namely, 1) lack of controllability, and 2) the need to add annotations to a large number of Chinese sentences during the parsing process to build a corpus. They eliminate the need for machine learning to solve polysemous and ambiguous problems, improve the efficiency of Chinese sentence structure analysis in human-computer interaction, and reduce manual costs. Attached Figure Description

[0036] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 An ambiguous parse tree provided for various embodiments of this application;

[0039] Figure 2 The main flowchart of the Chinese sentence structure analysis method provided in various embodiments of this application;

[0040] Figure 3 A block diagram of the components of the Chinese sentence structure analysis system provided in various embodiments of this application;

[0041] Figure 4 A flowchart illustrating an optional Chinese sentence structure analysis method provided for various embodiments of this application;

[0042] Figure 5 Analysis flowcharts of the Chinese analyzers provided in various embodiments of this application;

[0043] Figure 6 Flowcharts illustrating the analysis of polysemous words provided in various embodiments of this application;

[0044] Figure 7 A parse tree for ambiguous phrases provided for various embodiments of this application;

[0045] Figure 8 A flowchart illustrating the analysis of ambiguous phrases provided in various embodiments of this application. Detailed Implementation

[0046] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0047] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.

[0048] Definitions of some technical terms

[0049] In the following description, Chinese polysemous words (hereinafter referred to as polysemous words) refer to a word or phrase that has multiple meanings, and the meaning of a polysemous word in a specific sentence must be determined based on the context. For example, the word "quality" is a polysemous word. It can mean "quality" or a physical quantity.

[0050] Ambiguous structures in Chinese (hereinafter referred to as ambiguous structures) refer to phrases or sentences that can be parsed into multiple parse trees, meaning they have multiple meanings. This can include part-of-speech ambiguity and / or voice ambiguity. For example, "China's longest tunnel" and "the most capable student."

[0051] On the surface, both "China's longest tunnel" and "the most capable student" appear to be 'noun-adjective-of-noun' phrases. However, "China's longest tunnel" means "This tunnel is the longest in China," where "China" acts as an adverb, meaning "within China." Similarly, "the most capable student" means "This student has the strongest ability," where "ability" is indeed a noun. Therefore, as... Figure 1 As shown, the parse trees for "China's longest tunnel" and "the most capable student" are different. The left side shows the parse tree for "China's longest tunnel," where "China" is an adverb, "longest" is an adjective, and "tunnel" is a noun. The right side shows the parse tree for "the most capable student," where "capable" is a noun, "most capable" is an adjective, and "student" is a noun.

[0052] Accurate parse trees play a great role in natural language understanding. For example, in machine translation, based on these parse trees, a computer can accurately translate the longest tunnel in China into "the longest tunnel in China", and translate the most capable student into "student with the strongest skill".

[0053] Embodiment 1

[0054] An embodiment of the present invention provides a method for analyzing Chinese sentence structure, as shown in Figure 2 , the method for analyzing Chinese sentence structure comprises:

[0055] S101, parsing the structure of a Chinese sentence according to Chinese grammar by a preset Chinese parser to generate a first parse tree; and determining parsing objects existing in the first parse tree; the parsing objects are Chinese polysemous words and / or Chinese ambiguous structures;

[0056] S102, recommending a single Chinese meaning of the parsing object by a foreign language language model in a preset Chinese analyzer; the foreign language is a non-Chinese language;

[0057] S103, re-parsing the structure of the Chinese sentence according to the single Chinese meaning of the parsing object by the Chinese parser to generate a second parse tree.

[0058] Composition of Chinese sentence structure analysis system

[0059] In some embodiments, the method for analyzing Chinese sentence structure can be implemented in the form of a system to form a Chinese sentence structure analysis system, as shown in Figure 3 , the system mainly comprises a Chinese parser, a Chinese AI analyzer and a human-machine interface.

[0060] Wherein, the Chinese parser can adopt standard parser algorithms (for example, LL algorithm, LR algorithm, CYK algorithm, etc.) to analyze the structure of Chinese sentences. The analysis result is a parse tree of the Chinese sentence, and the polysemous words and ambiguous structures in the sentence are pointed out. The Chinese analyzer (hereinafter also described as the Chinese AI analyzer) mainly analyzes and solves ambiguities and polysemous words in Chinese sentences, and this component uses one or more foreign language language models when working, for example, Google BERT language model.

[0061] This invention provides accurate parse trees for Chinese sentences by parsing the structure of Chinese sentences using a Chinese parser, generating a first parse tree. It then identifies parsing objects within the first parse tree and further recommends single Chinese meanings for these objects using a foreign language model within the Chinese parser. This allows the Chinese parser to re-parse the Chinese sentence structure according to the single Chinese meaning of the parsing object, generating a second parse tree. This provides accurate parse trees for Chinese sentences while adhering to Chinese grammar (i.e., the user can control the analysis process) and without requiring annotations in a large number of Chinese sentences or the establishment of a corpus. It improves the efficiency of human-computer interaction in Chinese sentence structure analysis and reduces manual costs. It effectively solves two major technical problems inherent in rule-based methods, probabilistic statistics, and deep learning: 1) lack of control, and 2) the need to add annotations to a large number of Chinese sentences during the parsing process to build a corpus.

[0062] In this embodiment of the invention, the Chinese parser analyzes sentence structure entirely according to Chinese grammar, identifying ambiguous and polysemous words. If Chinese grammar requires a word to be a noun, the parser confirms that the word is a noun; therefore, this embodiment of the invention is a controllable AI method. In contrast, existing AI methods have hundreds of millions of parameters, making it impossible to find just a few parameters to determine the part of speech of a word.

[0063] The Chinese AI analyzer can analyze and resolve ambiguities and polysemous words in Chinese sentences. When the component operates, it uses one or more artificial intelligence language models, which can employ existing third-party software, such as the Google BERT language model. In this embodiment, the Chinese AI analyzer uses two language models: one for English and another for a different foreign language, such as French, Spanish, etc.

[0064] This implementation method determines the number of foreign language models in the Chinese sentence analysis system based on accuracy requirements. After the Chinese parser detects polysemy or ambiguity, the Chinese AI analyzer analyzes it. The analysis result may be 1) a definite word meaning or a definite sentence structure, or 2) indeterminate. If the system has multiple foreign language models, the English model can be used first. If the English model arrives at a definite conclusion, its conclusion is used; if it cannot arrive at a definite conclusion, other language models, such as the French language model, can be used. If the French model provides a definite conclusion, its conclusion is adopted. In other words, the more foreign language models the analyzer has, the higher the accuracy. In practice, one or two language models are usually used.

[0065] This embodiment uses the Google BERT language model as an example to illustrate the technical solution, but it does not mean that the present invention must use the Google BERT language model. Of course, models such as OpenAI GPT-3 can also be used in the present invention.

[0066] In some implementations, a single Chinese meaning of the parsed object is recommended through a foreign language model in a preset Chinese analyzer, including:

[0067] The Chinese analyzer translates the Chinese sentence into a translation composed of a foreign language; and recommends a single Chinese meaning for the parsed object based on the translation.

[0068] Detailed process of Chinese sentence structure analysis system

[0069] For example, such as Figure 4 As shown, the Chinese sentence structure analysis method provided in this embodiment includes:

[0070] Step 1. The structure of a Chinese sentence is parsed by a preset Chinese parser to generate a first parse tree; and the parsing objects existing in the first parse tree are identified; the parsing objects are Chinese polysemous words and / or Chinese ambiguous structures; in other words, the Chinese parser parses the Chinese sentence, generates a parse tree and finds the ambiguity and polysemous words in the sentence, and then sends these results to the Chinese AI analyzer.

[0071] Step 2. The Chinese AI analyzer uses an English language model to analyze ambiguities and polysemous words in Chinese sentences;

[0072] Step 3. If the analysis in step 2 yields a definite result, proceed to step 5; otherwise, proceed to step 4; that is, recommend a single Chinese meaning of the parsed object through the foreign language model in the preset Chinese analyzer;

[0073] Step 4. The Chinese AI analyzer uses other language models (such as a French language model) to analyze ambiguities and polysemous words in Chinese sentences. Optionally, the Chinese analyzer can include multiple foreign language models. When the current Chinese analyzer fails to recommend a single Chinese meaning for the parsed object, it calls the next foreign language model according to a pre-defined calling order to re-recommend the single Chinese meaning of the parsed object. The analysis of ambiguous structures fully adheres to Chinese grammar and is completely controllable. For example, if Chinese grammar determines that a word is a noun, the Chinese parser will confirm that word as a noun; the analyzer only analyzes words with polysemy or ambiguity if the grammar allows them.

[0074] Step 5. The Chinese parser uses the results of steps 2 and 4 to eliminate ambiguity and polysemous words, and then re-parses the Chinese sentence;

[0075] Step 6. The parse tree and the system's analysis of ambiguous and polysemous words are displayed on the human-computer interface. If the user is not satisfied with these analyses, they can modify them through the interface.

[0076] Optionally, the second parsing tree and / or the single Chinese meaning of the parsing object are displayed on a preset human-machine interface to prompt the user to confirm the single Chinese meaning of the second parsing tree and / or the parsing object; when a confirmation instruction is received, the user receives a modification operation on the single Chinese meaning of the second parsing tree and / or the parsing object on the human-machine interface.

[0077] In this implementation, a significant innovation is the use of English and other foreign language models to analyze ambiguity and polysemy in Chinese sentences. Current AI methods require training on massive amounts of corpora. This training can be fully automated, such as with the Google BERT language model, or it can require manual labeling. Ambiguity and polysemy in Chinese sentences, regardless of their multiple meanings, share the same morphology within the Chinese corpus. In other words, if AI is trained using Chinese corpora to analyze ambiguity and polysemy, manual labeling is essential; otherwise, training is impossible. This approach is extremely labor-intensive and cannot meet the requirements of large-scale language models. AI trained on Chinese corpora cannot resolve ambiguity and polysemy in Chinese. In this implementation, by comparing the morphologies of Chinese ambiguity and polysemy in foreign languages ​​and linking them to the translated text, language models such as Google BERT can be used to analyze Chinese ambiguity and polysemy.

[0078] Secondly, a pre-defined Chinese parser analyzes the structure of Chinese sentences, generating a first parse tree; and identifies ambiguous and polysemous words in the sentence. These analyses all adhere to Chinese grammar and are completely controllable. If Chinese grammar requires a word to be a noun, the Chinese parser confirms that the word is a noun. Only if Chinese grammar allows a word to be ambiguous or polysemous will the Chinese parser send that word to the Chinese analyzer for further analysis. In other words, the Chinese sentence structure analysis method of this invention is a controllable artificial intelligence method. In comparison, current deep learning methods have hundreds of millions of parameters and cannot adhere to the specific requirements of Chinese grammar.

[0079] Finally, this implementation uses multiple foreign language models to further improve the accuracy of sentence analysis.

[0080] The process of Chinese AI analyzer

[0081] In this embodiment of the invention, the Chinese AI analyzer is mainly used to eliminate ambiguity and polysemy. In some implementations, the process of the Chinese analyzer is as follows: Figure 5 As shown, it may include:

[0082] Step 1. Assume that an ambiguous or polysemous word has several Chinese meanings, for example, Chinese meanings A and B. The analyzer translates the sentence into English or other foreign languages ​​according to Chinese meanings A and B respectively, and identifies words that are translated differently in different translations. For example, the word in the translation of Chinese meaning A is foreign language word X, and the word in the translation of meaning B is foreign language word Y. Then, it is replaced with [MASK] in the translation. That is, the Chinese analyzer translates the Chinese sentence into a translation composed of the target language (English or other foreign languages), and replaces the translated word corresponding to the parsed object with the preset marker [MASK] in the translation.

[0083] Step 2. Run the English or other foreign language model to infer [MASK] as word C; that is, use the foreign language model to infer the token as the target word (word C);

[0084] Step 3. Use AI methods to determine whether the meanings of words C and X are similar; in other words, compare the meanings of the target word and the parsed object.

[0085] Step 4. If word C and X have similar meanings, recommend A; if C and Y have different meanings, recommend B; if the AI ​​method cannot determine the relationship between C and X, the analyzer's result is uncertain. That is, based on the comparison results, the single Chinese meaning of the parsed object is recommended.

[0086] In this embodiment, the first parsing tree is used to determine the multiple Chinese meanings of the parsing object; the target word is compared with each of the multiple Chinese meanings; when the comparison result is the same as or similar to one of the Chinese meanings, the Chinese meaning is recommended as the single Chinese meaning of the parsing object; when the comparison result is that the meanings are all different or not similar, no recommendation is made.

[0087] The following two examples illustrate the workflow of the Chinese AI analyzer.

[0088] Example 1

[0089] Example 1 primarily analyzes polysemous words. There are many polysemous words in Chinese. For example,

[0090] These shoes are of very good quality.

[0091] In this example, 'quality' is a polysemous word; it can refer to quality or a physical quantity.

[0092] like Figure 6 As shown, the Chinese AI analyzer uses the following steps to determine word meaning.

[0093] Step 1. The Chinese AI analyzer finds that there are two options for "质量" in the sentence "这种鞋的质量很好": Option A: "质量 refers to quality", or Option B: "质量 refers to physical quantity". Translate the sentence "这种鞋的质量很好" into English, and replace "质量" with [MASK]. The English phrase is 'the[MASK]of this kind of shoes is very good.'

[0094] Step 2. Run an English language model, for example, Google BERT English, to obtain the result that [MASK] is 'quality'.

[0095] Step 3. Compare the translation 'quality' of the ambiguous word '质量' with 'quality' by artificial intelligence method, and find that their meanings are consistent.

[0096] Step 4. Recommend to select Option A, that is, '质量' refers to quality.

[0097] Example 2

[0098] Example 2 mainly analyzes phrase ambiguity. In computer science, ambiguity means that a sentence or phrase has multiple parse trees (meanings). For example, "the longest tunnel in China" is a "noun-adjective-de-noun" phrase, as Figure 7 shown, this phrase may have two different parse trees. "中国" may be a noun, the name of a country; it may also mean "in China", which is equivalent to an adverb.

[0099] Option A: "中国" is a noun, "最长" is an adjective, and "隧道" is a noun.

[0100] Option B: "中国" acts as an adverb, "最长" acts as an adjective, and "隧道" acts as a noun.

[0101] As Figure 8 shown, the process of the Chinese AI analyzer is as follows.

[0102] Step 1. The Chinese AI analyzer finds that the difference between the two parse trees of "中国最长的隧道" is A: "中国 is a noun" or B: "中国 is an adverb"; it translates the phrase "中国最长的隧道" into English, and replaces "中国" with [MASK]. The English phrase is 'the tunnel with the longest[MASK]'.

[0103] Step 2. Run an English language model, for example, Google BERT English, to obtain the result that [MASK] is 'span' or 'entrance'.

[0104] Step 3. Comparing the translation "China" of "span" and "China" by an artificial intelligence method, and then comparing "entrance" and "China", it is found that their meanings are neither the same nor similar.

[0105] Step 4. Issuing a prompt, recommending to select B, "中国" is an adverb.

[0106] Embodiment 2

[0107] An embodiment of the present invention provides a Chinese sentence structure analysis device, the Chinese sentence structure analysis device comprises: a memory, a processor, and a computer program stored on the memory and executable on the processor;

[0108] when the computer program is executed by the processor, the steps of the Chinese sentence structure analysis method according to any one of Embodiment 1 are implemented.

[0109] Embodiment 3

[0110] An embodiment of the present invention provides a computer-readable storage medium, a Chinese sentence structure analysis program is stored on the computer-readable storage medium, and when the Chinese sentence structure analysis program is executed by a processor, the steps of the Chinese sentence structure analysis method according to any one of Embodiment 1 are implemented.

[0111] For the specific implementation of Embodiment 2 to Embodiment 3, reference may be made to Embodiment 1, which has corresponding technical effects.

[0112] It should be noted that, in this document, the terms "comprise", "include" or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. Without more restrictions, an element defined by the sentence "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0113] The serial numbers of the embodiments of the present invention described above are for description only, and do not represent the advantages or disadvantages of the embodiments.

[0114] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0115] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A method for analyzing the structure of Chinese sentences, characterized in that, The Chinese sentence structure analysis method includes: A first parse tree is generated by parsing the structure of a Chinese sentence according to Chinese grammar using a preset Chinese parser; and the parsing objects existing in the first parse tree are identified; the parsing objects are Chinese polysemous words and / or Chinese ambiguous structures. The step of recommending a single Chinese meaning of the parsed object by using a foreign language model in a preset Chinese analyzer specifically includes: translating the Chinese sentence into a translation composed of the target language using the Chinese analyzer, and replacing the translated word corresponding to the parsed object with a preset marker in the translation; inferring the marker as a target word using the foreign language model; comparing the meaning of the target word and the parsed object; and recommending a single Chinese meaning of the parsed object based on the comparison result; wherein the foreign language is a non-Chinese language. The Chinese parser re-parses the structure of the Chinese sentence according to the single Chinese meaning of the parsed object, generating a second parse tree.

2. The Chinese sentence structure analysis method according to claim 1, characterized in that, The method of recommending a single Chinese meaning for the parsed object through a foreign language model in a preset Chinese analyzer includes: The Chinese analyzer translates the Chinese sentence into a translation composed of a foreign language to avoid building a corpus and labeling the data in the corpus; and recommends a single Chinese meaning for the parsed object based on the translation.

3. The Chinese sentence structure analysis method according to claim 1, characterized in that, The ambiguous structures include part-of-speech ambiguity and / or voice ambiguity.

4. The Chinese sentence structure analysis method according to claim 1, characterized in that, After the Chinese parser re-parses the structure of the Chinese sentence according to the single Chinese meaning of the parsed object to generate a second parse tree, the process includes: The second parsing tree and / or the parsing object are displayed in a preset human-machine interface to prompt the user to confirm the single Chinese meaning of the second parsing tree and / or the parsing object; Upon receiving an unconfirmed instruction, the human-machine interface receives a modification operation on the single Chinese meaning of the second parse tree and / or the parsed object.

5. The Chinese sentence structure analysis method according to claim 1, characterized in that, The meaning of the target word and the parsed object are compared; Based on the comparison results, the recommended single Chinese meaning of the parsed object includes: The first parse tree is used to determine the multiple Chinese meanings of the parsed object; The target word is compared with the multiple Chinese meanings respectively; When the comparison result is the same as or similar to one of the Chinese meanings, that Chinese meaning is recommended as the single Chinese meaning of the parsed object. No recommendation will be made if the comparison results show that the meanings are all different or not very similar.

6. The Chinese sentence structure analysis method according to claim 1, characterized in that, The Chinese analyzer comprises multiple foreign language models.

7. The Chinese sentence structure analysis method according to claim 1, characterized in that, The method of recommending a single Chinese meaning for the parsed object through a foreign language model in a preset Chinese analyzer also includes: If the current Chinese analyzer does not recommend a single Chinese meaning of the parsed object, the next foreign language model is called to re-recommend the single Chinese meaning of the parsed object according to the preset calling order of the multiple foreign language models in the Chinese analyzer.

8. A Chinese sentence structure analysis device, characterized in that, The Chinese sentence structure analysis device includes: a memory, a processor, and a computer program stored in the memory and capable of running on the processor; When the computer program is executed by the processor, it implements the steps of the Chinese sentence structure analysis method as described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a Chinese sentence structure analysis program, which, when executed by a processor, implements the steps of the Chinese sentence structure analysis method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Establishing device and method for multilingual dictionary

    CN102789461A

  • Grammar analysis method and device, equipment and storage medium

    CN113901798A