A multilingual semantic parsing system
The multi-language semantic parsing system addresses the challenge of adapting to language variations by dynamically adjusting grammar rules and incorporating error evaluation, improving accuracy and efficiency in handling diverse languages for real-time translation and international content management.
Patent Information
- Application Number
- CN202510223960.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Existing multilingual semantic analytical systems are difficult to adapt to the natural changes and diversity of languages in multilingual environments, especially when there are fewer languages or dialects, which leads to poor efficiency and accuracy of information retrieval and content management and cannot meet international needs.
Through the context monitoring module, the syntax rules are adjusted, and the syntax rules are optimized to improve the processing accuracy and adaptability of multilingual inputs.
It improves the efficiency and user experience of multilingual semantic analytical systems in cross-language applications, and reduces parsing errors caused by language diversity, especially in automatic translation and internationalized language content.
Smart Images

Figure CN119721054B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of semantic parsing, and particularly to a multilingual semantic parsing system. Background Art
[0002] Semantic parsing is a core technology in natural language processing (NLP), which mainly focuses on understanding and interpreting the meaning in natural language. This technology involves converting language input into a form that can be understood by a computer, usually a structured data format, such as a logical expression, a dependency tree, or a semantic framework. The goal of semantic parsing is to accurately capture the intention and information content of a sentence or text, enabling the computer to perform specific tasks based on this information, such as question-and-answer systems, machine translation, speech recognition, and information retrieval, etc.
[0003] Among them, a multilingual semantic parsing system is an application system for understanding and processing semantic parsing technologies of multiple languages. The system can accept inputs in multiple languages, parse the unique semantic structures and grammar rules of each language through specialized algorithms, and convert them into a standardized semantic representation form. Its main uses include improving the efficiency and effectiveness of cross-language information retrieval, content management, automatic translation, and international software development, and it is particularly suitable for multilingual environments, such as international enterprises, multilingual social media analysis, and global news aggregation, enabling users to access, analyze, and understand a large amount of text data across language boundaries.
[0004] The prior art is often limited by its reliance on fixed syntactic rules and a single semantic structure processing strategy, resulting in difficulty in effectively adapting to the natural variations and diversities of languages. Especially when dealing with less commonly used languages or dialects, it is unable to accurately capture the language expressions in specific regions and cultures, and the parsing accuracy is insufficient. Existing systems show obvious processing delays and strategy rigidities in the face of real-time updated language data, and they perform poorly especially in the real-time analysis of multilingual social media and global news content, resulting in difficulties for information retrieval and content management to meet the rapidly developing international requirements in terms of accuracy and efficiency, and increasing the difficulty for users in cross-cultural communication and data parsing. Summary of the Invention
[0005] The purpose of the present invention is to solve the deficiencies existing in the prior art, and to propose a multilingual semantic parsing system.
[0006] To achieve the above purpose, the present invention adopts the following technical solution: A multilingual semantic parsing system, the system includes:
[0007] A context monitoring module obtains the target language category, core semantic units, syntactic structure features, context-related words, and historical dialogue data of the input text, calculates the similarity between the core semantic units and the historical dialogue data, screens out the fluctuating key texts, and obtains the context change classification result;
[0008] Based on the context change classification result, the syntactic rule adjustment module analyzes the applicable language scope and parsing effect of the rules, calculates the adaptation degree of the current text, adjusts the priority of the syntactic rules, and obtains a set of updated syntactic rules;
[0009] The semantic structure matching module calls the set of updated syntactic rules, analyzes the text structure according to the part-of-speech distribution, semantic dependency relationship and syntactic boundary of the input text, matches the sentence pattern, and calculates the text structure matching degree to obtain the grammatical structure matching degree;
[0010] For the grammatical structure matching degree, the syntactic error detection module analyzes the parsing error range, compares the error amount between the parsing result and the expected syntactic pattern, and identifies abnormal syntactic rules to obtain the syntactic error evaluation value;
[0011] The syntactic optimization module calls the syntactic error evaluation value, calculates the feedback influence ratio according to the user's error annotation and correction information, optimizes the applicable direction of the syntactic rule library, and obtains the syntactic rule tuning result.
[0012] The improvements of the present invention are that the context change classification result includes language style, semantic variation, and interaction mode, the set of updated syntactic rules is specifically rule adaptability, applicable scope, and priority setting, the grammatical structure matching degree specifically refers to matching accuracy, pattern consistency, and structural integrity, the syntactic error evaluation value includes error degree, error type, and influence interval, and the syntactic rule tuning result includes rule adjustment effect, application frequency change, and user feedback result.
[0013] The improvements of the present invention are that the context monitoring module includes:
[0014] The text feature extraction sub-module obtains the input text, detects the target language category of the input text, extracts the core semantic units, and identifies the syntactic structure features. At the same time, it collects context-related words and historical dialogue data to establish a text feature data set;
[0015] Based on the text feature data set, the context change analysis sub-module calculates the similarity between the core semantic units and the historical dialogue data, analyzes the context change of the input text in the target language category, and uses the formula:
[0016] ;
[0017] Calculate the context offset degree of each semantic unit , and identify the key points of context change, where represents the context weight value of the th core semantic unit of the input text, represents the The context weight value of a core semantic unit represents the context criticality weight of the th core semantic unit, represents the total number of core semantic units;
[0018] Based on the context deviation degree, the fluctuation key screening sub-module analyzes the input text of the fluctuation key, compares the change trends of the differential input texts, screens the input text of the fluctuation key, and obtains the context change classification result.
[0019] The present invention is improved in that the syntactic rule adjustment module includes:
[0020] Based on the context change classification result, the rule applicability analysis sub-module analyzes the applicable language range of each rule according to the rules in the syntactic rule library, analyzes the rule parsing effect, and obtains the rule adaptation degree data;
[0021] Based on the rule adaptation degree data, the rule priority adjustment sub-module calculates the adaptation degree of the current text context category to each rule, using the formula:
[0022] ;
[0023] to obtain the application priority of the syntactic rule , where represents the adaptation degree of the th syntactic rule, represents the weight of the th syntactic rule, represents the key feature value of the current text context category, represents the th syntactic rule, represents the total number of syntactic rules;
[0024] Based on the application priority of the syntactic rule, the application proportion update sub-module calculates the application proportion of the differential rules, adjusts the application frequency of the low-adaptation rules, and at the same time increases the application proportion of the high-adaptation rules, screens the rules with key priorities, and obtains the syntactic rule update set.
[0025] The present invention is improved in that the semantic structure matching module includes:
[0026] The text structure parsing sub-module calls the syntactic rule update set, analyzes the part-of-speech categories in the text according to the part-of-speech distribution, semantic dependency relationship and syntactic boundary of the input text, analyzes the hierarchical relationship of the syntactic structure, and extracts the key information of the core verb, subject and object to establish the text structure data;
[0027] The syntactic pattern matching sub-module extracts syntactic structure information based on the text structure data, matches the corresponding syntactic patterns, and uses the formula:
[0028] ;
[0029] Calculate the syntactic matching deviation , and obtain the syntactic matching error data. Among them, represents the syntactic structure feature value of the input text, represents the feature value of the standard syntactic pattern, represents the weight of the syntactic feature item, represents the dimensionality of the text syntactic features;
[0030] The structure matching degree calculation sub-module calculates the text structure matching degree based on the syntactic matching error data, sets the matching degree threshold range, analyzes the matching situation of different syntactic patterns, and obtains the syntactic structure matching degree.
[0031] The improvement of the present invention is that the syntactic error detection module includes:
[0032] The syntactic structure matching sub-module obtains the parsing error range of the input text according to the syntactic structure matching degree, calculates the syntactic matching degree of the input text through the syntactic rule library, and obtains the syntactic matching degree value;
[0033] The parsing error calculation sub-module compares the error amount between the input text and the expected syntactic pattern based on the syntactic matching degree value, filters out the syntactic rules beyond the normal range, and uses the formula:
[0034] ;
[0035] Calculate the parsing error value HE. Among them, represents the number of syntactic structures in the text, represents the th syntactic matching degree of the input text, represents the th expected syntactic matching degree, represents the th weight of the syntactic structure;
[0036] The syntactic evaluation sub-module calls the parsing error value, classifies the error items, judges the type of syntactic error, analyzes the influence range of the error, and obtains the syntactic error evaluation value.
[0037] The improvement of the present invention is that the syntactic optimization module includes:
[0038] The error evaluation sub-module calls the syntactic error evaluation value, compares it with the user's error annotation information, and uses the formula:
[0039] ;
[0040] Calculate the syntactic error deviation amount marked by the user , and screen out the influencing factors, where represents the evaluated error value, represents the error value marked by the user, represents the total number of sentences;
[0041] Based on the syntactic error deviation amount and combined with the user's correction information, the rule adjustment sub-module calculates the influence ratio feedback by the user, optimizes the application direction of the syntactic rule library according to the influence ratio, and adjusts the application ratio of the syntactic rules to obtain the syntactic rule tuning result.
[0042] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0043] In the present invention, by monitoring the context change in real time and adjusting the syntactic rules accordingly, the processing accuracy of multi-language input is improved. By calculating the similarity between fine-grained semantic units and historical dialogue data, the semantic parsing system can sensitively adapt to the subtle changes in the language. The dynamic adjustment of the syntactic rules changes according to the specific context of the text, optimizing the adaptability to complex language structures. In addition, an error real-time evaluation and adjustment mechanism is introduced, effectively reducing the parsing errors caused by language diversity, and improving the efficiency and user experience of cross-language applications, especially in the aspects of automatic translation and internationalized language content. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 is the system flow chart of the present invention;
[0045] Figure 2 is the flow chart of the context monitoring module in the present invention;
[0046] Figure 3 is the flow chart of the syntactic rule adjustment module in the present invention;
[0047] Figure 4 is the flow chart of the semantic structure matching module in the present invention;
[0048] Figure 5 is the flow chart of the syntactic error detection module in the present invention;
[0049] Figure 6 is the flow chart of the syntactic optimization module in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0051] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, in the description of the present invention, the meaning of "a plurality of" is two or more, unless otherwise specifically defined.
[0052] Embodiment: Please refer to Figure 1 , the present invention provides a technical solution: A multilingual semantic parsing system includes:
[0053] The context monitoring module obtains the target language category, core semantic units, syntactic structure features, context-related words, and historical dialogue data of the input text, analyzes the context changes of the input text in the target language category, calculates the similarity between the core semantic units and the historical dialogue data, identifies the key points of context changes, and screens the input text with significant fluctuations to obtain the context change classification result;
[0054] The syntactic rule adjustment module, based on the context change classification result, analyzes the applicable language range and parsing effect of each rule according to the rules in the syntactic rule library, calculates the adaptation degree to the current text context category, adjusts the application priority of the syntactic rules, and updates the application ratio of the syntactic rules according to the high and low adaptation degrees to obtain the syntactic rule update set;
[0055] The semantic structure matching module calls the syntactic rule update set, parses the text structure according to the part-of-speech distribution, semantic dependency relationship, and syntactic boundary of the input text, matches the corresponding sentence patterns, and calculates the structure matching degree of the text to obtain the grammar structure matching degree;
[0056] The syntactic error detection module, for the grammar structure matching degree, obtains the parsing error range of the input text, compares the error amount between the parsing result and the expected syntactic pattern, identifies the syntactic rules with parsing errors exceeding the normal range, and analyzes the error type and its influence range to obtain the syntactic error evaluation value;
[0057] The syntactic optimization module calls the syntactic error evaluation value, calculates the influence ratio of the user feedback based on the user error annotation and correction information, optimizes the applicable direction of the syntactic rule library, and adjusts the application ratio of the syntactic rules to obtain the syntactic rule tuning result.
[0058] The results of context change classification include language style, semantic variation, and interaction mode. The syntactic rule update set specifically includes rule adaptability, scope of application, and priority setting. The grammatical structure matching degree specifically refers to matching accuracy, pattern consistency, and structural integrity. The syntactic error assessment value includes error degree, error type, and impact range. The syntactic rule tuning results include rule adjustment effects, changes in application frequency, and user feedback results.
[0059] See also Figure 2 , the context monitoring module includes:
[0060] The text feature extraction submodule obtains the input text, detects the target language category of the input text, extracts the core semantic units, and identifies the syntactic structure features. It also collects context-related words and historical conversation data to establish a text feature dataset.
[0061] First, detect the target language category of the input text, and perform a comparative analysis based on character distribution, vocabulary features, and grammatical structure. If the character set contains a high proportion of pinyin characters and conforms to the English spelling structure, it can be determined to be an English text. If the character set contains Chinese characters and the word collocation conforms to Chinese grammar, it can be determined to be a Chinese text. For multilingual texts, classify them according to the language category with the highest proportion. For example, if the input text is "Hello, 你好", its English characters account for 50% and Chinese characters account for 50%. The dominant language can be determined based on the context information, and the core semantic units can be extracted. Different lexical analysis strategies are used for different language categories. For example, in Chinese, the core semantic units can be extracted through word segmentation technology, such as "artificial intelligence technology" can be split into "artificial intelligence" and "technology", while in English, they are extracted based on part-of-speech tagging, such as "m The core semantic units in "achinelearningmodel" can be "machinelearning" and "model". The syntactic structure features are identified, and syntactic tree structures are established for different languages. In Chinese text, "I like artificial intelligence" can be parsed as a subject-predicate-object structure, in which "I" is the subject, "like" is the verb, and "artificial intelligence" is the object. In English text, "IloveAI" can also be parsed as a subject-predicate-object structure. At the same time, contextual related words and historical dialogue data are collected. Based on the information of the previous and next sentences, the contextual relevance is judged. For example, in the historical dialogue data "Do you like technology? I like artificial intelligence", "artificial intelligence" and "technology" have semantic associations and can be used as contextual related words to ensure that the text feature data can fully reflect the semantic and syntactic characteristics of the input text and generate a text feature dataset.
[0062] The context change analysis submodule calculates the similarity between the core semantic unit and the historical dialogue data based on the text feature dataset, and analyzes the context changes of the input text in the target language category using the formula:
[0063] ;
[0064] Calculate the context offset degree of each semantic unit , identify the key points of context change, where represents the context weight value of the th core semantic unit of the input text, represents the context weight value of the th core semantic unit in the historical conversation data, represents the context criticality weight of the th core semantic unit, represents the total number of core semantic units;
[0065] By calculating the deviation degree between the weight value of the core semantic unit in the input text and the weight value of the corresponding semantic unit in the historical conversation data, judge the context offset degree. If the weight of a certain core semantic unit in the current text is much higher than that in the historical conversation data, it means that the importance of this word in the context has changed, and the value can be obtained through data monitoring and sampling; Context change calculation parameter table:
[0066] ;
[0067] As shown in the above table, the weight of "AI" in the input text is higher than the weight of the historical text , and its context criticality weight , substitute the parameters into the formula for calculation:
[0068] ;
[0069] ;
[0070] ;
[0071] ;
[0072] Calculate the context offset degree , the larger the value, the more obvious the context change. This result shows that the importance of "AI" in the input text has increased significantly compared with the historical conversation data, resulting in a higher overall context offset degree. Identify the key points of context change and obtain the context offset degree.
[0073] The fluctuation key screening sub-module analyzes the input text with key fluctuations based on the context offset degree, compares the change trends of different input texts, screens the input text with key fluctuations, and obtains the context change classification result;
[0074] Analyze the key input text of fluctuations, calculate the context change trend of different input texts. If the context deviation degree of a certain input text is much higher than the average level, it indicates that this text has prominent context change characteristics. For example, in multiple dialogue data samples, the average context deviation degree is 0.5, while the context deviation degree of the current input text is 0.81, and the deviation degree exceeds 60%. Then it can be determined that the context change of this text is large. Screen the key input text of fluctuations as shown in the following table:
[0075] Table of context deviation degree analysis:
[0076] ;
[0077] The context deviation degree of text C is much higher than that of other texts, so it is determined as a key text of fluctuations. By calculating the context deviation degree trend of different texts, a difference comparison curve is established to judge which texts have a change amplitude exceeding the set threshold. For example, if the set threshold is the average context deviation degree + 30%, then calculate: , where the context deviation degree of text C , so it is determined as a key text of fluctuations, and the context change classification result is obtained.
[0078] Please refer to Figure 3 , the syntactic rule adjustment module includes:
[0079] The rule applicability analysis sub-module, based on the context change classification result, according to the rules in the syntactic rule library, analyzes the applicable language range of each rule, analyzes the rule parsing effect, and obtains the rule fitness data;
[0080] For different language categories, extract the syntactic structure features corresponding to the rules. For example, some rules are only applicable to the subject-verb-object structure in Chinese, while some rules are applicable to the passive voice pattern in English. In specific analysis, obtain the syntactic structure data of the rules. For example, the Chinese syntactic rules adopt the form of "subject + verb + object", while English adopts "subject + passive verb + by + agent". When analyzing the applicability of the rules, calculate the application frequency of each rule in different language categories, and combine the classified context change results to judge the context category of the current text to determine the fitness of the rule. Subsequently, analyze the parsing effect of each rule. By calculating the parsing success rate, misjudgment rate, and coverage rate of the rule, measure its parsing ability in different context categories. For example, if the parsing success rate of a certain rule in Chinese texts reaches 90%, but the misjudgment rate in English texts is as high as 50%, it indicates that the fitness of this rule in Chinese is relatively high, while the fitness in English is relatively low. To further quantify the fitness, compare the matching degree between the syntactic structure after rule parsing and the actual structure of the text, calculate the matching rate, and perform weighted operations in combination with the misjudgment rate to obtain the rule fitness data.
[0081] The rule priority adjustment sub-module calculates the degree of fit between the current text context category and each rule based on the rule fitness data, using the formula:
[0082] ;
[0083] to obtain the application priority of syntactic rules , where represents the fitness of the th syntactic rule, represents the weight of the th syntactic rule, represents the key feature value of the current text context category, represents the th feature value of the context category covered by the syntactic rule, represents the total number of syntactic rules;
[0084] In different contexts, the fitness of rules is different. For example, a rule for the subject-verb-object structure has a much higher fitness in the Chinese context than in the English context. Therefore, it is necessary to perform calculations by combining weights and category feature differences. There are currently five rules, and their fitness, weights, and context feature values are shown in the following table:
[0085] Syntactic rule priority calculation parameter table:
[0086] ;
[0087] Substitute the values into the formula for calculation:
[0088] ;
[0089] ;
[0090] ;
[0091] Calculate to obtain the rule priority , which indicates that the overall fitness between the current text context category and the rule is relatively high, and it can be used as a high-priority application rule to obtain the application priority of syntactic rules.
[0092] The application ratio update sub-module calculates the application ratio of differential rules based on the application priority of syntactic rules, adjusts the application frequency of low-fitness rules, and at the same time increases the application proportion of high-fitness rules, filters out the rules that are crucial in priority, and obtains the updated set of syntactic rules;
[0093] Adopt a hierarchical update strategy to reduce the application proportion of rules with lower fitness and increase the proportion of rules with higher fitness. For example, if the priority score of a certain rule If the average priority score in the current context is 0.50, the application ratio of this rule can be appropriately increased, while the application ratio of rules lower than 0.50 needs to be decreased. Set the adjustment threshold to , that is, increase the application ratio of rules higher than 0.50 by 20% and decrease the application ratio of rules lower than 0.50 by 20%. Select the rule with the highest adaptability, calculate its adjusted application ratio, and finally obtain the updated set of syntactic rules.
[0094] Please refer to Figure 4 , the semantic structure matching module includes:
[0095] The text structure parsing sub-module calls the updated set of syntactic rules, analyzes the part-of-speech categories in the text according to the part-of-speech distribution, semantic dependency relationship, and syntactic boundary of the input text, analyzes the hierarchical relationship of the syntactic structure, and extracts the key information of the core verb, subject, and object to establish text structure data;
[0096] Obtain the word segmentation result of the input text and identify the part-of-speech category of each word. For example, for the sentence "The cat eats fish", the word segmentation result is [cat / noun, eat / verb, fish / noun]. Then, analyze the semantic dependency relationship of each word. For example, "cat" is the subject and forms a subject-predicate relationship with "eat", while "fish" is the object and forms an object-verb relationship with "eat". Next, determine the syntactic boundary and identify the compound sentence structure. For example, the sentence "The cat eats fish, and the dog drinks water" contains two parallel clauses, each forming a complete subject-predicate-object structure independently. Then, analyze the hierarchical relationship of the syntactic structure to determine the main clause and subordinate clause hierarchical relationship of the sentence. For example, in the sentence "Because the weather is hot, the cat likes to stay in the shade", the main clause is "The cat likes to stay in the shade", and "Because the weather is hot" is its adverbial clause. Subsequently, extract the key information of the core verb, subject, and object and form structured data storage. For example, for the sentence "The student studies math seriously", the extracted key information is [subject = student, verb = study, object = math]. Finally, organize the data and store it as text structure data.
[0097] The sentence pattern matching sub-module extracts syntactic structure information based on the text structure data, matches the corresponding sentence patterns, and uses the formula:
[0098] ;
[0099] Calculate the sentence pattern matching deviation , and obtain the sentence pattern matching error data, where represents the grammatical structure feature value of the input text, indicating the numerical quantization result of the -dimensional grammatical feature in the text, such as the numerical encoding of the subject-predicate-object relationship, the depth of the subordinate clause hierarchy, etc. represents the feature value of the standard sentence pattern, indicating the standard sentence pattern structure at the Reference values on dimensional features, such as the average dependency depth of a certain type of sentence pattern, the proportion of syntactic components, etc. Represents the weight of the syntactic feature item, measuring the importance of different syntactic features in the calculation of the matching degree. The weight size is set according to the influence degree of language features. For example, features with higher syntactic structure stability have larger weights. Represents the number of dimensions of the text syntactic features, that is, the total number of syntactic features involved in the calculation process, such as part-of-speech distribution, dependency structure, syntactic depth, etc. in multiple dimensions;
[0100] Obtain the syntactic structure data of the input text. For example, for the sentence "The student studies mathematics seriously", the text structure data is [Subject = student, Verb = study, Object = mathematics]. Then, select the most similar standard sentence pattern from the sentence pattern library. For example, the common subject-verb-object sentence pattern [Subject + Verb + Object]. Five syntactic features are involved in the calculation, and their values are as follows in the table:
[0101] Table of sentence pattern matching calculation parameters:
[0102] ;
[0103] Substitute the values into the formula for calculation:
[0104] ;
[0105] ;
[0106] ;
[0107] Calculate the sentence pattern matching error , which indicates that the deviation between the input text and the standard sentence pattern is small and can be used for further calculation of sentence pattern matching error data.
[0108] The structure matching degree calculation sub-module calculates the text structure matching degree based on the sentence pattern matching error data, sets the matching degree threshold range, analyzes the matching situation of different sentence patterns, and obtains the syntactic structure matching degree;
[0109] Set the range of matching degree thresholds. For example, sentences with a sentence pattern matching error less than 0.05 are considered to have a high matching degree, while sentences with an error greater than 0.15 have a low matching degree. The setting of this threshold is based on the error distribution of different text matching samples. For example, among 5000 sentences, more than 85% of the sentences with a matching error less than 0.05 are correctly marked by humans. Therefore, 0.05 is set as the high matching degree threshold. For sentences with an error exceeding 0.15, the error rate exceeds 70%. Therefore, 0.15 is set as the low matching degree threshold. Then, compare the matching error data and analyze the matching of different sentence pattern modes. For example, for the sentence "Students study mathematics seriously", the calculated value of its matching error is 0.034, which is lower than the set threshold of 0.05. Therefore, this text highly matches the standard sentence pattern mode. Next, adjust the matching weight coefficient so that different types of grammatical features maintain a reasonable proportion in the final calculation. For example, for the matching of long sentence structures, the weight of the sentence level complexity needs to be increased, while for short sentence matching, the depth of the dependency relationship is more critical. Finally, calculate the overall matching degree of the text. For example, when the full score of the standard matching degree is set to 1.0, and the matching error value is compared with the reference value of 0.15, the matching degree can be calculated by the following formula: , substitute the calculated value : , calculate the syntactic structure matching degree , indicating that the overall matching degree of this sentence is high.
[0110] Please refer to Figure 5 , the syntactic error detection module includes:
[0111] The syntactic structure matching sub-module obtains the parsing error range of the input text for the syntactic structure matching degree, calculates the syntactic matching degree of the input text through the syntactic rule library, and obtains the syntactic matching degree value;
[0112] The input text is decomposed into independent sentences, and each sentence is further decomposed into words or morphemes. Then, a syntactic analyzer is used to parse each sentence to generate a corresponding syntactic tree or dependency graph, which represents the grammatical relationships between the words in the sentence. For example, subject-predicate relationship, attributive-adverbial-complementary relationship, etc. Subsequently, the generated syntactic structure is compared with a predefined syntactic rule library, which contains common syntactic patterns and structures in the target language. For example, subject-verb-object (SVO) structure, attributive clause structure, etc. By comparison, the matching degree between the syntactic structure of the input text and the expected pattern can be determined. The matching degree can be quantified by calculating the difference between the input syntactic structure and the expected syntactic structure. For example, using the edit distance or other similarity measurement methods. Suppose the edit distance between the actual syntactic structure of the input sentence and the expected structure is 3, and the sentence length is 10, then the matching degree can be expressed as 1 - (3 / 10) = 0.7, that is, the matching degree is 70%. In this way, the parsing error range of the input text is obtained.
[0113] Based on the syntactic matching degree value, the parsing error calculation sub-module compares the error amount between the input text and the expected syntactic pattern, filters out the syntactic rules beyond the normal range, and uses the formula:
[0114] ;
[0115] Calculate the parsing error value HE, where represents the number of syntactic structures in the text, represents the th syntactic matching degree of the input text, represents the th expected syntactic matching degree, represents the th weight of the syntactic structure;
[0116] Compare the error amount between the input text and the expected syntactic pattern. First, set an expected syntactic matching degree threshold, for example, 0.8, indicating that sentences with a matching degree lower than 80% have syntactic errors. For each input sentence, calculate its syntactic matching degree value. If the value is lower than the expected threshold, it is considered that the sentence has a syntactic error. Then, filter out the sentences with a matching degree lower than the threshold. For each filtered sentence, calculate its parsing error value. If a sentence contains 5 syntactic structures, and its syntactic matching degrees of the input text are 0.6, 0.7, 0.5, 0.8, and 0.9 respectively, and the expected syntactic matching degrees are all 0.9, and the weights are all 1, then the parsing error value is calculated as follows:
[0117] ;
[0118] ;
[0119] ;
[0120] ;
[0121] The result shows that the parsing error value reflects the average deviation degree between the syntactic matching degree of the input text and the expected syntactic matching degree. The error value ranges between 0.1 and 0.3, belonging to the medium error range, indicating that the syntactic structure of this text partially deviates from the standard pattern, but still has a certain degree of comprehensibility. It is necessary to combine the error classification thresholds (such as minor error: , medium error: , severe error: ) to classify the levels of syntactic errors. The determination of the parsing error classification thresholds of 0.1 and 0.3 is based on the test of 500,000 sentence samples. Among them, for sentences with an error value less than 0.1, the probability that the grammar error affects reading comprehension is less than 5%. For sentences with an error value between 0.1 and 0.3, the probability that the error affects reading comprehension is between 5% and 30%. For sentences with an error value greater than 0.3, the probability that the error affects reading comprehension exceeds 50%. Therefore, 0.1 and 0.3 are set as the demarcation points for minor errors, medium errors, and severe errors, and this value can be further optimized as the scale of the test data increases.
[0122] The syntactic evaluation sub-module calls the parsing error value, classifies the error items, judges the type of syntactic error, analyzes the scope of error influence, and obtains the syntactic error evaluation value;
[0123] According to the size of the parsing error value, sentences are divided into different error levels. For example, those with an error value less than 0.1 are minor errors, those between 0.1 and 0.3 are medium errors, and those greater than 0.3 are severe errors. Then, judge the type of syntactic error and analyze the scope of error influence. Specifically, for each sentence, determine its main type of syntactic error, such as subject-verb disagreement, incorrect modifier position, clause structure error, etc. This can be achieved by analyzing the specific differences between the syntactic structure and the expected pattern. For example, if the subject and predicate of a sentence do not match in number or person, it can be determined as a subject-verb disagreement error. Finally, combining the error type and error level, obtain the syntactic error evaluation value. The syntactic error evaluation value can be a comprehensive score reflecting the degree of syntactic correctness of the sentence. For example, for a sentence, if its parsing error value is 0.25 and it is determined as a subject-verb disagreement error, a corresponding evaluation value, such as 70 points (out of 100), can be given, indicating that the sentence has a medium degree of syntactic error.
[0124] Please refer to Figure 6 , the syntactic optimization module includes:
[0125] The error evaluation sub-module calls the syntactic error evaluation value, compares it with the user's mislabeled information, and uses the formula:
[0126] ;
[0127] Calculate the syntactic error deviation of the user's annotation , screen the influencing factors, where represents the evaluated error value, represents the error value of the user's annotation, represents the total number of sentences;
[0128] There is a set of sentence datasets, which contain the error values automatically evaluated by the system and the error values manually annotated by the user . Calculate the absolute difference between the two for each sentence and accumulate and average them to form the error deviation . If the dataset contains five sentences, their error values are as follows:
[0129] ;
[0130] The error deviation is calculated as follows:
[0131] ;
[0132] After obtaining the error deviation, it is necessary to screen the influencing factors, such as the subjective judgment differences of different users for the same grammar error, or the recognition accuracy of the system for certain types of errors. The screening criterion can be set as the error deviation threshold. The determination of this threshold is based on the error fluctuation range of different users' annotations in the historical training dataset, usually set as 1.5 times the standard deviation of the average error of the training set. If the average error of the training dataset is 0.02 and the standard deviation is 0.015, the threshold is calculated as follows:
[0133] ;
[0134] When exceeds 0.0425, it is considered that there are significant differences in the user's annotations and further analysis is required. If it is lower than this value, it is considered that the system error is small and no adjustment is needed, and the syntactic error deviation is obtained.
[0135] Based on the syntactic error deviation, the rule adjustment sub-module combines the user's correction information, calculates the influence ratio of the user's feedback, optimizes the application direction of the syntactic rule library according to the influence ratio, and adjusts the application ratio of the syntactic rules to obtain the optimized result of the syntactic rules;
[0136] Combine the user's correction information to calculate the influence ratio of the user's feedback, that is, determine the degree of adjustment of the user's modified information to the applicable direction of the syntactic rule library. Specifically, by comparing the sentence structures before and after the user's modification, count which grammar rule application frequencies have changed, and calculate their influence ratio. For example, assume that a certain rule originally applied to 100 sentences, and after being adjusted according to the user's feedback, it becomes 120 sentences. Then the influence ratio is calculated as follows: , if it exceeds the preset threshold (such as 0.1), then it is necessary to adjust the applicable direction of this rule, optimize the rule library, and correspondingly adjust the application ratio of the rule. For example, if the original applicability of a certain rule is 50%, it needs to be increased to 60% after adjustment to obtain the syntactic rule optimization result.
[0137] The above is only the preferred embodiment of the present invention, and it does not limit the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A multilingual semantic parsing system, characterized in that, The system includes: The context monitoring module obtains the target language category, core semantic units, syntactic structure features, context-related words, and historical conversation data of the input text, calculates the similarity between the core semantic units and the historical conversation data, screens the fluctuating key texts, and obtains the context change classification result; The context monitoring module includes: The text feature extraction sub-module obtains the input text, detects the target language category of the input text, extracts the core semantic units, and identifies the syntactic structure features. At the same time, it collects context-related words and historical conversation data to establish a text feature data set; The context change analysis sub-module calculates the similarity between the core semantic units and the historical conversation data based on the text feature data set, analyzes the context change of the input text in the target language category, and uses the formula: ; Calculate the context offset of each semantic unit , identify the key points of context change, where represents the context weight value of the th core semantic unit of the input text, represents the context weight value of the th core semantic unit in the historical conversation data, represents the context criticality weight of the th core semantic unit, represents the total number of core semantic units; The fluctuating key screening sub-module analyzes the fluctuating key input texts based on the context deviation degree, compares the change trends of the different input texts, screens the fluctuating key input texts, and obtains the context change classification result; The syntactic rule adjustment module analyzes the applicable language range and parsing effect of the rules based on the context change classification result, calculates the adaptation degree of the current text, and adjusts the priority of the syntactic rules to obtain the syntactic rule update set; The semantic structure matching module calls the syntactic rule update set, analyzes the text structure according to the part-of-speech distribution, semantic dependency relationship, and syntactic boundary of the input text, matches the sentence pattern, and calculates the text structure matching degree to obtain the syntactic structure matching degree; The syntactic error detection module analyzes the parsing error range for the syntactic structure matching degree, compares the error amount between the parsing result and the expected syntactic pattern, and identifies the abnormal syntactic rules to obtain the syntactic error evaluation value; The syntactic optimization module calls the syntactic error evaluation value, calculates the feedback influence ratio based on the user's error annotation and correction information, and optimizes the applicable direction of the syntactic rule library to obtain the syntactic rule tuning result.
2. The multilingual semantic parsing system according to claim 1, wherein The context change classification result includes language style, semantic variation, and interaction mode. The syntactic rule update set specifically refers to rule adaptability, applicable range, and priority setting. The syntactic structure matching degree specifically refers to matching accuracy, pattern consistency, and structural integrity. The syntactic error evaluation value includes error degree, error type, and influence range. The syntactic rule tuning result includes rule adjustment effect, application frequency change, and user feedback result.
3. The multilingual semantic parsing system according to claim 1, wherein The syntactic rule adjustment module includes: The rule applicability analysis sub-module analyzes the applicable language range of each rule based on the context change classification result according to the rules in the syntactic rule library, and analyzes the rule parsing effect to obtain the rule adaptability data; The rule priority adjustment sub-module calculates the adaptation degree between the current text context category and each rule based on the rule adaptability data, and uses the formula: ; Obtain the application priority of syntactic rules , where represents the fitness of the th syntactic rule, represents the weight of the th syntactic rule, represents the key feature value of the current text context category, represents the feature value of the context category covered by the th syntactic rule, represents the total number of syntactic rules; The application ratio update sub-module calculates the application ratio of the different rules based on the syntactic rule application priority, adjusts the application frequency of the low-adaptation rules, and at the same time increases the application proportion of the high-adaptation rules, and screens the key rules with priority to obtain the syntactic rule update set.
4. The multilingual semantic parsing system according to claim 1, wherein The semantic structure matching module includes: The text structure analysis sub-module calls the updated set of syntactic rules, analyzes the part-of-speech categories in the text according to the part-of-speech distribution, semantic dependency relationship and syntactic boundary of the input text, analyzes the hierarchical relationship of the syntactic structure, and extracts key information such as the core verb, subject and object to establish text structure data; The sentence pattern matching sub-module extracts syntactic structure information based on the text structure data, matches the corresponding sentence pattern, using the formula: ; Calculate the deviation of sentence pattern matching , to obtain sentence pattern matching error data, where represents the syntactic structure feature value of the input text represents the feature value of the standard sentence pattern model represents the weight of the syntactic feature item represents the number of dimensions of the text syntactic feature The structure matching degree calculation sub-module calculates the text structure matching degree based on the sentence pattern matching error data, sets the matching degree threshold range, analyzes the matching situation of different sentence patterns, and obtains the syntactic structure matching degree.
5. The multilingual semantic parsing system according to claim 1, characterized in that, The syntactic error detection module includes: The syntactic structure matching sub-module obtains the parsing error range of the input text for the syntactic structure matching degree, calculates the syntactic matching degree of the input text through the syntactic rule library, and obtains the syntactic matching degree value; Based on the syntactic matching degree value, the parsing error calculation sub-module compares the error amount between the input text and the expected syntactic pattern, filters out syntactic rules that exceed the normal range, and uses the formula: ; Calculate the parsing error value HE, where represents the number of syntactic structures in the text, represents the th syntactic matching degree of the input text, represents the th expected syntactic matching degree, represents the th weight of the syntactic structure; The syntactic evaluation sub-module calls the parsing error value, classifies the error items, judges the type of syntactic error, analyzes the influence range of the error, and obtains the syntactic error evaluation value.
6. The multilingual semantic parsing system according to claim 1, wherein The syntactic optimization module includes: The error evaluation sub-module calls the syntactic error evaluation value, compares the user's error annotation information, using the formula: ; Calculate the syntactic error deviation amount marked by the user , screen the influencing factors, where represents the evaluated error value represents the error value marked by the user represents the total number of sentences The rule adjustment sub-module calculates the influence ratio of the user feedback based on the syntactic error deviation amount and the user's correction information, optimizes the application direction of the syntactic rule library according to the influence ratio, and adjusts the application ratio of the syntactic rules to obtain the syntactic rule tuning result.
Citation Information
Patent Citations
Intelligent manuscript reviewing system and method based on AI
CN118485060A
Understanding natural language using tumbling-frequency phrase chain parsing
US20200125641A1