Grammar analysis method and device and electronic equipment
By adjusting the reading order of the terminator and modifying the automaton state, the structure of the terminator stream is realized, the problem of low syntax parsing efficiency in the existing technology is solved, and the compilation and development efficiency and syntax extension compatibility are improved.
Patent Information
- Application Number
- CN202510628445.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-12
AI Technical Summary
When reading the final word, existing syntax parsers must read it one by one in a certain order, and cannot fully utilize the implemented syntax rules, resulting in poor syntax extension and compatibility and low compilation and development efficiency.
By adjusting the order of reading of the terminator, allowing branch, loop and jump to read, modifying the automaton state and inserting preset terminators to perform specific actions, realizing the structure of the terminator stream.
Improves compilation and development efficiency and cost-effectiveness, supports richer syntax features and clearer syntax tree construction, and reduces computational overhead.
Smart Images

Figure CN120469694A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer and software development technology, and in particular to a grammar parsing method, device and electronic device. Background Art
[0002] In compiler theory, a compiler is conceptually a streaming text processor. The hardware-independent portion of a compiler primarily consists of two components: a lexical analyzer and a parser. In the parser, a finite state machine (FSM) is equivalent to a grammar; determining the FSM uniquely defines the language. Therefore, the FSM is the core component of the parser. The FSMs in the lexical analyzer and parser enable compilation of data streams (such as source programs).
[0003] Specifically, if Figure 1 As shown in the figure, during the compilation process, first, when the data stream passes through the lexical analyzer, specific character sequences in the data stream are converted into terminal words according to the determined lexical rules, thereby converting the character stream into a terminal word stream. After the lexical analyzer lexically analyzes the data stream, each word obtained is a terminal word.
[0004] The terminal word stream is then moved into the FSM in the parser. The FSM works as follows: each time a terminal word is moved in, the target state to transition to is determined based on the current state and the terminal word read in (and subsequent terminal words to be read), and an action is performed during the state transition. Thus, after the terminal word stream is moved into the FSM, the FSM constructs a syntax tree based on the terminal word stream. During this process, the syntax tree is output when a terminal symbol is read (or the current state is determined to be a terminal state) and no error is reported.
[0005] However, existing parsers must read terminal words one by one in a specific order. This means that the terminal word stream lacks structure (such as branches, loops, and jumps), making it impossible to fully utilize the implemented grammar rules. This hinders the reuse of implemented grammars in grammar extension and compatibility tasks. Summary of the Invention
[0006] This application provides a grammar parsing method, device, and electronic device to solve the problem that the terminal word stream has no structure (such as branches, loops, jumps, etc.), which makes it impossible to fully utilize the implemented grammar rules, resulting in the problem of being unfavorable for the reuse of the implemented grammar in grammar expansion, grammar compatibility, etc. The specific implementation scheme is as follows:
[0007] In a first aspect, the present application provides a grammar parsing method, the method comprising:
[0008] Determine multiple terminal words corresponding to the input text;
[0009] Determine a pending processing method for the terminal word, and perform corresponding processing on the terminal word according to the pending processing method;
[0010] Modifying a state of an automaton based on the processing of the terminal words, and adjusting a reading order of the plurality of terminal words based on the modification of the state of the automaton to obtain a new reading order;
[0011] According to the new reading order, the corresponding processing is continued for the next terminal word until the end condition is met to obtain the target result.
[0012] Through the above-mentioned application embodiments, the modification of the automaton state and the adjustment of the reading order of multiple terminal words are realized, so that according to the adjusted new reading order, the effect of selectively, cyclically or jumpily reading the terminal words can be achieved, that is, the structuring of the terminal word flow is realized, so that the grammatical parsing based on the terminal words is no longer a sequential process, but a structural process, that is, branching, looping, and jumping are allowed during grammatical parsing, and the original results can be reused in grammatical expansion and grammatical compatibility, so that the development cost of compilation can be reduced and the development efficiency of compilation can be improved.
[0013] In a possible implementation, determining the multiple terminal words corresponding to the input text includes:
[0014] Obtaining the input text to be processed;
[0015] Performing lexical analysis on the input text to obtain corresponding multiple terminal words, and determining the original reading order of the multiple terminal words.
[0016] Through the above-described application embodiment, based on the lexical analysis of the acquired input text, the input text is converted into multiple terminal words, so that the compiler can better understand the meaning of the input text, thereby facilitating further analysis of the terminal words to obtain the final target result. At the same time, the original reading order of the multiple terminal words is also determined, so that when reading the terminal words later, especially when reading the first terminal word, the terminal word to be read can be determined.
[0017] In a possible implementation, determining a pending processing method for a terminal word and performing corresponding processing on the terminal word according to the pending processing method includes:
[0018] Before reading the terminal word, determining whether the processing mode for the terminal word is reduction processing or shift processing according to the terminal word and the current state of the automaton;
[0019] If it is determined that the processing method to be processed is the reduction processing, then according to the target grammar production formula, multiple terminal words and non-terminal words before the corresponding terminal word are reduced to obtain the corresponding non-terminal words, and when reducing, the grammar production formula is modified or added.
[0020] Through the above-mentioned application embodiment, before reading the terminal word, it is determined whether to perform reduction processing or shift processing for the terminal word according to the current state of the terminal word and the automaton, so that the determination of the pending processing mode corresponding to the terminal word is more efficient and more accurate. In addition, after determining to perform reduction processing for the terminal word, according to the target grammar production formula, the multiple terminal words and non-terminal words before the terminal word are reduced to obtain the corresponding non-terminal words, thereby converting the more complex grammatical structure (i.e., the structure composed of multiple terminal words and non-terminal words) into simpler non-terminal words, which helps to reduce the complexity of grammatical analysis, enables the compiler to process the input text more efficiently, and enables the compiler to identify and process these complex structures through reduction, thereby enabling the compiler to support richer grammatical features. At the same time, the reduction based on the target grammar production formula reduces the number of symbols that the compiler needs to process, thereby reducing the computational overhead in the parsing process, and also enables the compiler to more easily build a clear-layered grammatical tree, which helps subsequent semantic analysis and code generation. In addition, when reducing, modifying or adding grammar production formulas can adapt to specific analysis needs or solve certain problems.
[0021] In a possible implementation, after determining that the to-be-processed manner for the terminal word is reduction processing or shift processing, the method further includes:
[0022] If it is determined that the processing method to be processed is the shift processing, the corresponding terminal word is shifted into the automaton;
[0023] According to the corresponding terminal word and the current state of the automaton, a target state of the automaton to be transferred is determined, and the state of the automaton is transferred from the current state to the target state.
[0024] Through the above-mentioned application embodiment, after determining to move the terminal word into the process, the terminal word is moved into the automaton, and the target state of the automaton to be transferred is determined based on the terminal word and the current state of the automaton, so that the determined target state is more accurate. Then, based on the target state, the state of the automaton is transferred, thereby realizing the state transfer of the automaton, allowing the automaton to distinguish different terminal words, avoiding errors and ambiguities in the analysis process, and further improving the accuracy of the entire compilation process.
[0025] In a possible implementation, adjusting the reading order of the multiple terminal words includes:
[0026] When relocating the lexer, a preset terminal word is inserted so that the lexer performs a specific action when reading the preset terminal word; wherein the specific action includes any one of a loop jump, a branch jump, and a modified terminal word, and the action parameters of the specific action are specified by the specification action or the shift action that creates the preset terminal word.
[0027] Through the above-mentioned application embodiment, when relocating the lexer, a preset terminal word is inserted, so that when the lexer reads the preset terminal word, any specific action including loop jump, branch jump, and terminal word modification is performed, so that the actual reading order of multiple terminal words is changed, thereby realizing the adjustment of the reading order of multiple terminal words, so that the terminal word stream has a structure.
[0028] In a possible implementation, the end condition is that an end word is read and the current state of the automaton is an end state, wherein the end word is the next end word of the last end word in the original reading order.
[0029] Through the above application embodiment, the end condition based on the end word and the end state is used to accurately determine whether to end the parsing process to complete the parsing of the input text.
[0030] In a second aspect, the present application further provides a grammar parsing device, the device comprising:
[0031] A determination module, configured to determine a plurality of terminal words corresponding to an input text;
[0032] A first processing module is used to determine a pending processing method for the terminal word and perform corresponding processing on the terminal word according to the pending processing method;
[0033] a modification module, configured to modify a state of the automaton based on processing of the terminal words, and adjust a reading order of the plurality of terminal words based on the modification of the state of the automaton to obtain a new reading order;
[0034] The second processing module is used to continue to perform corresponding processing on the next terminal word according to the new reading order until the end condition is met to obtain the target result.
[0035] In a possible implementation, the determination module is specifically configured to obtain the input text to be processed; perform lexical analysis on the input text to obtain a corresponding plurality of terminal words; and determine the original reading order of the plurality of terminal words.
[0036] In one possible implementation, the first processing module is specifically configured to determine, before reading a terminal word, whether the processing method for the terminal word is reduction processing or shift processing based on the terminal word and the current state of the automaton; if it is determined that the processing method is reduction processing, then, based on a target grammar production, multiple terminal words and non-terminal words preceding the corresponding terminal word are reduced to obtain corresponding non-terminal words, and during the reduction, the grammar production is modified or newly added.
[0037] In one possible embodiment, the first processing module is further used to move the corresponding terminal word into the automaton if it is determined that the processing method to be processed is the move-in processing; determine the target state of the automaton to be transferred based on the corresponding terminal word and the current state of the automaton, and change the state of the automaton from the current state to the target state.
[0038] In one possible implementation, the modification module is specifically used to insert a preset terminal word when relocating the lexer, so that the lexer performs a specific action when reading the preset terminal word; wherein the specific action includes any one of a loop jump, a branch jump, and a modified terminal word, and the action parameters of the specific action are specified by the specification action or the shift action that creates the preset terminal word.
[0039] In a possible implementation, the end condition is that an end word is read and the current state of the automaton is an end state, wherein the end word is the next end word of the last end word in the original reading order.
[0040] In a third aspect, the present application provides an electronic device, comprising:
[0041] Memory for storing computer programs;
[0042] The processor is configured to implement the above-mentioned grammar parsing method steps when executing the computer program stored in the memory.
[0043] In a fourth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned grammar parsing method are implemented.
[0044] For each of the above-mentioned aspects from the second to the fourth aspects and the technical effects that may be achieved by each of the aspects, please refer to the above-mentioned description of the technical effects that can be achieved by the first aspect or various possible solutions in the first aspect, and no further details will be given here. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 A logical diagram of an existing grammatical parsing method provided in an embodiment of the present application;
[0046] Figure 2 A schematic diagram of an application scenario provided in an embodiment of the present application;
[0047] Figure 3 A flowchart of a grammar parsing method provided in an embodiment of the present application;
[0048] Figure 4 A logical diagram of a grammar parsing method provided in an embodiment of the present application;
[0049] Figure 5a A schematic diagram of a GROUPING SETS syntax extension provided in an embodiment of the present application Figure 1 ;
[0050] Figure 5b A schematic diagram of a GROUPING SETS syntax extension provided in an embodiment of the present application Figure 2 ;
[0051] Figure 6a A WITH FUNCTION syntax extension provided in this application embodiment Figure 1 ;
[0052] Figure 6b A WITH FUNCTION syntax extension provided in this application embodiment Figure 2 ;
[0053] Figure 6c A WITH FUNCTION syntax extension provided in this application embodiment Figure 3 ;
[0054] Figure 7 A schematic diagram of a grammatical parsing device provided in an embodiment of the present application;
[0055] Figure 8 A schematic diagram of an electronic device provided in this application. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail with reference to the accompanying drawings. The specific operating methods in the method embodiments can also be applied to device embodiments or system embodiments. It should be noted that in the description of the present application, "multiple" is understood as "at least two". "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the following three situations: A exists alone, A and B exist at the same time, and B exists alone. A is connected to B, which can represent the following two situations: A is directly connected to B and A is connected to B through C. In addition, in the description of the present application, words such as "first" and "second" are only used to distinguish the purpose of description, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order.
[0057] The embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0058] In compilation theory, the complete process of compiling a program generally includes lexical analysis, syntax analysis, semantic analysis, intermediate code generation and other processes.
[0059] The lexical analysis described above is used to convert a specific character sequence into a terminal word according to a certain lexical rule. Therefore, the lexical analysis can also be considered as a process of converting a character stream into a terminal word stream.
[0060] The above grammar analysis is used to convert the terminal word stream into a syntax tree containing non-terminal words (non-leaf nodes) and terminal words (leaf nodes) through specification (or deduction). In this process, the terminal words are read sequentially until they are accepted or an error is reported.
[0061] The above-mentioned semantic analysis is used to translate the results of grammatical analysis into actual actions. In actual operation, semantic analysis is usually completely synchronized with grammatical analysis and both belong to the content of the grammatical parser. The implementation method of the semantic analysis can be: bind an additional action when the specification is made; or it can be: bind an additional action when shifting. For example, in the process of translating a+b, expression->identifier and expression->expression plus expression will be used for the specification. The action of the former grammar production is to find the address corresponding to the identifier and read its content. The action of the latter grammar production is to add the two expressions to generate a new expression. The specific content of the action is to generate intermediate code, but in actual application, different actions can be used as needed. When the grammatical parsing method proposed in the embodiment of the present application is applied in the structured query language (English: Structured Query Language, abbreviated as SQL), intermediate code can be not generated, so that the development efficiency of compilation can be improved.
[0062] Throughout the parsing process, the order in which the terminal words are read remains unchanged, and each terminal word is read only once. This means that the parser is a sequential process that can only read terminal words sequentially by shifting, without looping or jumping through the reading process. The parser can only make decisions about shifting and reducing (or deducing), and select the grammar productions used for the reduction (or deduction). Furthermore, throughout the parsing process, the structure of the FSM within the parser remains unchanged (only the current state changes, not the state transition function).
[0063] Although the above method is still Turing complete (that is, all functions can be implemented by the program), the terminal word flow without structure (such as branches, loops, jumps, etc.) cannot fully utilize the implemented grammatical rules, which is not conducive to the reuse of the implemented grammar in grammar expansion, grammar compatibility, etc., and ultimately leads to low compilation development efficiency.
[0064] To solve the above problem, an embodiment of the present application provides a grammatical parsing method to determine multiple terminal words corresponding to the input text, and then before reading the current terminal word, first determine the pending processing mode of the current terminal word, and based on the pending processing mode, perform corresponding processing on the current terminal word, and then based on the processing of the current terminal word, modify the state of the automaton, and based on the modification of the state of the automaton, adjust the reading order of multiple terminal words to obtain a new reading order, and then continue to perform corresponding processing on the next terminal word according to the new reading order until the end condition is reached and the target result is obtained, so that multiple terminal words can no longer be read in sequence according to the original reading order, but the terminal words can be read according to the adjusted new reading order, so that the reading of the terminal words can have a selective, cyclical, and jumpy effect, that is, the structuring of the terminal words is realized. In other words, the input text is no longer viewed as a sequential stream of terminal words, but the terminal words themselves are allowed to have structure, so that grammatical parsing is no longer a sequential process, but a structured process (that is, branching, looping, and jumping are allowed during grammatical parsing). This is conducive to reusing the original results in grammar expansion and grammar compatibility (such as reusing the original grammar), reducing the development cost and time cost of compilation, and significantly improving the development efficiency of compilation.
[0065] The preferred embodiments of the present application will be described and explained below with reference to the accompanying drawings.
[0066] See Figure 2 FIG2 is a schematic diagram of an application scenario of an embodiment of the present application, which includes a terminal device 210 and a compiler 220 .
[0067] In the embodiment of the present application, the terminal device 210 includes but is not limited to mobile phones, tablet computers, laptop computers, desktop computers, e-book readers, intelligent voice interaction devices, smart home appliances, car terminals and other devices; a client related to compilation can be installed on the terminal device, which can be software (such as a browser), or a web page, applet, etc. The compiler 220 can be a background server corresponding to the software or web page, applet, etc., or an electronic device (such as a server) specifically used for compilation, which is not specifically limited in this application. The compiler 220 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (English: Content Delivery Network, abbreviated as CDN), and big data and artificial intelligence platforms.
[0068] It should be noted that the grammar parsing method in the embodiment of the present application can be executed by an electronic device, which can be a compiler 220 or a terminal device 210. That is, the method can be executed by the compiler 220 or the terminal device 210 alone, or can be executed jointly by the compiler 220 and the terminal device 210.
[0069] In an optional implementation, the terminal device 210 and the compiler 220 may communicate with each other via a communication network.
[0070] In an optional implementation, the communication network is a wired network or a wireless network.
[0071] It should be noted that Figure 2 The figures are only examples. In fact, the number of terminal devices and compilers is not limited and is not specifically limited in the embodiments of this application.
[0072] The following describes the grammar parsing method provided by the exemplary embodiment of the present application in conjunction with the application scenarios described above and with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principles of the present application, and the implementation of the present application is not limited in this respect. In addition, the embodiments of the present application can be applied to various scenarios, including not only computer language compilation scenarios, but also various scenarios such as cloud technology, artificial intelligence, smart transportation, and assisted driving.
[0073] The following describes a grammar parsing method provided by an exemplary embodiment of the present application in combination with the application scenarios described above and with reference to the accompanying drawings.
[0074] See Figure 3As shown, the embodiment of the present application provides a grammar parsing method, including:
[0075] S301: Determine multiple terminal words corresponding to the input text.
[0076] In the embodiment of the present application, the input text may be an input language, such as a source program. The source program is a text using a specific encoding set. In the embodiment of the present application, the input language may be used as a data stream.
[0077] Specifically, first, an input text is obtained. Then, a lexical analysis is performed on the input text, thereby dividing the input text into multiple terminal words, and then determining multiple terminal words corresponding to the input text. The multiple terminal words can be a sequence (i.e., a terminal word sequence), that is, the multiple terminal words have an order, and the order can be an original reading order, so that the original reading order of the multiple terminal words can be determined, so that when reading multiple terminal words later, it can be determined in which order to read the terminal words first.
[0078] The above-mentioned lexical analysis can be performed by converting a specific character sequence into a terminal word according to a determined lexical rule. For example, the lexical rule can include a preset dictionary library, so that the input text can be segmented based on the preset dictionary library to obtain multiple terminal words. The terminal word can be a specific character sequence (i.e., a character stream) representing a specific specific meaning. Generally, each terminal word has a clear meaning.
[0079] Exemplarily, the terminal word can be any reserved word in the C language, an operator, a sequence starting with a letter or an underscore and containing letters, underscores or numbers (i.e., an identifier), a sequence starting with a number and containing only numbers (i.e., an integer); a sequence in the form of "integer.integer" (i.e., a floating-point number); a pair of single quotes ' and the content between them (i.e., a character); or a pair of double quotes " and the content between them (i.e., a string).
[0080] Each word obtained after lexical analysis is a terminal word. In other words, the deduction of the grammar ends here. The word is a concrete object of the language, not an abstract concept. Therefore, through lexical analysis, a data stream (such as a source program) can be converted into a stream of terminal words. In other words, the input text can be converted into a stream of terminal words.
[0081] A grammar can be defined as: mathematically, the set of all sequences of terminal words that can be written in a language (whether a programming language or a natural language). This grammar must contain a special grammar production, ε->E, whose left side represents an empty string (nothing), a terminal word (not restricted by grammar category). This production is used to generate a language out of thin air. All statements (i.e., sequences of terminal words) that satisfy this production are considered acceptable, or legal.
[0082] For example, for the C language, there can be the following grammar production: {(ε->program), (program->main function), (main function->main function function), (main function->main function name parameter function body), (function body->function body statement), (function body->statement), (statement->expression semicolon), (statement->ε), (expression->expression plus identifier), (expression->identifier)}. For this grammar, int main(){a+b;} is a syntactically correct program, but semantically illegal.
[0083] The above grammar can be described using grammar productions. This grammar production can be a mapping, such as A->B, which means that A can generate B, meaning that A, which has a relatively abstract or simple meaning, can generate B, which has a more specific or complex meaning. However, it should be noted that while using grammar productions to describe statements is convenient, it is difficult to determine whether a statement is legal based on the grammar productions. Mathematically, it can be proven that any grammar is equivalent to a certain finite automaton (DFA) (i.e., the set of legal statements is the same), and the existing technology has detailed algorithms for constructing such automatons, thus transforming the problem of determining whether a statement is legal into the problem of state transitions in the automaton.
[0084] Furthermore, grammars can be divided into three categories based on the constraints on the grammar productions: context-sensitive grammars, context-free grammars (i.e., A is restricted to a single non-terminal word), and regular grammars (i.e., B is restricted to one terminal word and at most one non-terminal word). These three types of grammars are equivalent, meaning they represent the same scope. In practical grammar applications, compilers typically use context-free grammars.
[0085] Non-terminal words represent abstract meanings. Due to their abstract nature, they cannot be defined as specific character sequences; their sequences can only be constrained to a certain degree by grammar productions. The process of gradual abstraction progresses from the previously defined but meaningless characters to the previously defined terminal words with specific sequences and definite meanings, and then to the non-terminal words without specific sequences but capable of expressing abstract meanings.
[0086] Furthermore, grammar analysis can include two methods: deduction and reduction. Deduction involves gradually converting abstract concepts into concrete statements. Reduction involves gradually abstracting concrete statements into concepts. For example, based on grammar productions, the operation of abstracting several currently read symbols (terminal or non-terminal words) into a more abstract non-terminal word is called reduction.
[0087] In the embodiments of the present application, the grammatical analysis involved is not limited to specification, but may also include deduction.
[0088] S302: Determine a pending processing method for the terminal word, and perform corresponding processing on the terminal word according to the pending processing method.
[0089] Specifically, before reading a terminal word, the corresponding processing mode for the terminal word, shift processing or reduction processing, is determined based on the terminal word and the current state of the automaton. That is, before reading a terminal word, a determination is made as to whether the terminal word should be shifted or reduced (i.e., using terminal word pre-reading technology) based on the terminal word and the current state of the automaton. The corresponding processing is then performed, thereby resolving shift / reduce conflicts.
[0090] For example, the action table is searched based on the current state of the automaton and the terminal word. If the action found is a shift, the corresponding pending processing method of the terminal word is determined to be a shift process; if the action found is a reduction, the corresponding pending processing method of the terminal word is determined to be a reduction process.
[0091] However, it should be noted that when using an LALR(1) (Look-Ahead-LR(1)) grammar compiler, if the terminal word to be read is changed, the pre-read terminal word must be cleared. When using a grammar compiler that pre-reads a single terminal word (such as LL(1) (Left-to-right, Top-down parsing with 1 symbol lookahead) or LR(1) (Left-to-right, Rightmost derivation in reverse, with 1 symbol lookahead)), if the terminal word to be read is changed, the pre-read terminal word must also be cleared. In addition, if the compiler theoretically pre-reads n terminal words, all of them should be cleared. LR(0) (Left-to-Right, Rightmost derivation, 0 lookahead) that does not pre-read terminal words does not do this.
[0092] In an embodiment of the present application, before reading the terminal word, the processing method corresponding to the terminal word is determined to be shift processing or reduction processing based on the terminal word and the current state of the automaton. It can be that before reading a terminal word each time, the processing method corresponding to the terminal word is determined to be shift processing or reduction processing based on the terminal word and the current state of the automaton.
[0093] Furthermore, before reading the initial terminal word, the processing method determined for the initial terminal word is typically a shift process. The initial terminal word is the first terminal word in the original reading order of multiple terminal words corresponding to the input text. In other words, when processing the initial terminal word, it is typically subjected to a shift process.
[0094] The above automaton can be a FSM.
[0095] Then, after determining the pending processing mode corresponding to the current terminal word, the current terminal word is processed accordingly according to the pending processing mode.
[0096] In one possible implementation, when it is determined that the processing method corresponding to a terminal word is reduction processing based on the current state of the terminal word and the automaton, the multiple terminal words and non-terminal words preceding the terminal word are reduced according to the target grammar production formula to obtain the corresponding non-terminal words, thereby converting a more complex grammatical structure (i.e., a structure composed of multiple terminal words and non-terminal words) into a simpler non-terminal word, which helps to reduce the complexity of grammatical analysis, enables the compiler to process input text more efficiently, and enables the compiler to identify and process these complex structures through reduction, thereby enabling the compiler to support richer grammatical features. In addition, the reduction based on the target grammar production formula reduces the number of symbols that the compiler needs to process, thereby reducing the computational overhead during the parsing process, and also makes it easier for the compiler to construct a clearly hierarchical syntax tree, which is helpful for subsequent analysis and code generation.
[0097] Furthermore, during the reduction, additional operations may be performed to more smoothly reduce a series of words (such as multiple terminal words and non-terminal words) to corresponding non-terminal words.
[0098] In the embodiment of the present application, the above additional operation can be to modify or add a grammar production formula. That is, during the specification, the grammar production formula can also be modified or added to meet specific analysis requirements or solve certain problems.
[0099] The modified or newly added grammar productions can be target grammar productions or non-target grammar productions, so that when the specification is further reduced, a new target grammar production can be selected from the modified or newly added grammar productions or the original grammar productions. In addition, the modified or newly added grammar productions can be multiple grammar productions or a single grammar production. In other words, when performing the specification, one or more grammar productions can be modified or newly added to facilitate better parsing of the input text.
[0100] In another possible implementation, when, based on the terminal word and the current state of the automaton, it is determined that the processing method corresponding to the terminal word is to be moved into, the terminal word is moved into the automaton. Then, based on the terminal word and the current state of the automaton, the target state to which the automaton is to be transferred is determined, and the state of the automaton is transferred from the current state to the target state, thereby achieving state transfer of the automaton. This allows the automaton to distinguish different terminal words, avoids errors and ambiguities during analysis, and further improves the accuracy of the entire parsing process.
[0101] S303: Based on the processing of the terminal words, the state of the automaton is modified, and based on the modification of the state of the automaton, the reading order of the plurality of terminal words is adjusted to obtain a new reading order.
[0102] Specifically, when the processing method is reduction processing, the state of the automaton can also be modified during the reduction, and based on the modification of the state of the automaton, the reading order of multiple terminal words can be adjusted to obtain a new reading order, thereby achieving the effect of selectively, cyclically or jumpily reading terminal words, that is, realizing the structuring of the terminal word stream.
[0103] The multiple terminal words may include all or part of the multiple terminal words corresponding to the input text. The multiple terminal words may also include repeated terminal words. The multiple terminal words may be set according to the specific application scenario.
[0104] The adjusted reading order may be the current reading order of the terminal words, which may be the original reading order or a non-original reading order.
[0105] In the embodiment of the present application, two components, an automaton editor and a lexical locator, can be added to the compiler (the components are as follows: Figure 4As shown), through the automaton editor and the lexical locator, the structure of the automaton can be modified and the current reading position of the lexer (the lexer can be a lexical analyzer) can be corrected in the action of the automaton state transfer, and then based on the correction of the current reading position of the lexer, the terminal words in the automaton flowing into the syntax parser can no longer flow in according to the original reading order, but flow in according to the current reading position corrected by the lexer (that is, the adjusted new reading order), so that the reading order of the terminal words is also changed, thereby achieving the effect of selectively, cyclically or jumpily reading the terminal words, that is, realizing the structuring of the terminal word flow, that is, realizing the adjustment of the reading order of multiple terminal words.
[0106] Specifically, in order to facilitate the structuring of terminal words (i.e., to facilitate the adjustment of the reading order of multiple terminal words), when the current reading position of the lexer is relocated by the lexer locator (i.e., when the lexer is relocated), some additional terminal words can be inserted, and the inserted additional terminal words are recorded as preset terminal words, that is, the preset terminal words can be inserted when the lexer is relocated. Then, when the lexer reads the preset terminal words, a type of specific action can be performed, so that through the execution of the specific action, the actual reading order of the multiple terminal words is changed, thereby achieving the adjustment of the reading order of the multiple terminal words and making the terminal words have structure.
[0107] In the embodiment of the present application, the above-mentioned preset terminal words will not change the original order (i.e., the original reading order) of multiple terminal words (i.e., the terminal word sequence), but will in fact change the reading order of multiple terminal words, which is in fact equivalent to deleting, adding, copying, and rearranging terminal words. Thus, by presetting terminal words, the reading order of multiple terminal words can be adjusted, and a new reading order of multiple terminal words can be obtained, so that the terminal words have a structure. Moreover, the above-mentioned preset terminal words are both control symbols for modifying the lexer and terminal words that actually participate in the state transfer of the automaton, and also have the effect of resolving specification / specification conflicts. In addition, the above-mentioned preset terminal words do not match any input text.
[0108] The specific action can be a jump action such as a loop or branch, or an action such as modifying a terminal word. The specific action of the specific action can be determined by the action parameters of the specific action. The action parameters can be specified by the specification action or the shift action that creates the preset terminal word.
[0109] That is to say, when performing the specification, it is also possible to set what action the preset terminal word should perform according to needs (such as based on the automaton state determination) and specify the parameters of the action, that is, specify the parameters of the specific action.
[0110] Furthermore, one or more predefined terminators can be created within a protocol. When a predefined terminator is subsequently entered, a corresponding protocol action can be performed in the next step. However, the predefined terminator can also be used only to convey information without corresponding to a protocol action.
[0111] In addition, in the aforementioned step S302 , when specifying, the modification or addition of the grammar production formula can be achieved by inserting the above-mentioned preset terminal words, so that the modification or addition of the grammar production formula can better meet the needs.
[0112] Optionally, when the processing method is shift processing, in the shift processing, after the state of the automaton is changed from the current state to the target state, the reading order of multiple terminal words can also be adjusted based on the modification of the state of the automaton to obtain a new reading order.
[0113] In the embodiment of the present application, the specific adjustment method for adjusting the reading order of multiple terminal words performed during the shift processing is consistent with the specific adjustment method for adjusting the reading order of multiple terminal words performed during the specification processing, and will not be repeated here.
[0114] In an embodiment of the present application, the reading order of the above-mentioned multiple terminal words can be adjusted by adjusting the reading order once each time a corresponding reduction processing and / or shift processing is performed on a terminal word; or the reading order can be adjusted when a specific terminal word (such as a preset terminal word) is subjected to corresponding reduction processing and / or shift processing, thereby achieving flexible adjustment of the reading order to better meet user needs and make the target result obtained more accurate.
[0115] S304: Continue to perform corresponding processing on the next terminal word according to the new reading order until the end condition is met and the target result is obtained.
[0116] After obtaining the new reading order, continue to perform corresponding processing on the next terminal word according to the new reading order (such as corresponding shift processing or corresponding reduction processing), and then continue to perform corresponding processing based on the processing performed on the terminal word (such as adjusting the state of the automaton, adjusting the reading order of multiple terminal words), until the end condition is reached and the target result of the input text is obtained.
[0117] It's important to note that when performing the corresponding reduction processing for the next terminal word, the reduction is also based on the target grammar production. The target grammar production can be a grammar production selected from the modified or newly added grammar productions, or the original grammar productions. When performing the reduction based on the target grammar production, the target grammar production can be selected based on the specific application scenario.
[0118] The target result may be the purpose that the source program (ie, input text) needs to achieve, such as outputting a text, outputting a program, outputting an instruction, outputting an action, etc.
[0119] The termination condition can be that a termination word is read and the current state of the automaton is the termination state. The termination word is the next termination word after the last termination word in the original reading order. The termination state is an acceptable state. Therefore, this termination condition can more accurately determine whether to terminate the aforementioned processing and output the target result.
[0120] It should be noted that when the end state does not match the end word, the state of the automaton can still be adjusted through shift processing or reduction processing. In addition, the above end word does not match any input text.
[0121] In addition, in an embodiment of the present application, after continuing to perform corresponding processing on the next terminal word according to the new reading order, it is possible to jump back to the original reading order according to the first requirement, and then obtain the target result when the end condition is reached; it is also possible to no longer jump back to the original reading order according to the second requirement, and then obtain the target result when the end condition is reached.
[0122] In summary, compared with the prior art method for parsing the grammar in which the structure of the automaton remains unchanged during the entire grammar parsing process (only the current state is changed, but the state transfer function is not changed), the grammar parsing method proposed in the present application adjusts the automaton as needed during the state transfer action and corrects the reading position of the lexer, so that the structure of the automaton is changed, and the reading order of multiple terminal words is also changed, thereby achieving the effect of selectively, cyclically or jumpily reading terminal words, that is, realizing the structuring of the terminal word stream.
[0123] The grammar parsing method proposed in this application no longer regards the input language (text) as a sequential stream of terminal words, but allows the terminal words themselves to have structure, making grammar parsing no longer a sequential process, but a structured process (i.e., branching, looping, and jumping are allowed during grammar parsing). This is conducive to effectively reusing existing results (such as reusing the original grammar) in grammar expansion and grammar compatibility, reducing the development cost and time cost of compilation, and significantly improving the development efficiency of compilation. At the same time, through dynamically variable automata, the design of grammar and syntax is more flexible, expanding the representation of grammar by orders of magnitude, completing tasks that were originally impossible to complete, and can be applied to fourth-generation computer languages.
[0124] The technical solution of this application is further explained below in conjunction with the specific application process.
[0125] For example, based on the grammar parsing method proposed in this application, the grammar of exponentiation can be implemented in C language. The implementation process of the grammar of exponentiation can be as follows:
[0126] First, add the terminal word power as a reserved word and use it as the exponentiation operator. Let the name of the terminal word power be P.
[0127] Then, a new grammar production is added: expression->expressionPexpression. This new grammar production is used as the target grammar production.
[0128] The grammar of exponentiation is generated through the newly added grammar production rules (ie target grammar production rules).
[0129] When a power b is read, and multiple terminal words and non-terminal words have been reduced to corresponding non-terminal words according to the target grammar production rule, that is, expression P expression, in the reduction action of further reducing it to the left expression in the target grammar production rule, the preset terminal word is inserted, and the multiple terminal words (that is, the terminal word sequence) is modified to: a power b repeat(*, goto(a), b)repeated.
[0130] The above terminal word sequence means: the preset terminal word jumps to a and reads b times in a loop. Each time a is read, a * is read and the parameter b is subtracted by 1 until it reaches 0 and the following repeated is read.
[0131] Furthermore, a new grammar production is added: expression->expressionrepeatexpressionrepeated. The value of the left-hand expression is assigned by the second expression on the right, and the value of the second expression on the right reuses the multiplication production: expression->expression*expression. This new grammar production is used as the new target grammar production.
[0132] Therefore, the above steps are equivalent to writing a series of multiplications. The effect is like dynamically editing the grammar productions and terminal word sequences, resulting in a production like this: expression->expressiona Pexpressionb repeatexpressiona*expressiona*expressiona*…*expressiona (a total of b a's). However, in practice, only a power b is written, thus reusing the multiplication productions, reducing the amount of code to write and significantly improving compilation efficiency.
[0133] Based on the grammar parsing method proposed in the embodiment of the present application, the grammar tree obtained by using the dynamically corrected automaton can reuse the implemented grammar elements to complete grammar expansion, thereby reducing the compilation time cost and development efficiency.
[0134] For example, Figure 5a As shown, based on the grammar parsing method proposed in this application, the two implemented grammars of UNION ALL and GROUP BY can be reused through grammar nesting to achieve the expansion of GROUPING SETS grammar. In this process, it is necessary to obtain and analyze the content of GROUPING SETS multiple times. The number of multiple constructions is determined by the GROUP BY query and UNION ALL operation determined by GROUPING SETS, as shown in the following example. Figure 5b shown.
[0135] The GROUP BY statement can be used to group data according to specified rules. UNION ALL can be used to combine the result sets of multiple SELECT statements into a single result set, returning all rows.
[0136] For example, Figure 6a As shown, based on the grammar parsing method proposed in this application, the WITH FUNCTION grammar extension can be implemented by reusing the call and parameter passing of local functions, where local functions are already implemented grammar. The WITH FUNCTION semantically requires that when it is called, it will be operated on by the defined expression, but it must use the semantics (context) of the called environment, and the name can be covered by a smaller visible scope.
[0137] like Figure 6b 、 Figure 6c As shown in the example, in WITH FUNCTION, the local function f(x,y) = x*100+y. Then, the local function f(x,y) can be called based on f(a,b) and f(c,d). For example, when a is 1 and b is 1, the call to f(x,y) determines that f(a,b) = a*100+b, resulting in f(a,b) being 101.
[0138] As can be seen from the above process, the grammar parsing method provided in the embodiment of the present application makes the design of grammar and syntax more flexible, expands the representation range of grammar by orders of magnitude, completes tasks that were originally impossible to complete, and further significantly improves the scalability of the compiled language.
[0139] Based on the same inventive concept, the present application also provides a grammar parsing device, such as Figure 7 FIG. 1 is a schematic diagram of the structure of a grammar parsing device provided by the present application, the device comprising:
[0140] Determination module 701, for determining a plurality of terminal words corresponding to the input text;
[0141] The first processing module 702 is used to determine a pending processing method for the terminal word and perform corresponding processing on the terminal word according to the pending processing method;
[0142] A modification module 703 is configured to modify the state of the automaton based on the processing of the terminal words, and adjust the reading order of the plurality of terminal words based on the modification of the state of the automaton to obtain a new reading order;
[0143] The second processing module 704 is used to continue to perform corresponding processing on the next terminal word according to the new reading order until the end condition is met to obtain the target result.
[0144] In a possible implementation, the determination module 701 is specifically configured to obtain the input text to be processed; perform lexical analysis on the input text to obtain corresponding multiple terminal words, and determine the original reading order of the multiple terminal words.
[0145] In one possible implementation, the first processing module 702 is specifically configured to determine, before reading the terminal word, whether the processing method for the terminal word is reduction processing or shift processing based on the terminal word and the current state of the automaton; if it is determined that the processing method is the reduction processing, then based on the target grammar production, multiple terminal words and non-terminal words before the corresponding terminal word are reduced to obtain corresponding non-terminal words, and during the reduction, the grammar production is modified or added.
[0146] In one possible implementation, the first processing module 702 is further used to move the corresponding terminal word into the automaton if it is determined that the processing method to be processed is the move-in processing; determine the target state of the automaton to be transferred based on the corresponding terminal word and the current state of the automaton, and change the state of the automaton from the current state to the target state.
[0147] In one possible implementation, the modification module 703 is specifically used to insert a preset terminal word when relocating the lexer, so that the lexer performs a specific action when reading the preset terminal word; wherein the specific action includes any one of a loop jump, a branch jump, and a modified terminal word, and the action parameters of the specific action are specified by the specification action or the shift action that creates the preset terminal word.
[0148] In a possible implementation, the end condition is that an end word is read and the current state of the automaton is an end state, wherein the end word is the next end word of the last end word in the original reading order.
[0149] Based on the same inventive concept, an electronic device is also provided in the embodiment of the present application. The electronic device can realize the function of the above-mentioned grammar parsing device, referring to Figure 8 , the above-mentioned electronic equipment includes:
[0150] At least one processor 801, and a memory 802 connected to the at least one processor 801. The specific connection medium between the processor 801 and the memory 802 is not limited in the embodiment of the present application. Figure 8 In the example, the processor 801 and the memory 802 are connected via a bus 800. Figure 8 The bus 800 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, Figure 8 The diagram is represented by only one thick line, but this does not mean that there is only one bus or one type of bus. Alternatively, the processor 801 may also be referred to as a controller, without limitation to the name.
[0151] In the embodiment of the present application, the memory 802 stores instructions that can be executed by at least one processor 801. The at least one processor 801 can execute the grammar parsing method discussed above by executing the instructions stored in the memory 802. The processor 801 can implement Figure 7 The functions of each module in the device shown.
[0152] Among them, the processor 801 is the control center of the device, which can use various interfaces and lines to connect the various parts of the entire control device, and monitor the device as a whole by running or executing instructions stored in the memory 802 and calling data stored in the memory 802, the various functions of the device and processing data.
[0153] In one possible design, processor 801 may include one or more processing units. Processor 801 may integrate an application processor and a modem processor. The application processor primarily processes the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 801. In some embodiments, processor 801 and memory 802 may be implemented on the same chip. In some embodiments, they may also be implemented on separate chips.
[0154] The processor 801 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the grammar parsing method disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor.
[0155] The memory 802 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory 802 may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (RAM), a static random access memory (SRAM), a programmable read-only memory (PROM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic memory, a disk, an optical disk, etc. The memory 802 is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 802 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.
[0156] By designing and programming the processor 801, the code corresponding to the grammar parsing method introduced in the above embodiment can be fixed into the chip, so that the chip can execute the code when running. Figure 3 The steps of the grammar parsing method of the embodiment shown are as follows: How to design and program the processor 801 is a technique well known to those skilled in the art and will not be described in detail here.
[0157] Based on the same inventive concept, an embodiment of the present application further provides a storage medium storing computer instructions. When the computer instructions are executed on a computer, the computer executes the grammar parsing method discussed above.
[0158] In some possible implementations, various aspects of the grammar parsing method provided in the present application can also be implemented in the form of a program product, which includes program code. When the program product is run on an apparatus, the program code is used to enable the control device to execute the steps of the grammar parsing method according to various exemplary implementations of the present application described above in this specification.
[0159] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0160] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0161] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0162] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0163] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. A grammar parsing method, characterized in that: include: Determine multiple terminal words corresponding to the input text; Determine a pending processing method for the terminal word, and perform corresponding processing on the terminal word according to the pending processing method; Modifying a state of an automaton based on the processing of the terminal words, and adjusting a reading order of the plurality of terminal words based on the modification of the state of the automaton to obtain a new reading order; According to the new reading order, the corresponding processing is continued for the next terminal word until the end condition is met to obtain the target result.
2. The method according to claim 1, wherein The step of determining the plurality of terminal words corresponding to the input text includes: Obtaining the input text to be processed; Performing lexical analysis on the input text to obtain corresponding multiple terminal words, and determining the original reading order of the multiple terminal words.
3. The method according to claim 1, wherein The determining of a pending processing method for the terminal word and performing corresponding processing on the terminal word according to the pending processing method includes: Before reading the terminal word, determining whether the processing mode for the terminal word is reduction processing or shift processing according to the terminal word and the current state of the automaton; If it is determined that the processing method to be processed is the reduction processing, then according to the target grammar production formula, multiple terminal words and non-terminal words before the corresponding terminal word are reduced to obtain the corresponding non-terminal words, and when reducing, the grammar production formula is modified or added.
4. The method according to claim 3, wherein After determining that the to-be-processed manner for the terminal word is reduction processing or shift processing, the method further includes: If it is determined that the processing method to be processed is the shift processing, the corresponding terminal word is shifted into the automaton; According to the corresponding terminal word and the current state of the automaton, a target state of the automaton to be transferred is determined, and the state of the automaton is transferred from the current state to the target state.
5. The method according to claim 1, wherein The adjusting the reading order of the multiple terminal words includes: When relocating the lexer, a preset terminal word is inserted so that the lexer performs a specific action when reading the preset terminal word; wherein the specific action includes any one of a loop jump, a branch jump, and a modified terminal word, and the action parameters of the specific action are specified by the specification action or the shift action that creates the preset terminal word.
6. The method according to claim 1, wherein The end condition is that an end word is read and the current state of the automaton is an end state, wherein the end word is the next end word of the last end word in the original reading order.
7. A grammar parsing device, characterized in that: include: A determination module, configured to determine a plurality of terminal words corresponding to an input text; A first processing module is used to determine a pending processing method for the terminal word and perform corresponding processing on the terminal word according to the pending processing method; a modification module, configured to modify a state of the automaton based on processing of the terminal words, and adjust a reading order of the plurality of terminal words based on the modification of the state of the automaton to obtain a new reading order; The second processing module is used to continue to perform corresponding processing on the next terminal word according to the new reading order until the end condition is met to obtain the target result.
8. The device according to claim 7, wherein The modification module is also used to insert a preset terminal word when relocating the lexer, so that the lexer performs a specific action when reading the preset terminal word; wherein the specific action includes any one of a loop jump, a branch jump, and a modified terminal word, and the action parameters of the specific action are specified by the specification action or the shift action that creates the preset terminal word.
9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the method steps of any one of claims 1 to 6 when executing the computer program stored in the memory.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps according to any one of claims 1 to 6 are implemented.