A stack-based SQL multi-statement parsing method, device, equipment and medium

CN122816638APending Publication Date: 2026-09-25SHANGHAI DAMENG DATABASE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610914271.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]有鉴于此,本发明的目的在于提供一种基于栈的SQL多语句解析方法、装置、设备及介质,能够解决分隔符误判、嵌套结构处理不足以及流式处理能力不足的问题,以实现对输入SQL脚本的准确、高效解析与断句

Benefits of technology

[0014]本申请中,通过预设词法分析器对输入的SQL字符流进行逐字符扫描,以扫描得到所述SQL字符流中的当前单词;若所述当前单词为非特征单词,则将所述当前单词追加至当前的预设缓冲区中;若所述当前单词为特征单词,则确定所述特征单词的单词类型;所述非特征单词为除所述特征单词之外的其他SQL关键字;所述特征单词为用于控制嵌套层级的SQL关键字;若所述单词类型为第一关键字,则将所述预设词法分析器切换至所述第一关键字对应的语义词法状态,以便根据相应的语法规则进行解析,并根据当前所述特征单词的出现次数,对当前的第一语法栈进行相应的语法标识压入,以及对所述预设缓冲区进行解析内容追加;所述第一关键字为用于表征复合语句块开始的关键字;若所述单词类型为第二关键字,则将所述预设词法分析器切换至所述第二关键字对应的语义词法状态,以便根据相应的语法规则进行解析,并对当前的第二语法栈进行相应的语法标识压入,以及对所述预设缓冲区进行解析内容追加;所述第二关键字为用于表征循环语句块或条件语句块开始的关键字;若所述单词类型为第三关键字,则弹出相应的目标语法栈的语法标识,并将表征块结束的内容追加至所述预设缓冲区;所述第三关键字为用于表征复合语句块、循环语句块或条件语句块结束的关键字;在完成对所述预设缓冲区的追加操作后,跳转至所述通过预设词法分析器对输入的SQL字符流进行逐字符扫描的步骤,直至满足预设结束条件,以输出所述预设缓冲区中的当前内容,得到当前的语句解析结果。由上可见,本申请通过预设词法分析器逐字符扫描输入的SQL字符流以获取当前单词,区分非特征单词与用于控制嵌套层级的特征单词,非特征单词直接追加至预设缓冲区,特征单词进一步判定单词类型,若为表征复合语句块开始的第一关键字,切换所述预设词法分析器的语义词法状态并根据单词出现次数对第一语法栈压入语法标识,若为表征循环或条件语句块开始的第二关键字,切换所述预设词法分析器的语义词法状态,并对第二语法栈压入语法标识,若为表征语句块结束的第三关键字,弹出对应语法栈的语法标识,并在完成任一单词类型的处理后,均进行预设缓冲区的内容追加,以及在完成追加操作后循环执行扫描步骤,直至满足预设结束条件输出所述预设缓冲区的内容,得到SQL语句解析结果。这样一来,通过本申请的上述过程,采用词法分析器逐字符扫描SQL字符流,能够精准拆解字符流获取有效单词,保证单词识别的细致度与准确性;区分非特征单词与控制嵌套层级的特征单词,实现普通单词与层级控制单词的分类处理,聚焦SQL嵌套结构解析核心需求;针对不同类型特征单词执行差异化处理,匹配复合、循环、条件语句块的开始与结束语法规则,管控语句嵌套层级;通过第一语法栈与第二语法栈分别压入、弹出语法标识,可动态记录并追踪多层级语句的嵌套关系,避免嵌套结构解析混乱;切换对应语义词法状态处理特征单词的技术特征,适配不同类型语句块的解析逻辑,提升复杂SQL语句解析的适配性,整体方案能精准识别并解析多层嵌套的复杂SQL语句,提升SQL语句解析的稳定性、准确性与嵌套结构处理能力,进而解决分隔符误判、嵌套结构处理不足以及流式处理能力不足的问题,以实现对输入SQL脚本的准确、高效解析与断句。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816638A_ABST
    Figure CN122816638A_ABST
Patent Text Reader

Abstract

The application discloses a stack-based SQL multi-statement parsing method and device, equipment and medium, and relates to the technical field of database management. The method comprises the following steps: a preset lexical analyzer is used to scan SQL character streams character by character to obtain a current word; if the word is a non-feature word, the word is added to a preset buffer; if the word is a feature word, the type of the word is determined; if the word is a keyword representing the start of a composite statement block, the semantic lexical state of the preset lexical analyzer is switched, and the syntax identifier of a first syntax stack is pressed in according to the number of occurrences of the current word; if the word is a keyword representing the start of a loop statement block or a conditional statement block, the semantic lexical state of the preset lexical analyzer is switched, and the syntax identifier of a second syntax stack is pressed in; if the word is a keyword representing the end of a block, the syntax identifier of a target syntax stack is popped; after the syntax stack operation is completed, the preset buffer is added, and the character-by-character scanning is continued until a preset ending condition is met, so that the parsing is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of database management technology, and in particular to a stack-based SQL multi-statement parsing method, apparatus, device, and medium. Background Technology

[0002] In current database development and management systems, it is often necessary to parse and segment SQL (Structured Query Language) script files for execution, formatting, or further semantic analysis. Existing technologies for parsing SQL statements typically suffer from the following problems: Delimiter misinterpretation: Traditional methods often use semicolons (;) as the end marker for SQL statements. However, in scenarios involving strings, comments, or complex objects such as stored procedures, triggers, and function definitions, statements may contain multiple semicolons, leading to premature truncation or incorrect segmentation. Insufficient support for nested and complex structures: Regular expressions and other keyword matching methods are typically linear scans without the concept of "nesting levels," and cannot correctly handle multi-level nested structures such as BEGIN…END, IF…ELSE, and LOOP. Syntax parsing has limitations: using a lexical analyzer like Flex, whose core function is to split the input string into tokens (syntactic units / words), but it lacks nested semantics and is prone to parsing errors; relying directly on a syntax parser requires writing complete rules in the grammar, and when encountering IF, CASE, LOOP, etc., the rules become increasingly large, resulting in high maintenance costs; at the same time, syntax parsing requires that the input SQL statement must fully conform to the syntax parsing rules, otherwise a syntax tree cannot be parsed.

[0003] In summary, how to solve the problems of misjudgment of delimiters, insufficient handling of nested structures, and insufficient streaming processing capabilities in order to achieve accurate and efficient parsing and sentence segmentation of input SQL scripts is an urgent problem to be solved. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a stack-based SQL multi-statement parsing method, apparatus, device, and medium that can solve the problems of misjudgment of delimiters, insufficient handling of nested structures, and insufficient streaming processing capabilities, so as to achieve accurate and efficient parsing and sentence segmentation of input SQL scripts. The specific solution is as follows: Firstly, this application provides a stack-based SQL multi-statement parsing method, including: The input SQL character stream is scanned character by character by a preset lexical analyzer to obtain the current word in the SQL character stream; If the current word is a non-feature word, then the current word is appended to the current preset buffer; if the current word is a feature word, then the word type of the feature word is determined; the non-feature word is any SQL keyword other than the feature word; the feature word is an SQL keyword used to control the nesting level. If the word type is the first keyword, the preset lexical analyzer is switched to the semantic lexical state corresponding to the first keyword so as to parse according to the corresponding grammar rules, and according to the occurrence frequency of the current feature word, the corresponding grammar identifier is pushed into the current first grammar stack, and the parsed content is appended to the preset buffer; the first keyword is a keyword used to represent the beginning of a compound statement block; If the word type is the second keyword, the preset lexical analyzer is switched to the semantic lexical state corresponding to the second keyword so as to parse according to the corresponding grammar rules, push the corresponding grammar mark into the current second grammar stack, and append the parsed content to the preset buffer; the second keyword is a keyword used to represent the beginning of a loop statement block or a conditional statement block; If the word type is a third keyword, the corresponding target syntax stack syntax identifier is popped out, and the content indicating the end of the block is appended to the preset buffer; the third keyword is a keyword used to indicate the end of a compound statement block, a loop statement block, or a conditional statement block; After completing the append operation to the preset buffer, the process jumps to the step of scanning the input SQL character stream character by character using a preset lexical analyzer until a preset termination condition is met, so as to output the current content in the preset buffer and obtain the current statement parsing result.

[0005] Optionally, the step of pushing corresponding grammatical identifiers into the current first grammar stack based on the occurrence frequency of the current feature words, and appending parsed content to the preset buffer, includes: If the occurrence count is one, then the compound statement block identifier and the keyword identifier of the first keyword are pushed sequentially into the current first syntax stack; The characters scanned before the first keyword are appended to the preset buffer, and the first keyword is appended to the preset buffer; The characters scanned after the first keyword are rolled back to the SQL character stream, and the corresponding characters in the preset buffer are deleted.

[0006] Optionally, the step of pushing corresponding grammatical identifiers into the current first grammar stack based on the occurrence frequency of the current feature words, and appending parsed content to the preset buffer, includes: If the number of occurrences is not one, then push the keyword identifier of the first keyword into the current first syntax stack; The first keyword is appended to the preset buffer; The characters scanned after the first keyword are rolled back to the SQL character stream, and the corresponding characters in the preset buffer are deleted.

[0007] Optionally, the step of pushing the corresponding syntax markers onto the current second syntax stack and appending parsed content to the preset buffer includes: Push the keyword identifier of the second keyword onto the current second syntax stack, and append the second keyword to the preset buffer; The characters scanned after the second keyword are rolled back to the SQL character stream, and the corresponding characters in the preset buffer are deleted.

[0008] Optionally, popping the syntax identifier of the corresponding target syntax stack and appending the content representing the end of the block to the preset buffer includes: If the third keyword is a keyword used to indicate the end of a compound statement block, then the syntax identifier at the top of the current first syntax stack is popped, and the third keyword is appended to the preset buffer. If the first syntax stack is not empty, then determine the current syntax identifier at the top of the first syntax stack; If the current syntax identifier is a compound statement block identifier, then the compound statement block identifier is popped up; if the current syntax identifier is not a compound statement block identifier, then the characters scanned after the third keyword are rolled back to the SQL character stream, and the corresponding characters in the preset buffer are deleted.

[0009] Optionally, after popping up the compound statement block identifier, the method further includes: If the first syntax stack is detected to be empty, then the preset termination condition is determined to be met.

[0010] Optionally, if the word type is a third keyword, then the corresponding grammar identifier of the target grammar stack is popped out, and the content at the end of the representation block is appended to the preset buffer, including: If the third keyword is a keyword used to indicate the end of a loop statement block or a conditional statement block, then the syntax identifier at the top of the current second syntax stack is popped, and the third keyword is appended to the preset buffer. If the second syntax stack is not empty, then the characters scanned after the third keyword are rolled back to the SQL character stream, and the corresponding characters in the preset buffer are deleted.

[0011] Secondly, this application provides a stack-based SQL multi-statement parsing apparatus, comprising: The character stream scanning module is used to scan the input SQL character stream character by character using a preset lexical analyzer to obtain the current word in the SQL character stream; The type determination module is used to append the current word to the current preset buffer if the current word is a non-feature word; and to determine the word type of the feature word if the current word is a feature word. The non-feature word is any SQL keyword other than the feature word; the feature word is any SQL keyword used to control the nesting level. The first content appending module is used to switch the preset lexical analyzer to the semantic lexical state corresponding to the first keyword if the word type is the first keyword, so as to parse according to the corresponding grammar rules, push the corresponding grammar mark into the current first grammar stack according to the occurrence frequency of the current feature word, and append the parsed content to the preset buffer; the first keyword is a keyword used to represent the beginning of a compound statement block; The second content appending module is used to switch the preset lexical analyzer to the semantic lexical state corresponding to the second keyword if the word type is the second keyword, so as to parse according to the corresponding grammar rules, push the corresponding grammar mark into the current second grammar stack, and append the parsed content to the preset buffer; the second keyword is a keyword used to represent the beginning of a loop statement block or a conditional statement block; The third content appending module is used to pop up the syntax identifier of the corresponding target syntax stack if the word type is a third keyword, and append the content indicating the end of the block to the preset buffer; the third keyword is a keyword used to indicate the end of a compound statement block, a loop statement block, or a conditional statement block; The step jump module is used to jump to the step of scanning the input SQL character stream character by character by character through the preset lexical analyzer after the append operation to the preset buffer is completed, until the preset termination condition is met, so as to output the current content in the preset buffer and obtain the current statement parsing result.

[0012] Thirdly, this application provides an electronic device, comprising: Memory, used to store computer programs; A processor is used to execute the computer program to implement the aforementioned stack-based SQL multi-statement parsing method.

[0013] Fourthly, this application provides a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned stack-based SQL multi-statement parsing method.

[0014] In this application, a preset lexical analyzer scans the input SQL character stream character by character to obtain the current word in the SQL character stream. If the current word is a non-feature word, it is appended to the current preset buffer. If the current word is a feature word, its word type is determined. The non-feature word is any SQL keyword other than the feature word. The feature word is an SQL keyword used to control nesting levels. If the word type is a first keyword, the preset lexical analyzer is switched to the semantic lexical state corresponding to the first keyword to parse according to the corresponding syntax rules. Based on the occurrence count of the current feature word, corresponding syntax markers are pushed onto the current first syntax stack, and parsed content is appended to the preset buffer. The first keyword is a keyword used to represent the beginning of a compound statement block. If the word type is the second keyword, the preset lexical analyzer is switched to the semantic lexical state corresponding to the second keyword to parse according to the corresponding grammar rules, push the corresponding grammar identifier into the current second grammar stack, and append the parsed content to the preset buffer; the second keyword is a keyword used to represent the beginning of a loop statement block or a conditional statement block. If the word type is the third keyword, the grammar identifier of the corresponding target grammar stack is popped, and the content representing the end of the block is appended to the preset buffer; the third keyword is a keyword used to represent the end of a compound statement block, a loop statement block, or a conditional statement block. After completing the append operation to the preset buffer, the process jumps to the step of scanning the input SQL character stream character by character by character through the preset lexical analyzer until the preset end condition is met, so as to output the current content in the preset buffer and obtain the current statement parsing result. As can be seen from the above, this application uses a preset lexical analyzer to scan the input SQL character stream character by character to obtain the current word, distinguishes between non-feature words and feature words used to control nesting levels, directly appends non-feature words to a preset buffer, and further determines the word type of feature words. If it is the first keyword indicating the start of a compound statement block, the semantic lexical state of the preset lexical analyzer is switched and a grammar identifier is pushed onto the first grammar stack according to the word's occurrence count. If it is the second keyword indicating the start of a loop or conditional statement block, the semantic lexical state of the preset lexical analyzer is switched and a grammar identifier is pushed onto the second grammar stack. If it is the third keyword indicating the end of a statement block, the grammar identifier corresponding to the grammar stack is popped. After processing any word type, the content of the preset buffer is appended, and after the append operation is completed, the scanning steps are executed cyclically until the preset end condition is met and the content of the preset buffer is output to obtain the SQL statement parsing result.In this way, through the process described above in this application, a lexical analyzer scans the SQL character stream character by character, accurately disassembling the character stream to obtain effective words, ensuring the detail and accuracy of word recognition; it distinguishes between non-feature words and feature words that control nesting levels, realizing the classification and processing of ordinary words and level control words, focusing on the core needs of SQL nested structure parsing; it performs differentiated processing for different types of feature words, matching the start and end syntax rules of compound, loop, and conditional statement blocks, and controlling the nesting level of statements; by pushing and popping syntax markers into the first syntax stack and the second syntax stack respectively, it can dynamically record and track the nesting relationship of multi-level statements, avoiding chaotic nested structure parsing; it switches the technical features of corresponding semantic lexical states to process feature words, adapting to the parsing logic of different types of statement blocks, improving the adaptability of complex SQL statement parsing; the overall solution can accurately identify and parse complex SQL statements with multiple nesting levels, improving the stability, accuracy, and nested structure processing capability of SQL statement parsing, thereby solving the problems of misjudgment of delimiters, insufficient nested structure processing, and insufficient streaming processing capability, so as to achieve accurate and efficient parsing and sentence segmentation of input SQL scripts. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0016] Figure 1 This is a flowchart of a stack-based SQL multi-statement parsing method disclosed in this application; Figure 2 This is a flowchart illustrating a stack-based SQL multi-statement parsing method disclosed in this application; Figure 3 This is a schematic diagram of the structure of a stack-based SQL multi-statement parsing device disclosed in this application; Figure 4 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Existing technologies for parsing SQL statements typically suffer from the following problems: **Delimiter misinterpretation:** Traditional methods often use semicolons (;) as the end marker for SQL statements. However, in scenarios involving strings, comments, or complex objects such as stored procedures, triggers, and function definitions, multiple semicolons may be present within the statement, leading to premature truncation or incorrect segmentation. **Insufficient support for nested and complex structures:** Regular expressions and other keyword matching methods are usually linear scans without the concept of "nesting levels," making it unable to correctly handle multi-level nested structures such as BEGIN…END, IF…ELSE, and LOOP. **Limitations of syntax parsing:** Lexical analyzers like Flex primarily segment the input string into tokens (syntactic units), but they lack nesting semantics, making them prone to parsing errors. Direct reliance on a syntax parser requires writing complete rules in the grammar, which become increasingly voluminous and costly to maintain when encountering IF, CASE, and LOOP structures. Furthermore, syntax parsing requires the input SQL statement to fully conform to the parsing rules; otherwise, a syntax tree cannot be generated.

[0019] To overcome the aforementioned technical problems, this application provides a stack-based SQL multi-statement parsing method that can solve the problems of misjudgment of delimiters, insufficient handling of nested structures, and insufficient streaming processing capabilities, so as to achieve accurate and efficient parsing and sentence segmentation of input SQL scripts.

[0020] See Figure 1 As shown, this embodiment of the invention discloses a stack-based SQL multi-statement parsing method, including: Step S11: Scan the input SQL character stream character by character using a preset lexical analyzer to obtain the current word in the SQL character stream.

[0021] In this embodiment, a preset lexical analyzer performs a character-by-character scan on the input SQL character stream to extract the current word. This preset lexical analyzer can be a lexical analyzer generated using the Flex tool, and it has different semantic lexical states. The analyzer can be switched to a specified lexical state to continue scanning, resulting in different token matching rules in different states, enabling accurate identification of various grammatical structures.

[0022] It should be noted that current database development and management systems often require parsing and segmenting SQL script files for execution, formatting, or further semantic analysis. Existing technologies typically suffer from problems such as misjudgment of delimiters, insufficient support for nested and complex structures, limitations in syntax parsing, and a lack of dynamic context control. Therefore, this application proposes a stack-based SQL multi-statement parsing method. By combining stack management and lexical state switching, it effectively solves the problems of misjudgment of delimiters, insufficient handling of nested structures, and insufficient streaming processing capabilities in existing technologies, thereby achieving accurate and efficient parsing and segmentation of input SQL scripts. Specifically, based on Flex lexical parsing, a stack mechanism is introduced for semantic-level control. During parsing, a marker is pushed onto the stack whenever a key syntax block such as BEGIN, WITH, or IF is encountered; when a closing marker such as END or END IF is encountered, a pop-out matching is performed. The overall system includes the following components: a lexical analyzer, a multi-stack management module, a lexical state switching module, and a statement caching module. The lexical analyzer, based on a Flex tool, can scan the SQL input character by character. Multi-stack management module: This module utilizes multiple stacks working collaboratively to record different syntax levels. These include the main syntax stack (for block structures), the ifLoopStack (for conditional and loop statements), and the procedureDefineStack (for stored procedures / functions). These stacks manage ordinary statements, loop statements, and procedure definition statements, enabling the parser to distinguish different syntax scopes and perform matching step-by-step. Lexical state switching module: By defining different lexical states (such as the general SQL state GENERAL_SQL_CLAUSE, the BEGIN block state BEGIN_SQL_CLAUSE, the IF / LOOP state IF_LOOP_BLOCK, the stored procedure definition state PROCEDURE_DEFINE_SQL_CLAUSE, and the initial state YYINITIAL), the lexer is switched to a specified lexical state to continue scanning. Different token matching rules exist in different states, allowing for accurate identification of various syntax structures. Statement caching module: This module uses a StringBuilder as a buffer (buf) to accumulate SQL characters. When a boundary condition is reached, the complete statement is output. Figure 2The diagram shows a flowchart of a stack-based SQL multi-statement parsing method provided in this application. It mainly includes the following key steps: Lexical analyzer scanning: The input SQL character stream is scanned character by character using a lexical analyzer to identify feature tokens, such as block start markers BEGIN, IF, WHILE, and block end marker END. Stack state management and lexical state switching: When a block start marker is first identified, the block identifier and keyword identifier are pushed onto the stack sequentially; if it is an inner block start marker, only the keyword identifier is pushed; when a block end marker is identified, the keyword identifier at the top of the stack is popped; the stack top is checked again to see if it is a block identifier, and if so, it continues to pop; different syntax parsing modes are switched according to the currently scanned feature tokens. Buffer management: Scanned non-feature tokens are stored in the buffer. Block integrity judgment: When a block end marker is identified and the stack is empty, the current SQL block is determined to be completely finished, and the complete statement is output; if the stack is not empty, the read characters are recycled into the buffer. In this way, this embodiment uses a preset lexical analyzer to provide a standardized and stable parsing foundation, ensuring the standardization and consistency of SQL character processing; by performing character-by-character scanning, it can accurately identify word boundaries and valid content in the character stream, avoid parsing omissions, and ensure accurate extraction of the current word, providing a basic unit for subsequent SQL syntax parsing and statement processing.

[0023] Step S12: If the current word is a non-feature word, then append the current word to the current preset buffer; if the current word is a feature word, then determine the word type of the feature word; the non-feature word is any SQL keyword other than the feature word; the feature word is an SQL keyword used to control the nesting level.

[0024] In this embodiment, the current words obtained from the scan are classified. Non-feature words are directly appended to a preset buffer, while feature words are determined by their word type. Non-feature words are ordinary SQL keywords, while feature words are SQL keywords that control nesting levels, such as BEGIN, IF, WHILE, and END. This approach distinguishes between feature words and non-feature words, focusing on keywords related to nesting level control and providing special processing for key syntax blocks. It adds a layer of dynamic context tracking to the SQL parsing process, ensuring that nested structures are correctly identified and matched. Appending non-feature words to the preset buffer and storing ordinary keywords in an orderly manner ensures a continuous and uninterrupted parsing process. Identifying feature words and determining their word types provides a basis for tracking and controlling SQL nesting levels, improving the accuracy of syntax parsing and the efficiency of hierarchical processing.

[0025] Step S13: If the word type is the first keyword, the preset lexical analyzer is switched to the semantic lexical state corresponding to the first keyword so as to parse according to the corresponding grammar rules, and according to the occurrence frequency of the current feature word, the corresponding grammar identifier is pushed into the current first grammar stack, and the parsed content is appended to the preset buffer; the first keyword is a keyword used to represent the beginning of a compound statement block.

[0026] In this embodiment, after identifying the first keyword representing the start of a compound statement block, namely BEGIN, the preset lexical analyzer is switched to the corresponding semantic lexical state BEGIN_SQL_CLAUSE. Syntax identifiers are pushed onto the first syntax stack based on the frequency of this word, and parsed content is appended to the buffer. It should be noted that the first syntax stack is not only used for compound statements, but also for ordinary syntax blocks (pushing GeneralSql onto the stack), packages (pushing package, begin, etc. onto the stack), and serves as the stack for determining whether a statement has ended.

[0027] It should be noted that the process of pushing corresponding syntax markers onto the current first syntax stack based on the occurrence count of the current feature word, and appending parsed content to the preset buffer, is as follows: If the occurrence count is one, the compound statement block marker and the keyword marker of the first keyword are pushed onto the current first syntax stack in sequence; the characters scanned before the first keyword are appended to the preset buffer, and the first keyword is appended to the preset buffer; the characters scanned after the first keyword are rolled back to the SQL character stream, and the corresponding characters in the preset buffer are deleted. That is, when the compound statement block start marker is first identified, that is, when the first keyword appears for the first time, the compound statement block marker beginClause and the keyword marker begin are pushed onto the first syntax stack in sequence, the characters before the first keyword are stored in the preset buffer, the first keyword is appended to the preset buffer, the scanned characters after the keyword are rolled back to the SQL character stream, and the corresponding redundant characters in the preset buffer are cleared.

[0028] It should be further pointed out that the process of pushing the corresponding grammar identifiers into the current first grammar stack based on the occurrence count of the current feature word, and appending the parsed content to the preset buffer, is as follows: If the occurrence count is not one, the keyword identifier of the first keyword is pushed into the current first grammar stack; the first keyword is appended to the preset buffer; the characters scanned after the first keyword are rolled back to the SQL character stream, and the corresponding characters in the preset buffer are deleted. That is, when the first keyword is not appearing for the first time, only the keyword identifier "begin" is pushed into the first grammar stack, the first keyword is stored in the preset buffer, and the subsequently scanned characters of the first keyword are rolled back to the original SQL character stream, and the corresponding redundant characters in the preset buffer are cleared. In this way, this embodiment switches the dedicated semantic lexical state, matches the grammar rules of compound statement blocks, ensures that various grammatical structures can be identified, and achieves high-precision parsing of complex grammatical structures; it combines word occurrence counts to operate the grammar stack, accurately records the nesting level of statements; and it synchronously appends content to the buffer, retains parsed data, and realizes hierarchical tracking and continuous parsing of SQL compound statement blocks.

[0029] Step S14: If the word type is a second keyword, the preset lexical analyzer is switched to the semantic lexical state corresponding to the second keyword so as to parse according to the corresponding grammar rules, push the corresponding grammar mark into the current second grammar stack, and append the parsed content to the preset buffer; the second keyword is a keyword used to represent the beginning of a loop statement block or a conditional statement block.

[0030] In this embodiment, after identifying the second keyword representing the start of a loop or conditional statement block, such as IF or LOOP, the preset lexical analyzer is switched to the corresponding semantic lexical state IF_LOOP_BLOCK, parsed according to the corresponding syntax rules, and simultaneously a syntax identifier is pushed onto the second syntax stack ifLoopStack, and parsed content is appended to the preset buffer. That is, the second syntax stack is used to determine whether the if...end if or loop...end loop structure has ended.

[0031] It should be noted that the process of pushing the corresponding syntax identifiers onto the current second syntax stack and appending parsed content to the preset buffer is as follows: Push the keyword identifier of the second keyword onto the current second syntax stack, and append the second keyword to the preset buffer; roll back the characters scanned after the second keyword to the SQL character stream, and delete the corresponding characters in the preset buffer. That is, push the second keyword identifier (if / loop, etc.) onto the second syntax stack, store the second keyword in the preset buffer, and simultaneously roll back the scanned characters after the second keyword to the original SQL character stream, and delete the corresponding redundant characters in the preset buffer. In this way, this embodiment switches to a dedicated semantic lexical state, matches the syntax rules of conditional and loop statements, and ensures that nested structures are correctly identified and matched; it independently uses the second syntax stack to push the identifiers, separately tracks the level of such statement blocks, continuously records the nesting level of conditional and loop statement blocks, and manages them separately from the level of compound statement blocks; it synchronously appends content to the buffer, completely retains the parsed data, and achieves parallel and accurate parsing of multiple types of statement blocks.

[0032] Step S15: If the word type is a third keyword, the corresponding target syntax stack syntax identifier is popped out, and the content representing the end of the block is appended to the preset buffer; the third keyword is a keyword used to represent the end of a compound statement block, a loop statement block, or a conditional statement block.

[0033] In this embodiment, after identifying the third keyword representing the end of various statement blocks, such as END, END IF, END LOOP, etc., the syntax identifier is popped from the corresponding syntax stack, and the relevant content of the end of the statement block is written into the preset buffer.

[0034] In one specific implementation, if the third keyword is a keyword indicating the end of a compound statement block, the syntax identifier at the top of the current first syntax stack is popped, and the third keyword is appended to the preset buffer. If the current first syntax stack is not empty, the current syntax identifier at the top of the first syntax stack is determined. If the current syntax identifier is a compound statement block identifier, the compound statement block identifier is popped. If the current syntax identifier is not a compound statement block identifier, the characters scanned after the third keyword are rolled back to the SQL character stream, and the corresponding characters in the preset buffer are deleted. That is, when the third keyword END indicating the end of a compound statement block is identified, the top identifier of the first syntax stack is popped and the third keyword is stored in the preset buffer. If the first syntax stack is not empty at this time, the current top identifier is determined. If it is the compound statement block identifier beginClause, it continues to pop; otherwise, the subsequent characters are rolled back and the corresponding content in the buffer is cleared. It should be noted that after popping the compound statement block identifier, it is also necessary to determine whether the preset termination condition is met to determine whether the current SQL character stream has finished scanning. The processing flow is as follows: if the first syntax stack is detected to be empty, it is determined that the preset termination condition is met. That is, after popping the compound statement block identifier, it is checked whether the first syntax stack is empty. If it is empty, it is determined that the preset termination condition has been met. It can be understood that the termination condition is passively detected. For example, in a normal SQL statement, when a semicolon is detected, a pop operation is performed on the stack. If it is found that the stack is empty after popping, it is considered that the normal SQL statement has ended, the lexical state is switched to YYINITIAL, the complete statement is output, the buffer is cleared, and the next statement is processed. If it is a compound statement, when the END keyword is detected, the main stack is also popped. If it is empty, it is considered that there is no compound block, which is consistent with the above operation. If it is not empty, it is considered that there is still an outer compound structure. That is, the preset termination condition is to determine whether the stack (first syntax stack) is empty when the END or semicolon keyword is detected.

[0035] In another specific implementation, if the third keyword is a keyword used to indicate the end of a loop or conditional statement block, then the syntax identifier at the top of the current second syntax stack is popped, and the third keyword is appended to the preset buffer; if the current second syntax stack is not empty, then the characters scanned after the third keyword are rolled back to the SQL character stream, and the corresponding characters in the preset buffer are deleted. That is, when a third keyword of the loop or conditional statement block type is identified, such as END IF or END LOOP, the top identifier of the second syntax stack is popped, and the third keyword is written to the preset buffer; if the second syntax stack is not empty, then the subsequent scanned characters are rolled back, and the corresponding contents of the buffer are cleared. In this way, this embodiment pops the stack top marker layer by layer, unregisters nested statement block levels in an orderly manner, matches the closing logic of nested structures, and matches the start and end structures of statements; multi-stack joint detection uses stack emptiness as the end judgment criterion, uniformly verifies whether all nested levels of statement blocks are closed, determines the end point of SQL statement parsing, ensures that the entire statement structure is complete and there are no unclosed blocks, and improves the effectiveness of parsing results; by tracing the context through the stack, conflicts caused by relying solely on syntax rules are avoided, and there is no need to write a large number of complex syntax rules.

[0036] Step S16: After completing the append operation to the preset buffer, jump to the step of scanning the input SQL character stream character by character by character through the preset lexical analyzer until the preset termination condition is met, so as to output the current content in the preset buffer and obtain the current statement parsing result.

[0037] In this embodiment, after each append to the preset buffer, the process returns to the character scanning step and repeats until a preset termination condition is triggered. Finally, the content of the preset buffer is output as the statement parsing result. This embodiment achieves uninterrupted segment-by-segment processing of the SQL character stream through cyclical jumps, ensuring complete parsing of the entire statement. Context tracing via stack avoids conflicts arising from relying solely on syntax rules and eliminates the need to write numerous complex syntax rules. Even if user SQL is not entirely standard, the stack can help quickly locate errors and specify error information. Extensibility is also enhanced; simply expanding the stack tag types allows for rapid support of new syntax blocks without large-scale modifications to syntax rules. Thus, while maintaining the efficiency of Flex lexical analysis, high-precision parsing of complex syntax structures is achieved.

[0038] As can be seen from the above, this embodiment of the application uses a preset lexical analyzer to scan the input SQL character stream character by character to obtain the current word, distinguishes between non-feature words and feature words used to control the nesting level, directly appends non-feature words to the preset buffer, and further determines the word type of feature words. If it is the first keyword representing the start of a compound statement block, the semantic lexical state of the preset lexical analyzer is switched and a grammar identifier is pushed onto the first grammar stack according to the number of times the word appears. If it is the second keyword representing the start of a loop or conditional statement block, the semantic lexical state of the preset lexical analyzer is switched and a grammar identifier is pushed onto the second grammar stack. If it is the third keyword representing the end of a statement block, the grammar identifier of the corresponding grammar stack is popped. After processing any word type, the contents of the preset buffer are appended, and after the append operation is completed, the scanning steps are executed cyclically until the preset end condition is met and the contents of the preset buffer are output to obtain the SQL statement parsing result. In this way, through the above process of the embodiments of this application, the lexical analyzer scans the SQL character stream character by character, which can accurately decompose the character stream to obtain effective words, ensuring the detail and accuracy of word recognition; it distinguishes between non-feature words and feature words that control the nesting level, realizing the classification and processing of ordinary words and level control words, focusing on the core needs of SQL nested structure parsing; it performs differentiated processing for different types of feature words, matching the start and end syntax rules of compound, loop, and conditional statement blocks, and controlling the nesting level of statements; by pushing and popping syntax markers into the first syntax stack and the second syntax stack respectively, it can dynamically record and track the nesting relationship of multi-level statements, avoiding chaotic nested structure parsing; it switches the technical features of corresponding semantic lexical state to process feature words, adapting to the parsing logic of different types of statement blocks, improving the adaptability of complex SQL statement parsing, and the overall solution can accurately identify and parse complex SQL statements with multiple nesting, improving the stability, accuracy and nesting structure processing capability of SQL statement parsing, thereby solving the problems of misjudgment of delimiters, insufficient nesting structure processing and insufficient streaming processing capability, so as to achieve accurate and efficient parsing and sentence segmentation of input SQL scripts.

[0039] As can be seen from the previous embodiment, this application discloses a stack-based SQL multi-statement parsing method, which can solve the problems of misjudgment of delimiters, insufficient handling of nested structures, and insufficient streaming processing capabilities, so as to achieve accurate and efficient parsing and sentence segmentation of input SQL scripts. Next, taking a nested BEGIN...END structure as an example, the stack-based SQL multi-statement parsing method will be described in detail. The example SQL is as follows: BEGIN INSERT INTO t VALUES(1); BEGIN UPDATE t SET id = 2; END; END; This SQL statement contains two nested BEGIN...END blocks. The outer block contains an INSERT statement and a complete inner BEGIN...END block, which in turn contains an UPDATE statement.

[0040] Phase 1: Initial State. The system is in its initial state, all stacks are empty, buffers are empty, and the lexical analyzer is in its initial state, ready to begin scanning the input SQL character stream.

[0041] Phase Two: Identifying the Outer BEGIN (First Layer). When the lexical analyzer scans the first BEGIN keyword, it triggers the outer BEGIN block identification logic as follows: The lexical analyzer switches to BEGIN block processing state -> clears the main syntax stack and control structure stack -> pushes the "beginClause" block identifier and the "begin" keyword identifier onto the main syntax stack in sequence -> initializes the current SQL buffer -> appends the preceding invalid character to the buffer -> appends BEGIN to the buffer -> rolls back the whitespace characters after BEGIN to the input stream -> deletes the last character (\n or space) from the buffer. At this time, the main syntax stack contains two elements, "beginClause" and "begin", indicating that it is currently in the first level of a BEGIN block.

[0042] Phase 3: Scanning the outer block content (INSERT statement). The lexical analyzer continues scanning under the outer BEGIN block state BEGIN_SQL_CLAUSE, sequentially identifying tokens such as INSERT, INTO, t, and VALUES(1). Since these tokens are all non-feature tokens, the system directly appends them to the current SQL buffer. The buffer content gradually accumulates, becoming "BEGIN\nINSERT INTO t VALUES(1);". The main syntax stack state remains unchanged, still containing the hierarchy identifier of the outer BEGIN block.

[0043] Phase Four: Identifying the Inner BEGIN (Second Layer). The lexical analyzer continues scanning. When it identifies the second BEGIN keyword, the following logic for identifying the inner BEGIN block is triggered: The lexical analyzer remains in the BEGIN block processing state (because the inner and outer blocks use the same state) -> pushes the "begin" keyword identifier back onto the main syntax stack -> appends BEGIN to the buffer -> rolls back the whitespace characters following BEGIN to the input stream -> deletes the last character (\n or space) from the buffer. The state of the main syntax stack changes as follows: Before pushing: ["beginClause", "begin"] (outer block identifier); After pushing: ["beginClause", "begin", "begin"] (outer + inner block identifier).

[0044] Key code snippet (block start processing): / / When the BEGIN keyword is detected, perform the following operations. stack.push("begin"); / / Push the keyword identifier In this way, multiple levels of nested block identifiers accumulate in the main syntax stack, with the topmost element always representing the innermost block currently being processed.

[0045] Phase 5: Scanning the inner block content (UPDATE statement). The lexical analyzer continues scanning in the inner BEGIN block state, identifying tokens such as UPDATE, t, SET, id, =, and 2. Since these tokens are non-feature tokens, the system directly appends them to the current SQL buffer. At this time, the buffer content continues to accumulate, becoming: BEGIN \nINSERT INTO t VALUES(1); \nBEGIN \nUPDATE t SET id = 2; The main syntax stack state remains unchanged, containing identification information for two layers of blocks.

[0046] Phase Six: Identifying the Inner END (End of Second Layer). When the lexical analyzer detects the first END keyword, it triggers the following block end identification logic for the inner block: First, append END; to the buffer. The buffer becomes: BEGIN \nINSERT INTO t VALUES(1); \nBEGIN \nUPDATE t SET id = 2; \nEND; —> Pop the top element of the main syntax stack (matching the inner "begin") —> Check the top of the stack again and find it is not empty —> Backtrack one character: semicolon ";" —> Delete the last character (semicolon) from the buffer. The state of the main syntax stack changes as follows: Before popping: ["beginClause", "begin", "begin"]; After popping: ["beginClause", "begin"] (outer block identifier is preserved).

[0047] Key code snippet (block end handling): / / When the END keyword is detected, perform the following operations. stack.pop(); / / Pops "begin" from the stack. if (stack.peek().equals("beginClause")) { stack.pop(); / / Pops the block identifier "beginClause" } / / Check if there are any outer blocks if (stack.isEmpty()) { / / The stack is empty, the current SQL block has completely ended, and the SQL is output. } else { / / The stack is not empty; there are still unfinished blocks on the outer layer. Backtrack semicolon. yypushback(1); buf.deleteCharAt(buf.length() - 1); } Phase 7: Continue scanning the outer block content. The lexical analyzer rereads the semicolon and appends it to the buffer —> the END keyword is detected again —> the top element of the main syntax stack is popped (matching the outer "begin") —> the top of the stack is checked again and found to be the "beginClause" block identifier, so it is popped again —> END; is appended to the buffer: BEGIN \nINSERT INTO tVALUES(1); \nBEGIN \nUPDATE t SET id = 2; \nEND; \nEND; The state of the main syntax stack changes as follows: Before popping: ["beginClause", "begin"]; After popping: [] (the stack is empty, and the outer block has also ended). Key point: This is the second processing of the END keyword. During the first processing, there was still an outer block; during the second processing, the outer block has also been completed, and the stack is empty, allowing the complete SQL statement to be output.

[0048] Phase 8: Output the complete SQL. Since the main syntax stack is empty, the system determines that the current SQL block (including nested inner blocks) has completely ended: Switch the lexical analyzer back to its initial state —> Call the SQL statement output method to output the complete SQL statements accumulated in the buffer —> Output the result: BEGIN INSERT INTO t VALUES(1); BEGIN UPDATE t SET id = 2; END; END; Accordingly, see Figure 3 As shown in the illustration, this application also provides a stack-based SQL multi-statement parsing device, including: The character stream scanning module 11 is used to scan the input SQL character stream character by character using a preset lexical analyzer to obtain the current word in the SQL character stream; The type determination module 12 is used to append the current word to the current preset buffer if the current word is a non-feature word; and to determine the word type of the feature word if the current word is a feature word; the non-feature word is an SQL keyword other than the feature word; the feature word is an SQL keyword used to control the nesting level. The first content appending module 13 is used to switch the preset lexical analyzer to the semantic lexical state corresponding to the first keyword if the word type is the first keyword, so as to parse according to the corresponding grammar rules, push the corresponding grammar mark into the current first grammar stack according to the occurrence frequency of the current feature word, and append the parsed content to the preset buffer; the first keyword is a keyword used to represent the beginning of a compound statement block; The second content appending module 14 is used to switch the preset lexical analyzer to the semantic lexical state corresponding to the second keyword if the word type is the second keyword, so as to parse according to the corresponding grammar rules, push the corresponding grammar mark into the current second grammar stack, and append the parsed content to the preset buffer; the second keyword is a keyword used to represent the beginning of a loop statement block or a conditional statement block. The third content appending module 15 is used to pop up the syntax identifier of the corresponding target syntax stack if the word type is a third keyword, and append the content indicating the end of the block to the preset buffer; the third keyword is a keyword used to indicate the end of a compound statement block, a loop statement block or a conditional statement block; The step jump module 16 is used to jump to the step of scanning the input SQL character stream character by character by character through the preset lexical analyzer after the append operation of the preset buffer is completed, until the preset end condition is met, so as to output the current content in the preset buffer and obtain the current statement parsing result.

[0049] In some specific embodiments, the first content appending module 13 may specifically include: The first identifier push unit is used to push the compound statement block identifier and the keyword identifier of the first keyword into the current first syntax stack in sequence if the occurrence count is one. A character appending unit is used to append the characters scanned before the first keyword to the preset buffer, and to append the first keyword to the preset buffer; The first character deletion unit is used to roll back the characters scanned after the first keyword to the SQL character stream and delete the corresponding characters in the preset buffer.

[0050] In some specific embodiments, the first content appending module 13 may specifically include: The second identifier pushing unit is used to push the keyword identifier of the first keyword into the current first syntax stack if the number of occurrences is not one. A keyword appending unit is used to append the first keyword to the preset buffer; The second character deletion unit is used to roll back the characters scanned after the first keyword to the SQL character stream and delete the corresponding characters in the preset buffer.

[0051] In some specific embodiments, the second content appending module 14 may specifically include: The third identifier push unit is used to push the keyword identifier of the second keyword into the current second syntax stack and append the second keyword to the preset buffer; The third character deletion unit is used to roll back the characters scanned after the second keyword to the SQL character stream and delete the corresponding characters in the preset buffer.

[0052] In some specific embodiments, the third content appending module 15 may specifically include: The first identifier pop-up unit is used to pop up the syntax identifier at the top of the current first syntax stack if the third keyword is a keyword used to indicate the end of a compound statement block, and to append the third keyword to the preset buffer. The identifier determination unit is used to determine the current syntax identifier located at the top of the first syntax stack if the current first syntax stack is not empty. The fourth character deletion unit is used to pop up the compound statement block identifier if the current syntax identifier is a compound statement block identifier; if the current syntax identifier is not a compound statement block identifier, it rolls back the characters scanned after the third keyword to the SQL character stream and deletes the corresponding characters in the preset buffer.

[0053] In some specific embodiments, the stack-based SQL multi-statement parsing device may further include: The condition determination unit is used to determine that the preset termination condition is met if the first syntax stack is detected to be empty.

[0054] In some specific embodiments, the third content appending module 15 may specifically include: The second identifier pop-up unit is used to pop up the syntax identifier at the top of the current second syntax stack if the third keyword is a keyword used to indicate the end of a loop statement block or a conditional statement block, and to append the third keyword to the preset buffer. The fifth character deletion unit is used to, if the current second syntax stack is not empty, backtrack the characters scanned after the third keyword to the SQL character stream and delete the corresponding characters in the preset buffer.

[0055] Furthermore, embodiments of this application also disclose an electronic device, Figure 4 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the stack-based SQL multi-statement parsing method disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0056] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0057] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0058] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the stack-based SQL multi-statement parsing method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.

[0059] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned stack-based SQL multi-statement parsing method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0060] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0061] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0062] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0063] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0064] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A stack-based SQL multi-statement parsing method, characterized in that, include: The input SQL character stream is scanned character by character by a preset lexical analyzer to obtain the current word in the SQL character stream; If the current word is a non-feature word, then the current word is appended to the current preset buffer; If the current word is a feature word, then the word type of the feature word is determined; the non-feature words are other SQL keywords besides the feature word. The feature words are SQL keywords used to control nesting levels; If the word type is the first keyword, the preset lexical analyzer is switched to the semantic lexical state corresponding to the first keyword so as to parse according to the corresponding grammar rules, and according to the occurrence frequency of the current feature word, the corresponding grammar mark is pushed into the current first grammar stack, and the parsed content is appended to the preset buffer. The first keyword is a keyword used to indicate the beginning of a compound statement block; If the word type is the second keyword, the preset lexical analyzer is switched to the semantic lexical state corresponding to the second keyword so as to parse according to the corresponding grammar rules, push the corresponding grammar mark into the current second grammar stack, and append the parsed content to the preset buffer. The second keyword is a keyword used to indicate the beginning of a loop statement block or a conditional statement block; If the word type is a third keyword, the corresponding target syntax stack syntax identifier is popped out, and the content indicating the end of the block is appended to the preset buffer; the third keyword is a keyword used to indicate the end of a compound statement block, a loop statement block, or a conditional statement block; After completing the append operation to the preset buffer, the process jumps to the step of scanning the input SQL character stream character by character using a preset lexical analyzer until a preset termination condition is met, so as to output the current content in the preset buffer and obtain the current statement parsing result.

2. The stack-based SQL multi-statement parsing method according to claim 1, characterized in that, The steps of pushing corresponding grammatical identifiers into the current first grammar stack based on the occurrence frequency of the current feature words, and appending parsed content to the preset buffer, include: If the occurrence count is one, then the compound statement block identifier and the keyword identifier of the first keyword are pushed sequentially into the current first syntax stack; The characters scanned before the first keyword are appended to the preset buffer, and the first keyword is appended to the preset buffer; The characters scanned after the first keyword are rolled back to the SQL character stream, and the corresponding characters in the preset buffer are deleted.

3. The stack-based SQL multi-statement parsing method according to claim 1, characterized in that, The steps of pushing corresponding grammatical identifiers into the current first grammar stack based on the occurrence frequency of the current feature words, and appending parsed content to the preset buffer, include: If the number of occurrences is not one, then push the keyword identifier of the first keyword into the current first syntax stack; The first keyword is appended to the preset buffer; The characters scanned after the first keyword are rolled back to the SQL character stream, and the corresponding characters in the preset buffer are deleted.

4. The stack-based SQL multi-statement parsing method according to claim 1, characterized in that, The steps of pushing the corresponding syntax markers onto the current second syntax stack and appending parsed content to the preset buffer include: Push the keyword identifier of the second keyword onto the current second syntax stack, and append the second keyword to the preset buffer; The characters scanned after the second keyword are rolled back to the SQL character stream, and the corresponding characters in the preset buffer are deleted.

5. The stack-based SQL multi-statement parsing method according to claim 1, characterized in that, The step of popping the syntax identifier of the corresponding target syntax stack and appending the content representing the end of the block to the preset buffer includes: If the third keyword is a keyword used to indicate the end of a compound statement block, then the syntax identifier at the top of the current first syntax stack is popped, and the third keyword is appended to the preset buffer. If the first syntax stack is not empty, then determine the current syntax identifier at the top of the first syntax stack; If the current syntax identifier is a compound statement block identifier, then the compound statement block identifier is popped up; if the current syntax identifier is not a compound statement block identifier, then the characters scanned after the third keyword are rolled back to the SQL character stream, and the corresponding characters in the preset buffer are deleted.

6. The stack-based SQL multi-statement parsing method according to claim 5, characterized in that, After the compound statement block identifier is displayed, the following is also included: If the first syntax stack is detected to be empty, then the preset termination condition is determined to be met.

7. The stack-based SQL multi-statement parsing method according to any one of claims 1 to 6, characterized in that, If the word type is a third keyword, then the corresponding target syntax stack syntax identifier is popped out, and the content at the end of the representation block is appended to the preset buffer, including: If the third keyword is a keyword used to indicate the end of a loop statement block or a conditional statement block, then the syntax identifier at the top of the current second syntax stack is popped, and the third keyword is appended to the preset buffer. If the second syntax stack is not empty, then the characters scanned after the third keyword are rolled back to the SQL character stream, and the corresponding characters in the preset buffer are deleted.

8. A stack-based SQL multi-statement parsing device, characterized in that, include: The character stream scanning module is used to scan the input SQL character stream character by character using a preset lexical analyzer to obtain the current word in the SQL character stream; The type determination module is used to append the current word to the current preset buffer if the current word is a non-feature word. If the current word is a feature word, then the word type of the feature word is determined; the non-feature words are other SQL keywords besides the feature word. The feature words are SQL keywords used to control nesting levels; The first content appending module is used to switch the preset lexical analyzer to the semantic lexical state corresponding to the first keyword if the word type is the first keyword, so as to parse according to the corresponding grammar rules, push the corresponding grammar mark into the current first grammar stack according to the occurrence frequency of the current feature word, and append the parsed content to the preset buffer. The first keyword is a keyword used to indicate the beginning of a compound statement block; The second content appending module is used to switch the preset lexical analyzer to the semantic lexical state corresponding to the second keyword if the word type is the second keyword, so as to parse according to the corresponding grammar rules, push the corresponding grammar mark into the current second grammar stack, and append the parsed content to the preset buffer. The second keyword is a keyword used to indicate the beginning of a loop statement block or a conditional statement block; The third content appending module is used to pop up the syntax identifier of the corresponding target syntax stack if the word type is a third keyword, and append the content indicating the end of the block to the preset buffer; the third keyword is a keyword used to indicate the end of a compound statement block, a loop statement block, or a conditional statement block; The step jump module is used to jump to the step of scanning the input SQL character stream character by character by character through the preset lexical analyzer after the append operation to the preset buffer is completed, until the preset termination condition is met, so as to output the current content in the preset buffer and obtain the current statement parsing result.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the stack-based SQL multi-statement parsing method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store computer programs; wherein, when the computer programs are executed by a processor, they implement the stack-based SQL multi-statement parsing method as described in any one of claims 1 to 7.