Regular expression synthesis method and device based on big language model demand understanding
By combining a large language model with input completion, requirement understanding, and reflective iterative optimization, this regular expression generation method solves the problem of insufficient natural language requirement understanding in existing technologies. It achieves high-precision, high-coverage automated synthesis of regular expressions and has the ability to self-diagnose and progressively optimize.
Patent Information
- Application Number
- CN202511484994.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-10-17
AI Technical Summary
Existing regular expression generation technologies are insufficient in terms of understanding natural language requirements, accuracy, and stability of generated results. In particular, they are difficult to achieve high coverage and accuracy when dealing with complex business rules and high-precision matching requirements, and they lack effective feedback and iteration mechanisms.
By combining large language models for input completion, requirement understanding, and iterative optimization, a multi-strategy local repair method is adopted, including syntax tree splitting, model semantic splitting, reorganization repair, and single-point repair, to generate high-precision and high-coverage regular expressions.
It achieves automated synthesis of regular expressions with high precision and high coverage, significantly reducing the cost of manual intervention, improving the accuracy and robustness of the generated results, and possessing the ability for autonomous diagnosis and progressive optimization.
Smart Images

Figure CN120950416A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of software testing technology, and more specifically to a method and apparatus for synthesizing regular expressions based on large language model requirements understanding. Background Technology
[0002] Regular expressions (Regex) are a highly efficient text pattern matching tool widely used in fields such as data validation, information extraction, log analysis, data cleaning, and network security. Writing traditional regular expressions relies on the developer's professional experience and understanding of the target data patterns. The design and debugging process is often time-consuming and error-prone, especially when dealing with complex business rules, large character sets, or high-precision matching requirements, where the difficulty of manual writing increases significantly.
[0003] Existing automated regular expression generation technologies mainly include template matching, procedural induction, and deep learning-based methods. For example, the RegexGen method based on Domain Specific Language (DSL) can generate regular expressions given positive and negative examples; the Seq2Regex method based on neural networks attempts to directly map natural language descriptions to regular expressions, achieving end-to-end automated generation. However, these methods still have the following shortcomings: insufficient semantic understanding of natural language requirements, resulting in generated results that may not satisfy all positive and negative examples; and a lack of effective feedback iteration mechanisms, making it difficult to progressively correct generation errors.
[0004] When the regular expression is long or the structure is complex, local errors are difficult to locate and fix accurately; when only partial information is provided (such as only natural language description or only regular expression use cases), the generation accuracy drops significantly.
[0005] In recent years, Large Language Models (LLMs) have made significant progress in natural language understanding and code generation. However, directly applying LLMs to regular expression synthesis still faces problems of insufficient accuracy and stability. Therefore, there is an urgent need for a regular expression synthesis method that combines input completion, demand understanding, reflexive mechanisms, and local repair strategies to improve the accuracy, robustness, and adaptability of the generated code. Summary of the Invention
[0006] To address the aforementioned issues, this invention proposes a regular expression synthesis and apparatus based on the understanding of large language model requirements. By combining input completion, understanding and reflective iterative optimization, and multi-strategy local repair techniques, it achieves automated synthesis of high-precision, high-coverage, and iteratively optimizable regular expressions, significantly reducing the cost of manual intervention and improving the accuracy and robustness of the generated results.
[0007] The specific plan is as follows:
[0008] On the one hand, regular expression synthesis methods based on understanding the needs of large language models include:
[0009] S1: Obtain the natural language description and regular expression test case set input by the user; when either the natural language description or the regular expression test case set is missing, call the large language model to complete the missing information, and then call the large language model again to refine the natural language description and regular expression test case set, and output the refined natural language description and regular expression test case set.
[0010] S2: Based on the refined natural language description and regular use case set, construct requirement understanding prompt words, and input them into the large language model to generate multiple candidate understandings. The scorer scores the matching performance of each candidate understanding on the regular use case and selects the optimal understanding.
[0011] S3 inputs the optimally understood and refined natural language description and regular expression test case set into the large language model to generate candidate regular expressions, performs positive and negative example verification on the candidate regular expressions, and outputs the candidate regular expressions and the regular expression test cases that fail.
[0012] S4: Construct reflection prompts based on failed candidate regular expressions and failed regular expression use cases, and input them into a large language model to generate multiple candidate reflections; use a scorer to score the consistency performance of each candidate reflection on failed regular expression use cases, and select the optimal reflection;
[0013] S5 inputs the optimal reflection, the failed candidate regular expressions and the failed regular expression test cases into the large language model for repair, generates the repaired regular expressions, performs positive and negative example verification, and outputs the regular expression test cases that still fail.
[0014] S6. Determine if there are still any regular expression test cases that fail. If so, start the local repair module to split the current regular expression into a syntax tree or model semantics to obtain a sub-pattern list. By replacing each item in the sub-pattern list with a wildcard, determine if the positive example passes, thereby locating the erroneous sub-pattern. Then, use reorganization repair, single-point repair, adjacent window repair, non-adjacent merge block repair, or multi-point independent repair strategies to call the sub-pattern refinement module to accurately repair the local sub-pattern, generate the repaired regular expression, and perform overall verification until all regular expression test cases or the maximum number of repairs is reached, confirming that the regular expression synthesis is complete.
[0015] Furthermore, in S1, the large language model is invoked to refine the natural language description and regular use case set, outputting the refined natural language description and regular use case set. Specifically, the large language model actively identifies and proposes potential boundary conditions, format constraints, and implicit semantic assumptions based on the completed natural language description and regular use case set, generating extended positive and negative examples to eliminate requirement ambiguity and improve the completeness and consistency of input information. It also supports users to add, delete, and modify the completed and generated content in a unified interactive interface, ultimately outputting the refined natural language description and regular use case set.
[0016] Furthermore, in S2, the requirement understanding prompts include at least one of the following: format requirements, field ranges, boundary conditions, or special rules, used to guide the large language model to generate structured understanding.
[0017] Furthermore, in S2, the scorer uses the matching degree of candidate understandings on regular use cases as the evaluation index to score and rank each candidate understanding, and selects the optimal understanding for the current round. If the optimal understanding for the current round still fails to pass all regular use cases, the large language model is driven to iterate and generate a new round of candidate understandings. The scoring and selection process is repeated until the generated understandings satisfy all regular use cases or reach the preset maximum number of iterations, and finally the optimal understanding is output.
[0018] Furthermore, in S4, the reflection prompts include candidate regular expressions that failed, the types of failed test cases, and the regular expression test cases of the failed test cases. The reasons for failure are analyzed through the large language model to obtain the corresponding correction logic. Then, the scorer uses the consistency score between the corresponding correction logic and the failed test cases as the core evaluation index to score and rank each candidate reflection, and select the optimal reflection for the current round. If the optimal reflection for the current round still fails to pass all regular expression test cases, the large language model is driven to iterate and generate a new round of candidate reflections. The evaluation and selection process is repeated until all regular expression test cases are satisfied or the preset maximum number of iterations is reached, and finally the optimal reflection is output.
[0019] Furthermore, in S6, the syntax tree splitting specifically involves: generating an abstract syntax tree through a regular expression parser and extracting each node into an independent sub-pattern; the model semantic splitting specifically involves: calling a large language model to split the regular expression according to semantic boundaries, and verifying whether the recombined sub-patterns after splitting are equivalent to the original regular expression, thereby obtaining a list of sub-patterns.
[0020] Furthermore, in S6, by replacing each item in the subpattern list with a wildcard, it is determined whether the positive example passes, thereby locating the erroneous subpattern. Specifically, this includes: splitting the candidate regular expression to obtain the subpattern list; replacing one or more subpatterns in the subpattern list with wildcards to form a temporary regular expression; if the temporary regular expression can pass all the regular positive examples, it is determined that there is a repairable erroneous subpattern or erroneous combination in the replaced area.
[0021] Furthermore, in S6, the recombination repair includes sequential rearrangement, enumerated deletion, and enumerated insertion;
[0022] The specific process of the order rearrangement is as follows: perform full permutation and combination of the sub-patterns, and verify whether the new sequence passes all regular expression test cases;
[0023] The enumeration deletion specifically involves: attempting to delete each sub-pattern and verifying whether the remaining combinations pass all regular expression test cases; if they do, the deleted sub-pattern is determined to be redundant.
[0024] The enumeration insertion specifically involves: enumerating and inserting specified wildcards between sub-patterns, and calling single-point or multi-point repair modules to refine the wildcard region.
[0025] On the other hand, a regular expression synthesis device based on large language model requirement understanding includes:
[0026] The refinement module is used to obtain the natural language description and regular expression test case set input by the user. When either the natural language description or the regular expression test case set is missing, the large language model is called to complete the missing information. Then, the large language model is called again to refine the natural language description and the regular expression test case set, and the refined natural language description and regular expression test case set are output.
[0027] The optimal understanding filtering module is used to construct requirement understanding prompt words based on refined natural language descriptions and regular use case sets, and input them into a large language model to generate multiple candidate understandings. The scoring unit scores the matching performance of each candidate understanding on regular use cases and filters out the optimal understanding.
[0028] The validation module is used to input the optimally understood and refined natural language description and regular expression test case set into the large language model, generate candidate regular expressions, perform positive and negative example validation on the candidate regular expressions, and output the candidate regular expressions and the regular expression test cases that fail.
[0029] The optimal reflection filtering module is used to construct reflection prompt words based on the failed candidate regular expressions and failed regular expression use cases, and input them into the large language model to generate multiple candidate reflections; the scorer scores the consistency performance of each candidate reflection on the failed regular expression use cases, and filters out the optimal reflection.
[0030] The repair module is used to input the best reflection, the failed candidate regular expressions and the failed regular expression test cases into the large language model for repair, generate the repaired regular expressions, perform positive and negative example verification, and output the regular expression test cases that still fail.
[0031] The synthesis module is used to determine if there are still any regular expression test cases that fail. If so, the local repair module is activated to split the current regular expression into a syntax tree or model semantics to obtain a sub-pattern list. By replacing each item in the sub-pattern list with a wildcard, it is determined whether the positive example passes, thereby locating the erroneous sub-pattern. Then, it uses reorganization repair, single-point repair, adjacent window repair, non-adjacent merge block repair, or multi-point independent repair strategies to call the sub-pattern refinement module to accurately repair the local sub-pattern, generate the repaired regular expression, and perform overall verification until all regular expression test cases or the maximum number of repairs is reached, confirming that the regular expression synthesis is complete.
[0032] The present invention adopts the above technical solution and has the following beneficial effects:
[0033] (1) This invention realizes the ability of a large language model to extract structured semantic rules from fuzzy natural language descriptions by constructing requirement understanding prompts that include format requirements, field ranges and boundary conditions, and by using a scorer to quantify and filter the matching degree of multiple candidate understandings on regular use cases.
[0034] (2) This invention obtains a list of independently verifiable sub-patterns by generating an abstract syntax tree using a regular syntax parser and splitting the syntax tree, or by calling a large language model to perform equivalence verification model semantic splitting according to semantic boundaries. After locating erroneous sub-patterns by combining regular positive examples, strategies such as reorganization repair, single-point repair or multi-point independent repair are used to call the sub-pattern refinement module for local correction, thereby achieving accurate location and safe repair of local errors in complex regular expressions.
[0035] (3) This invention constructs reflection prompt words based on the failed candidate regular expressions and their failed test cases, drives the large language model to analyze the cause of the error and generate correction logic, and then sorts and filters the corrected logic and the failed test cases by the scorer with the consistency of the correction logic and the failed test cases as the core indicator, forming a multi-round iterative reflection-optimization closed loop, realizing the system's autonomous diagnosis and incremental optimization capabilities after the first generation failure, effectively dealing with the challenges of regular expression synthesis in complex situations such as boundary scenarios and implicit constraints. Attached Figure Description
[0036] Figure 1 This is a flowchart of the regular expression synthesis method based on large language model requirement understanding in an embodiment of the present invention;
[0037] Figure 2This is a schematic diagram of the overall structure of an embodiment of the present invention;
[0038] Figure 3 This is a schematic diagram illustrating the structure for understanding and reflection on embodiments of the present invention;
[0039] Figure 4 This is a diagram of a regular expression synthesis device based on large language model requirement understanding, according to an embodiment of the present invention. Detailed Implementation
[0040] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0041] like Figure 1 As shown, the present invention relates to regular expression synthesis based on large language model requirements understanding, comprising:
[0042] S1: Obtain the natural language description and regular expression test case set input by the user; when either the natural language description or the regular expression test case set is missing, call the large language model to complete the missing information, and then call the large language model again to refine the natural language description and regular expression test case set, and output the refined natural language description and regular expression test case set.
[0043] Specifically, the large language model is invoked to refine the natural language description and regular expression test case set, outputting a refined natural language description and regular expression test case set. This includes: the large language model actively identifying and proposing potential boundary conditions, format constraints, and implicit semantic assumptions based on the completed natural language description and regular expression test case set, generating extended positive and negative examples to eliminate requirement ambiguity and improve the completeness and consistency of input information; and supporting users to add, delete, and modify the completed and generated content in a unified interactive interface, ultimately outputting the refined natural language description and regular expression test case set.
[0044] Specifically, the system obtains the initial natural language description and / or regular expression test cases input by the user. For example, if the user inputs the initial natural language description: "Extract dates in the form of 2023-08-14, but do not match illegal dates such as 'abc' or '99 / 99 / 9999'", the system detects a missing set of regular expression test cases and calls the Large Language Model (LLM) to generate an initial set of test case suggestions (e.g., positive examples: "2023-08-14", "1902-11-01"; negative examples: "abc", "99 / 99 / 9999", "2023-13-45", etc.), and displays them back to the user for adjustment and confirmation. After user confirmation, the system sends the final description and test cases back to the LLM to generate a more accurate natural language description and more comprehensive regular expression test cases, and finally displays them back to the user for a second review and lock.
[0045] S2 constructs requirement understanding prompts based on refined natural language descriptions and regular use case sets, and inputs them into a large language model to generate multiple candidate understandings. A scorer scores the matching performance of each candidate understanding on regular use cases and selects the optimal understanding.
[0046] Specifically, the requirement understanding prompts include at least one of the following: format requirements, field ranges, boundary conditions, or special rules, used to guide the large language model to generate structured understanding.
[0047] Specifically, the scorer uses the matching degree of candidate understandings on regular use cases as the evaluation metric to score and rank each candidate understanding, and selects the best understanding for the current round. If the best understanding for the current round still fails to pass all regular use cases, the large language model is driven to iterate and generate a new round of candidate understandings. The scoring and selection process is repeated until the generated understandings satisfy all regular use cases or reach the preset maximum number of iterations, and finally the best understanding is output.
[0048] Specifically, the understanding prompt is generated by concatenating natural language descriptions with positive and negative examples to form the requirement understanding prompt word Prompt0. This requires the large model to output an understanding that includes the constraints of regular expressions, such as format requirements, field ranges, leap year rules, and boundary conditions. Candidate understanding generation and scoring involve the LLM outputting N candidate understandings {U_i}. The scorer performs consistency and coverage scoring on {U_i} on the test case set. If the highest score does not reach a threshold, iterative prompts (including error guidance and missing points) are provided until the best understanding U0 is obtained or the maximum number of iterations is reached. Then, understanding U0 is combined with the task input to generate candidate regular expressions {R_i}, which are quickly validated on the test case set, and the regular expression with the highest score, R0, is retained.
[0049] S3 inputs the optimally understood and refined natural language description and regular expression test case set into the large language model, generates candidate regular expressions, performs positive and negative example verification on the candidate regular expressions, and outputs the candidate regular expressions and the regular expression test cases that fail.
[0050] S4: Construct reflection prompts based on failed candidate regular expressions and failed regular expression use cases, input them into a large language model to generate multiple candidate reflections; use a scorer to score the consistency performance of each candidate reflection on failed regular expression use cases, and select the optimal reflection.
[0051] Specifically, the reflection prompts include candidate regular expressions that failed, the types of failed test cases, and the regular expression test cases of the failed test cases. The reasons for failure are analyzed through the large language model to obtain the corresponding correction logic. Then, the scorer uses the consistency score between the corresponding correction logic and the failed test cases as the core evaluation index to score and rank each candidate reflection, and select the optimal reflection for the current round. If the optimal reflection for the current round still fails to pass all regular expression test cases, the large language model is driven to iterate and generate a new round of candidate reflections. The evaluation and selection process is repeated until all regular expression test cases are satisfied or the preset maximum number of iterations is reached, and finally the optimal reflection is output.
[0052] Specifically, during the reflection process, if there are still unsuccessful samples in R0, the system constructs a reflection prompt Prompt1 based on the failed samples, requiring the large model output to include the constraint format that regular expressions should have (such as field range, leap year rules, boundary conditions, etc.). Then, the scorer performs consistency scoring and coverage scoring on {I_i} on the set of failed test cases. If the highest score does not reach the threshold, iterative prompts (including error guidance and omissions) are given until the best reflection I is obtained or the maximum number of iterations is reached.
[0053] S5 inputs the optimal reflection, the current regular expression, and the failed regular expression test cases into the large language model for repair, generates the repaired regular expression, performs positive and negative example verification, and outputs the regular expression test cases that still fail.
[0054] S6. Determine if there are still any regular expression test cases that fail. If so, start the local repair module to split the current regular expression into a syntax tree or model semantics to obtain a sub-pattern list. By replacing each item in the sub-pattern list with a wildcard, determine if the positive example passes, thereby locating the erroneous sub-pattern. Then, use reorganization repair, single-point repair, adjacent window repair, non-adjacent merge block repair, or multi-point independent repair strategies to call the sub-pattern refinement module to accurately repair the local sub-pattern, generate the repaired regular expression, and perform overall verification until all regular expression test cases or the maximum number of repairs is reached, confirming that the regular expression synthesis is complete.
[0055] Specifically, the syntax tree splitting involves generating an abstract syntax tree using a regular expression parser and extracting each node into an independent sub-pattern; the model semantic splitting involves calling a large language model to split the regular expression according to semantic boundaries and verifying whether the recombined sub-patterns are equivalent to the original regular expression to obtain a list of sub-patterns; the recombining and repair includes sequential rearrangement, enumeration deletion, and enumeration insertion.
[0056] The specific process of the order rearrangement is as follows: perform full permutation and combination of the sub-patterns, and verify whether the new sequence passes all regular expression test cases;
[0057] The enumeration deletion specifically involves: attempting to delete each sub-pattern sequentially, and verifying whether the remaining combinations pass all regular expression test cases; if they do, the deleted sub-pattern is determined to be redundant.
[0058] Specifically, in the local repair phase, this embodiment employs multiple refined strategies to precisely repair each sub-pattern. First, in the reorder_repair phase, the system optimizes the structure in three ways: First, sequential reordering, which involves performing full permutations of the sub-patterns of the regular expression to be repaired, generating a new sequence, and verifying positive and negative examples. If all positive examples are completely matched and all negative examples are excluded, the repair is complete. Second, enumerated deletion, where the system sequentially removes each potentially redundant sub-pattern, recombines the remaining parts, and verifies them. If deleting a sub-pattern allows the test to pass, it is considered redundant and is removed. Third, enumerated insertion, where the system sequentially inserts a general capture pattern (?P) into each potential location where a key sub-pattern might be missing. <lang>After verifying the possibility of repair in the regular expression positive example, the single-point or multi-point repair module is called to further refine the wildcard region.
[0059] Secondly, single-point repair precisely corrects the area of the regular expression to be repaired, replacing the sub-pattern with (?P <lang>After `.*)`, the subpattern refinement module `repair2one` is called. This module receives the currently replaced subpattern, a list of all positive and negative example blocks, and failed test cases. It uses a large language model to generate candidate repair solutions and iteratively verifies them to ensure that the repaired subpattern matches all positive example blocks and excludes all negative example blocks, until success or the maximum number of attempts is reached. When a single-point repair fails, the system starts single-point adjacent window repair (`adjacent_window_repair`), which expands several adjacent subpatterns to the left and right of the error location to form a joint repair block, replacing the entire block with `(?P`. <lang>The first part (.*) is then handed over to the repair2one module for overall refinement, suitable for handling problems caused by coordination errors of multiple adjacent sub-patterns. To further address complex cross-region errors, the system introduces non-adjacent merged block repair (merged_block_repair). This selects the beginning and end positions of multiple discontinuous but semantically related erroneous sub-patterns as boundaries, merging all sub-patterns within the interval into a single repair block, uniformly replacing them with (?P). <lang>After `.*)`, `repair2one` is invoked for integrated repair, effectively resolving scattered but logically dependent errors and improving the consistency and efficiency of the repair process. Furthermore, non-adjacent multi-repair independently replaces and repairs multiple discontinuous error sub-patterns independently. The correctness of each sub-pattern is guaranteed by the `repair2one` module, avoiding conflicts between positive and negative examples during the repair process. Finally, all repaired sub-patterns are recombined to form a complete regular expression. Through this multi-layered, multi-strategy collaborative local repair mechanism, the system achieves high-precision location and safe correction of various structural errors in complex regular expressions, significantly improving the repair success rate and generation quality.
[0060] Specifically, the local repair (single-point repair) system splits the regular expression to be repaired into three segments (e.g., a date format of "year-month-day"): Syntax tree splitting: Taking the regular expression as input, a syntax parser parses it into an abstract syntax tree, including character nodes, character class nodes, connection nodes, selection nodes, repetition nodes, grouping nodes, and quantifier nodes; a depth-first traversal is performed on the abstract syntax tree to extract independent sub-patterns node by node; the resulting independent sub-patterns are output sequentially to form a list of split sub-patterns. Model semantic splitting: The LLM is required to split the regular expression according to semantic boundaries, and after splitting, the sub-pattern lists need to be merged and verified to be consistent with those before splitting. The subpattern recombination and repair process involves permuting and combining the subpatterns of the regular expression to be repaired to form new subpattern sequences. For potentially redundant subpatterns, the system removes each subpattern sequentially and then permutes and combines the remaining subpatterns to form new subpattern sequences. All new subpattern sequences are merged and verified against positive and negative examples. If the rearranged sequences match all positive examples and exclude negative examples, the repair is complete. Using the subpattern list as input, with a list length of n, there are n+1 possible insertion positions, where the 0th position is before the first subpattern, the i-th position is after the i-th subpattern, and 1 ≤ i ≤ n. For positions where subpatterns may be missing, the system sequentially inserts a general capture pattern (?P) into each possible position. <lang>The candidate regular expressions are output sequentially to form a complete set of regular expressions after wildcard insertion; and the single-point repair or multi-point repair module is called to refine the wildcard region; the single-point replacement and verification part temporarily replaces p_i with wildcard capture (?P) one by one. <lang>The temporary expression R_tmp is formed using .*), and its repairability is verified on the set of positive test cases. The repair2one refinement part calls the subpattern refinement module repair2one for p_i that is determined to be repairable: input the current subpattern p_i, the list of positive and negative example capture values, and the list of failed sample capture values; construct a prompt containing "no matching positive example / mismatched negative example", and call LLM to produce the repair subpattern p_i'; iterate p_i' until it passes or reaches the maximum number of iterations; output the compliant repair subpattern p_i'; the backfill and overall verification part uses p_i' to backfill the original position and then performs overall verification; if it passes, the process ends, otherwise, multi-point repair is performed.
[0061] Specifically, during the local repair phase, the system implements a multi-strategy collaborative fine-grained repair mechanism for different types of regular expression errors. When a single-point repair fails and the error involves joint logic problems between adjacent sub-patterns, the system initiates single-point adjacent window repair (adjacent_window_repair):
[0062] First, locate the subpattern that mismatches on the negative example (e.g., when the regular expression R2 = (?:https?: / / )?(?:www\.)?([a-z0-9-]+)(\.[az]{2,})+ mismatches on the negative example "https: / / 123-.com", locate ([a-z0-9-]+) as the suspected error area). Using this location as the center, expand to the left and right sides by a preset window width w to form a repair block containing the context (e.g., (?:www\.)?([a-z0-9-]+)(\.[az]{2,})). Replace the entire repair block with the wildcard capture mode (?P). <lang>*) Construct a temporary regular expression R_tmp for repairability verification; if the verification passes, use the original repair block text as input to call the subpattern refinement module repair2one to generate the repaired subpattern B', and then fill it back into the original expression for overall verification. If the repair fails, the system will adaptively adjust the window size w and try asymmetric expansion strategies (such as expanding only to the left or only to the right) until the repair is successful or the next repair strategy is adopted. For cases where errors are distributed in multiple discontinuous positions but have semantic dependencies, the system uses non-adjacent merged block repair (merged_block_repair): select the minimum and maximum indices among all error positions as the start and end boundaries of the repair interval, merge all subpatterns within the interval into a continuous overall repair block, and use (?P <lang>After replacing the .*) block, verify its repairability; if the conditions are met, input the original text of the merged block into the repair2one module for integrated fine-tuning, and then backfill and verify the repair results. This strategy effectively solves the logical coupling problem between cross-region errors by using the "start-end encirclement, middle-segment merging" method, avoiding mutual interference caused by multiple independent repairs. For multiple discontinuous and independent error sub-patterns (such as "(cat|dog)s?" and "pet" in R4 = (cat|dog)s?|pet), the system performs non-adjacent multi-point independent repair (non_adjacent_multi_repair): each error sub-pattern is replaced with an independent wildcard (?P)<lang_k> The process begins by verifying whether the replaced R_tmp still matches all positive examples. Then, the repair2one module is called in parallel to independently refine each replaced sub-pattern, obtaining a repaired sub-pattern set {p_k'}. Finally, {p_k'} is backfilled and combined into a complete regular expression in its original order and validated overall until all test cases pass or the maximum number of iterations is reached. The core execution unit of the above repair process is the sub-pattern refinement module repair2one, which receives the sub-pattern to be repaired (sub_re), the capture value set of positive and negative examples, and a list of failed regular expression test cases. During initialization, it first verifies whether the current sub-pattern meets the matching requirements; if so, it returns directly. Otherwise, it constructs a hint word containing unmatched positive examples, mismatched negative examples, and contextual constraints, calls the large language model to generate a repair candidate sub_re', and continuously optimizes it through iterative verification until the generated result correctly matches all positive examples and excludes all negative examples, or the maximum number of attempts is reached. For complex regular expressions containing explicit branching structures "|", the system uses the branch refactoring module (branch_refactor_repair) for specialized processing: First, the outermost "|" operator is identified by the syntax parser, splitting the expression into multiple independent branches {b1,b2,…,bn}; then, full positive and negative example verification is performed on each branch bi, calculating the set of matching positive examples P(bi) and the set of mismatched negative examples N(bi), and marking negative pollution branches (N(bi) is not empty) and invalid branches (P(bi) is empty) accordingly; the top k valid branches that are neither pollution nor invalid and cover the most positive examples are retained (k is determined by a preset threshold or heuristic rule); for uncovered positive examples and all negative examples, sub-problems are generated and the branch splitting and verification process is recursively executed; finally, all valid branches are deduplicated and merged, reassembled into a new branch structure using "|", and returned to the upper-level strategy for backfilling and overall verification.Through the collaborative work of the aforementioned multi-granularity, multi-path local repair mechanisms, this invention can automatically generate high-precision, high-coverage regular expressions starting from the natural language descriptions and regular expression use cases input by the user. Furthermore, by using reflection guidance and multi-level repair strategies, it ensures the correctness and robustness of these expressions in positive and negative example verification, significantly reducing the cost of manual intervention and improving the automation level and overall efficiency of regular expression generation and optimization.
[0063] Specifically, Figure 2 This embodiment illustrates the overall structure of a regular expression synthesis method based on large language model (LLM) requirement understanding: First, the user inputs a natural language description and a set of regular expression use cases. The LLM then completes the missing information and refines the input. Next, an initial regular expression is generated based on the refined input and validated. If it fails, requirement understanding prompts are constructed to guide the LLM to generate multiple candidate understandings, and a scorer is used to select the optimal understanding. Subsequently, a new regular expression is generated using the optimal understanding and validated again. If it still fails, reflection prompts are constructed to analyze the reasons for failure, generate reflections, and further optimize the regular expression. Finally, the regular expression use cases that still fail are partially adjusted, such as splitting into sub-patterns, reorganizing and repairing, single-point repair, or multi-point repair, until all use cases pass or the maximum number of repairs is reached, thereby achieving accurate and efficient regular expression synthesis.
[0064] Specifically, Figure 3 This demonstrates the iterative optimization process of understanding and reflection in regular expression generation. The "Understanding" section on the left starts with natural language descriptions and regular expression use cases, generates and stores understanding through a large model, and then verifies it using a scorer. If it fails, feedback is provided to optimize the understanding until the iteration termination condition is met or the optimal understanding is generated, and then a new regular expression is generated for further verification. The "Reflection" section on the right addresses regular expressions that fail verification by evaluating the above conditions to determine whether reflection is needed. If so, the reflection is recorded and generated, and a new regular expression is generated for verification until it passes, at which point the final regular expression is output. Otherwise, feedback continues to optimize the reflection, forming a closed-loop iterative optimization mechanism.
[0065] Specifically, in this embodiment, the method is applied in a concrete way. First, the system obtains the natural language description and / or regular expression test cases input by the user. When either of these pieces of information is missing, the system immediately calls the large language model to complete the description according to the context and displays it back, allowing the user to add, delete, or modify it within a unified interface. After user confirmation, the system sends the description and test cases back into the large model, which refines a more accurate natural language description and expands it with a wider range of regular expression test cases. Finally, the system displays the results back to the user for secondary review and locks them, before proceeding to the next step. Secondly, requirement understanding prompts are constructed based on natural language descriptions and regular expression use cases, and then input into a large language model to generate structured understanding. A scorer filters for the optimal understanding until all use cases are satisfied or the maximum number of iterations is reached. Next, the optimal understanding and requirements are input into the large language model to generate regular expressions. Positive and negative examples are used to validate the generated results, and failed use cases are output. Then, reflection prompts are constructed based on the failed use cases, and input into the large language model to generate reflections. A scorer filters for the optimal reflections until all failed use cases are satisfied or the maximum number of iterations is reached. Next, the optimal reflections, along with erroneous regular expressions and failed regular expression use cases, are input into the large language model to require regular expression repair. Positive and negative examples are used to validate the repair results, and failed use cases are output. If failed use cases still exist, the reflection and repair process can be repeated, or a partial repair phase can be initiated. Finally, when the generated regular expression still fails to pass all regular expression use cases, the system initiates a partial repair module. The regular expression to be repaired is first split by syntax tree or semantically split based on the large language model. Then, the partial repair phase is performed.
[0066] Specifically, in this specification, "understanding and reflection" refers to the structured constraint descriptions (such as field ranges, formats, boundary conditions, etc.) for the target matching task; "positive and negative example use cases" refer to the sets of strings expected to match / not match, respectively; "sub-pattern" refers to the small functional units or logical blocks after semantic segmentation of the regular expression; and "scorer" refers to the evaluation module used to measure the performance of candidate understanding and reflection on the use case set. At each stage of this invention, unless otherwise stated, the passing criteria for regular expressions or sub-patterns shall at least satisfy: (a) all positive examples should be "fully matched"; and (b) all negative examples should be "fully rejected".
[0067] In summary, this invention, by introducing input completion and requirement optimization mechanisms, can automatically complete and refine input information using a large language model when users only provide incomplete natural language descriptions or missing regular expression test cases. This generates structured requirement descriptions and regular expression test case sets, and, combined with echo display and user interaction confirmation, ensures the completeness and accuracy of the input content. Compared to the traditional method that relies entirely on manual writing and repeated adjustments, this significantly lowers the user's barrier to entry and technical requirements, improves the initial input quality, and lays a reliable foundation for subsequent high-precision regular expression generation. Based on this, a two-layer optimization mechanism combining understanding and reflection is proposed: in the generation stage, requirement understanding prompts containing elements such as format constraints, field ranges, and boundary conditions are constructed to guide the large language model in generating semantically clear and structurally sound candidate regular expressions; after verification failure, reflection is built based on the failed test cases, driving the model to analyze the causes of errors and generate corrective logic. This process uses a scorer to quantitatively evaluate and iteratively screen candidate reflections, forming a closed-loop process of "generation-verification-reflection-optimization." This achieves continuous evolution and self-correction capabilities for regular expressions, effectively overcoming the problems of low accuracy and difficulty in convergence in one-time generation, and significantly reducing the cost of manual intervention and repeated debugging. A multi-strategy collaborative local repair method is designed, covering various refined repair strategies such as reorganization repair, single-point repair, adjacent window repair, non-adjacent merged block repair, non-adjacent multi-point independent repair, and branch reconstruction. The system first splits the regular expression to be repaired into sub-pattern units, accurately locates the error area, and adaptively selects the optimal repair path according to the error type, adjusting only the local area while preserving the correct structure. Among them, the repair2one module is responsible for performing fine-grained sub-pattern refinement, and the branch_refactor module is specifically responsible for handling complex logic containing "|" branch structures. The two work together to achieve comprehensive coverage from simple patterns to complex structures. This local repair mechanism avoids the uncertainty brought by overall rewriting, significantly improves repair efficiency, matching accuracy, and negative example exclusion capability, and ultimately achieves the goal of high-precision, low-perturbation, and iterative automated regular expression synthesis.
[0068] like Figure 4 As shown, this embodiment also discloses a regular expression synthesis device based on large language model requirement understanding, including:
[0069] The refinement module 41 is used to obtain the natural language description and regular expression test case set input by the user. When either the natural language description or the regular expression test case set is missing, the large language model is called to complete the missing information, and then the large language model is called again to refine the natural language description and regular expression test case set, and the refined natural language description and regular expression test case set are output.
[0070] The optimal understanding filtering module 42 is used to construct requirement understanding prompt words based on the refined natural language description and regular use case set, and input them into the large language model to generate multiple candidate understandings. The scoring unit scores the matching performance of each candidate understanding on the regular use case and filters out the optimal understanding.
[0071] The verification module 43 is used to input the optimally understood and refined natural language description and regular use case set into the large language model, generate candidate regular expressions, perform positive and negative example verification on the candidate regular expressions, and output the candidate regular expressions and the regular use cases that fail.
[0072] The optimal reflection filtering module 44 is used to construct reflection prompt words based on the failed candidate regular expressions and failed regular expression use cases. It inputs the large language model to generate multiple candidate reflections. The scorer scores the consistency performance of each candidate reflection on the failed regular expression use cases and filters out the optimal reflection.
[0073] Repair module 45 is used to input the optimal reflection, the current regular expression and the failed regular expression test cases into the large language model for repair, generate the repaired regular expression, perform positive and negative example verification, and output the regular expression test cases that still fail.
[0074] The synthesis module 46 is used to determine whether there are still regular expression test cases that fail. If so, the local repair module is started to split the current regular expression into a syntax tree or a model semantics to obtain a sub-pattern list. Based on the sub-pattern list, the erroneous sub-pattern is located and repaired using strategies such as reorganization repair, single-point repair, adjacent window repair, non-adjacent merge block repair, or multi-point independent repair. The sub-pattern refinement module is called to accurately repair the local sub-pattern, generate the repaired regular expression, and perform overall verification until all regular expression test cases pass or the maximum number of repairs is reached, confirming that the regular expression synthesis is complete.
[0075] The specific implementation of the regular expression synthesis system based on large language model requirement understanding is the same as the regular expression synthesis method based on large language model requirement understanding, and will not be described again in this embodiment.
[0076] Although the invention has been specifically shown and described in conjunction with preferred embodiments, those skilled in the art should understand that various changes in form and detail may be made to the invention without departing from the spirit and scope of the invention as defined in the appended claims, all of which shall be within the scope of protection of the invention.< / lang> < / lang> < / lang> < / lang> < / lang> < / lang> < / lang> < / lang>
Claims
1. A regular expression synthesis method based on large language model requirement understanding, characterized in that, Includes the following steps: S1: Obtain the natural language description and regular expression test case set input by the user; when either the natural language description or the regular expression test case set is missing, call the large language model to complete the missing information, and then call the large language model again to refine the natural language description and regular expression test case set, and output the refined natural language description and regular expression test case set. S2: Based on the refined natural language description and regular use case set, construct requirement understanding prompt words, and input them into the large language model to generate multiple candidate understandings. The scorer scores the matching performance of each candidate understanding on the regular use case and selects the optimal understanding. S3 inputs the optimally understood and refined natural language description and regular expression test case set into the large language model to generate candidate regular expressions, performs positive and negative example verification on the candidate regular expressions, and outputs the candidate regular expressions and the regular expression test cases that fail. S4: Construct reflection prompts based on failed candidate regular expressions and failed regular expression use cases, and input them into a large language model to generate multiple candidate reflections; use a scorer to score the consistency performance of each candidate reflection on failed regular expression use cases, and select the optimal reflection; S5 inputs the optimal reflection, the failed candidate regular expressions and the failed regular expression test cases into the large language model for repair, generates the repaired regular expressions, performs positive and negative example verification, and outputs the regular expression test cases that still fail. S6. Determine if there are still any regular expression test cases that fail. If so, start the local repair module to split the current regular expression into a syntax tree or model semantics to obtain a sub-pattern list. By replacing each item in the sub-pattern list with a wildcard, determine if the positive example passes, thereby locating the erroneous sub-pattern. Then, use reorganization repair, single-point repair, adjacent window repair, non-adjacent merge block repair, or multi-point independent repair strategies to call the sub-pattern refinement module to accurately repair the local sub-pattern, generate the repaired regular expression, and perform overall verification until all regular expression test cases or the maximum number of repairs is reached, confirming that the regular expression synthesis is complete.
2. The regular expression synthesis method based on large language model requirement understanding according to claim 1, characterized in that, In S1, the large language model is invoked to refine the natural language description and regular expression test case set, outputting the refined natural language description and regular expression test case set. Specifically, the large language model actively identifies and proposes potential boundary conditions, format constraints, and implicit semantic assumptions based on the completed natural language description and regular expression test case set, generating extended positive and negative examples to eliminate requirement ambiguity and improve the completeness and consistency of input information; and supports users to add, delete, and modify the completed and generated content in a unified interactive interface, finally outputting the refined natural language description and regular expression test case set.
3. The regular expression synthesis method based on large language model requirement understanding according to claim 1, characterized in that, In S2, the requirement understanding prompts include at least one of the following: format requirements, field ranges, boundary conditions, or special rules, which are used to guide the large language model to generate structured understanding.
4. The regular expression synthesis method based on large language model requirement understanding according to claim 1, characterized in that, In S2, the scorer uses the matching degree of candidate understandings on regular use cases as the evaluation index to score and rank each candidate understanding, and selects the best understanding in the current round. If the best understanding in the current round still fails to pass all regular use cases, the large language model is driven to iterate and generate a new round of candidate understandings. The scoring and selection process is repeated until the generated understandings meet all regular use cases or reach the preset maximum number of iterations, and finally the best understanding is output.
5. The regular expression synthesis method based on large language model requirement understanding according to claim 1, characterized in that, In S4, the reflection prompts include the candidate regular expressions that failed, the types of the failed test cases, and the regular expression test cases of the failed test cases. The reasons for failure are analyzed through the large language model to obtain the corresponding correction logic. Then, the scorer uses the consistency score between the corresponding correction logic and the failed test cases as the core evaluation index to score and rank each candidate reflection, and select the optimal reflection for the current round. If the optimal reflection for the current round still fails to pass all regular expression test cases, the large language model is driven to iterate and generate a new round of candidate reflections. The evaluation and selection process is repeated until all regular expression test cases are satisfied or the preset maximum number of iterations is reached, and finally the optimal reflection is output.
6. The regular expression synthesis method based on large language model requirement understanding according to claim 1, characterized in that, In S6, the syntax tree splitting specifically involves generating an abstract syntax tree through a regular expression parser and extracting each node into an independent sub-pattern; the model semantic splitting specifically involves calling a large language model to split the regular expression according to semantic boundaries, and verifying whether the recombined sub-patterns after splitting are equivalent to the original regular expression, thereby obtaining a list of sub-patterns.
7. The regular expression synthesis method based on large language model requirement understanding according to claim 1, characterized in that, In S6, by replacing each item in the subpattern list with a wildcard, it is determined whether the positive examples pass, thereby locating the erroneous subpatterns. Specifically, this includes: splitting the candidate regular expression to obtain the subpattern list; replacing one or more subpatterns in the subpattern list with wildcards to form a temporary regular expression; if the temporary regular expression can pass all the regular positive examples, it is determined that there are repairable erroneous subpatterns or incorrect combinations in the replaced area.
8. The regular expression synthesis method based on large language model requirement understanding according to claim 1, characterized in that, In S6, the recombination repair includes sequential rearrangement, enumerated deletion, and enumerated insertion; The specific process of the order rearrangement is as follows: perform full permutation and combination of the sub-patterns, and verify whether the new sequence passes all regular expression test cases; The enumeration deletion specifically involves: attempting to delete each sub-pattern and verifying whether the remaining combinations pass all regular expression test cases; if they do, the deleted sub-pattern is determined to be redundant. The enumeration insertion specifically involves: enumerating and inserting specified wildcards between sub-patterns, and calling single-point or multi-point repair modules to refine the wildcard region.
9. A regular expression synthesis device based on large language model requirement understanding, characterized in that, include: The refinement module is used to obtain the natural language description and regular expression test case set of user input; When either the natural language description or the set of regular expression test cases is missing, the large language model is called to complete the missing information. Then, the large language model is called again to refine the natural language description and the set of regular expression test cases, and the refined natural language description and the set of regular expression test cases are output. The optimal understanding filtering module is used to construct requirement understanding prompt words based on refined natural language descriptions and regular use case sets, and input them into a large language model to generate multiple candidate understandings. The scoring unit scores the matching performance of each candidate understanding on regular use cases and filters out the optimal understanding. The validation module is used to input the optimally understood and refined natural language description and regular expression test case set into the large language model, generate candidate regular expressions, perform positive and negative example validation on the candidate regular expressions, and output the candidate regular expressions and the regular expression test cases that fail. The optimal reflection filtering module is used to construct reflection prompt words based on the failed candidate regular expressions and failed regular expression use cases, and input them into the large language model to generate multiple candidate reflections; the scorer scores the consistency performance of each candidate reflection on the failed regular expression use cases, and filters out the optimal reflection. The repair module is used to input the best reflection, the failed candidate regular expressions and the failed regular expression test cases into the large language model for repair, generate the repaired regular expressions, perform positive and negative example verification, and output the regular expression test cases that still fail. The synthesis module is used to determine if there are still any regular expression test cases that fail. If so, the local repair module is activated to split the current regular expression into a syntax tree or model semantics to obtain a sub-pattern list. By replacing each item in the sub-pattern list with a wildcard, it is determined whether the positive example passes, thereby locating the erroneous sub-pattern. Then, it uses reorganization repair, single-point repair, adjacent window repair, non-adjacent merge block repair, or multi-point independent repair strategies to call the sub-pattern refinement module to accurately repair the local sub-pattern, generate the repaired regular expression, and perform overall verification until all regular expression test cases or the maximum number of repairs is reached, confirming that the regular expression synthesis is complete.
Citation Information
Patent Citations
Text classification method of regular expression generated based on large language model
CN117556049A
Auxiliary teaching interaction method and device for regular language and automaton theory
CN118627940A
Medical document regular expression generation method and device based on large language model
CN118709661A
Data set keyword generation and screening method based on large language model
CN119474339A
SQL (Structured Query Language) statement generation method based on large language model, medium and equipment
CN119862200A