Large model driven time series database user-defined function test method and system

CN122019389BActive Publication Date: 2026-09-25CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610167586.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-02-05
Publication Date
2026-09-25
Estimated Expiration
2046-02-05

AI Technical Summary

Technical Problem

LLM在预训练阶段可能未涵盖特定时序数据库的API文档、UDF最佳实践和编程模式,导致生成用例不准确或无效

Benefits of technology

[0058]第五方面,本申请提供一种计算机程序产品,包括计算机程序,该计算机程序被处理器执行时实现如上述第一方面及第一方面各种可能的实现方式所述的大模型驱动的时序数据库用户自定义函数测试方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019389B_ABST
    Figure CN122019389B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a large model driven time series database user defined function test method and system. Applied to the software test technical field, the method extracts the domain knowledge related to the user defined function from the pre-constructed time series database knowledge base by using a retrieval enhancement generation strategy, and constructs a static analysis instruction based on the domain knowledge; the source code of the user defined function and the static analysis instruction are input into a first large language model for analysis and processing, and a static analysis result is obtained; the time sequence characteristics of the test data are determined according to the static analysis result, and the test data meeting the time sequence characteristics are generated by using a second large language model after fine tuning; test cases are constructed according to the test data, and the user defined function is tested according to a preset execution strategy by using the test cases, and a dynamic test result is obtained. The method improves the efficiency and accuracy of the user defined function test.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of software testing technology, and in particular to a method and system for testing user-defined functions in a time-series database driven by a large model. Background Technology

[0002] With the rapid development of the Internet of Things (IoT) and edge computing, time-series databases, as core infrastructure for processing time-series data, are widely used in various industries, such as smart manufacturing, environmental monitoring, and intelligent transportation. Time-series databases (such as Apache IoTDB) allow users to define custom functions to extend data processing capabilities to meet specific business needs, such as complex aggregation calculations, anomaly detection, and feature analysis. However, existing tests on user-defined functions still have the following shortcomings and limitations:

[0003] Inefficient test case generation: Traditional user-defined function (UDF) testing relies heavily on manually writing test cases, which is time-consuming and prone to errors. Furthermore, some time-series data have unique characteristics, such as continuous timestamps and periodic sampling. Manually constructing this data requires a deep understanding of the database kernel and business scenarios, resulting in long test preparation cycles and high costs.

[0004] Automation tools have significant limitations: Existing automated testing tools, including those based on symbolic execution, fuzzing, or genetic algorithms, cannot effectively handle the characteristics of time-series data. Furthermore, these tools are typically designed for general software testing and lack knowledge of time-series databases, resulting in test cases that do not conform to the user-defined function specifications of time-series databases (such as return value type constraints and time window processing rules).

[0005] Large models lack domain adaptation: In recent years, large language models (LLMs) have been used for automated test case generation, but directly applying LLMs to generate test cases for time-series database UDFs suffers from a lack of knowledge. During the pre-training phase, LLMs may not cover the API documentation, UDF best practices, and programming patterns for specific time-series databases, leading to inaccurate or invalid generated test cases. Summary of the Invention

[0006] To address the aforementioned issues, this application provides a method and system for testing user-defined functions in a large model-driven time-series database.

[0007] Firstly, this application provides a method for testing user-defined functions in a large model-driven time-series database, the method comprising:

[0008] Receive and decode the code package of the user-defined function to obtain the source code of the user-defined function;

[0009] The source code of the user-defined function is pre-validated, and it is determined whether the pre-validation passes.

[0010] If the pre-verification passes, a retrieval-enhanced generation strategy is used to extract domain knowledge related to the user-defined function from a pre-built time-series database knowledge base, and static analysis instructions are constructed based on the domain knowledge.

[0011] The source code of the user-defined function and the static analysis instructions are input into the first language model for analysis and processing to obtain the static analysis results;

[0012] The temporal characteristics of the test data are determined based on the static analysis results, and test data conforming to the temporal characteristics are generated using the fine-tuned second language model.

[0013] Test cases are constructed based on the test data, and the user-defined function is tested using the test cases according to a preset execution strategy to obtain dynamic test results.

[0014] Optionally, the pre-verification includes:

[0015] Perform an integrity check on the source code of the user-defined function to verify whether the user-defined function contains the components necessary for its operation;

[0016] Perform a general syntax check on the source code of the user-defined function to verify whether the user-defined function conforms to the preset syntax rules;

[0017] When both the integrity check and the syntax check pass, the user-defined function is determined to be executable.

[0018] Optionally, a time-series database knowledge base may be constructed, including:

[0019] Collect text and code related to the user-defined function from the official repository, technical manuals, API documentation, and open-source community of the target time series database to form an original knowledge set;

[0020] The text and code in the original knowledge set are segmented into fragments to obtain the original knowledge fragment set;

[0021] Generate a corresponding index summary for each original knowledge fragment in the original knowledge fragment set;

[0022] The index summary and the corresponding original knowledge fragment identifier are combined to form an ordered pair of index knowledge, and the ordered pair of index knowledge is stored in the index knowledge base;

[0023] Each original knowledge fragment and its corresponding original knowledge fragment identifier are combined into an ordered content pair, and the ordered content pair is stored in the content knowledge base.

[0024] Establish a mapping relationship between the index knowledge base and the content knowledge base, and form a complete time-series database knowledge base based on the index knowledge base, the content knowledge base, and the mapping relationship.

[0025] Optionally, generating a corresponding index summary for each original knowledge fragment in the original knowledge fragment set includes:

[0026] If the original knowledge fragment is a text fragment, the core semantic points of the text fragment are extracted using natural language processing techniques to generate an index summary of the text fragment;

[0027] If the original knowledge fragment is a code fragment, the logical structure of the code fragment is parsed to generate an index summary of the code fragment.

[0028] Optionally, the step of extracting domain knowledge related to the user-defined function from a pre-built time-series database knowledge base using a retrieval enhancement generation strategy includes:

[0029] The source code of the user-defined function is analyzed to determine the query requirements of the user-defined function;

[0030] The query requirement is converted into a query vector, and the query vector is matched with the index summary in the index knowledge base based on similarity.

[0031] The target index summary most relevant to the query vector is determined based on the similarity matching results;

[0032] Based on the mapping relationship between the index knowledge base and the content knowledge base, the target original knowledge fragment corresponding to the target index summary in the content knowledge base is determined, and the target original fragment is used as the domain knowledge query result.

[0033] Optionally, the step of determining the temporal characteristics of the test data based on the static analysis results, and generating test data conforming to the temporal characteristics using the fine-tuned second language model, includes:

[0034] Extract the functional category of the user-defined function from the static analysis results;

[0035] The user-defined function's functional category is mapped to the temporal features of the test data through a mapping function, which is instantiated from a predefined rule base.

[0036] Based on the temporal features, a prompt instruction for the second language model is constructed, and test data conforming to the temporal features is generated using the second language model based on the prompt instruction.

[0037] Optionally, the step of constructing test cases based on the test data and using the test cases to test the user-defined function according to a preset execution strategy to obtain dynamic test results includes:

[0038] An execution script is generated based on the execution requirements of the user-defined function;

[0039] Test cases are constructed based on the test data, the execution script, test scenario requirements, expected output, and test case metadata.

[0040] Register the user-defined function in the test environment of the time-series database, load and execute the test cases in the test environment according to the preset execution strategy, and obtain dynamic test results.

[0041] Optionally, the method further includes:

[0042] A correlation analysis is performed on the static analysis results and the dynamic test results to obtain the correlation analysis results;

[0043] Based on the correlation analysis results, a structured report is generated for the user-defined function according to the preset quality assessment dimensions.

[0044] Secondly, this application provides a large model-driven time-series database user-defined function testing system, including:

[0045] A preliminary analysis unit is used to receive and decode the code package of a user-defined function to obtain the source code of the user-defined function;

[0046] The preliminary analysis unit is also used to pre-verify the source code of the user-defined function and determine whether the pre-verification passes.

[0047] The knowledge retrieval unit is used to extract domain knowledge related to the user-defined function from a pre-built time-series database knowledge base by employing a retrieval enhancement generation strategy, provided that the pre-verification passes, and to construct static analysis instructions based on the domain knowledge.

[0048] The static analysis unit is used to input the source code of the user-defined function and the static analysis instructions into the first language model for analysis and processing to obtain static analysis results;

[0049] The test management unit is used to determine the timing characteristics of test data based on static analysis results.

[0050] The test data generation unit is used to generate test data that conforms to the time sequence characteristics using the fine-tuned second language model;

[0051] The test execution unit constructs test cases based on the test data and uses the test cases to test the user-defined function according to a preset execution strategy to obtain dynamic test results.

[0052] Thirdly, this application provides a large model-driven time-series database user-defined function testing device, comprising:

[0053] Memory;

[0054] processor;

[0055] The memory stores computer-executed instructions;

[0056] The processor executes computer execution instructions stored in the memory to implement the large model-driven time-series database user-defined function testing method as described in the first aspect and various possible implementations of the first aspect above.

[0057] Fourthly, this application provides a computer storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the large model-driven time-series database user-defined function testing method as described in the first aspect and various possible implementations of the first aspect above.

[0058] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the large model-driven time-series database user-defined function testing method as described in the first aspect and various possible implementations of the first aspect above.

[0059] This application provides a large-model-driven testing method and system for user-defined functions in time-series databases. The method receives and decodes the code package of a user-defined function to obtain its source code; pre-verifies the source code and determines whether the pre-verification passes; if the pre-verification passes, it uses a retrieval-enhanced generation strategy to extract domain knowledge related to the user-defined function from a pre-built time-series database knowledge base, and constructs static analysis instructions based on the domain knowledge; it inputs the source code of the user-defined function and the static analysis instructions into a first large-scale language model for analysis and processing to obtain static analysis results; it determines the temporal characteristics of the test data based on the static analysis results, and uses a fine-tuned second large-scale language model to generate test data that conforms to the temporal characteristics; it constructs test cases based on the test data, and uses the test cases to test the user-defined function according to a preset execution strategy to obtain dynamic test results. This method improves the efficiency of test case generation and the coverage of time-series scenarios, and has higher testing accuracy and adaptability in complex scenarios. Attached Figure Description

[0060] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0061] Figure 1 This is a flowchart illustrating the user-defined function testing method for a large model-driven time series database provided in this application embodiment;

[0062] Figure 2 This is a schematic diagram of the process for constructing a time-series database knowledge base provided in an embodiment of this application;

[0063] Figure 3 This is a flowchart illustrating the retrieval of a time-series database knowledge base provided in an embodiment of this application;

[0064] Figure 4 This is a schematic diagram of the process for generating test data provided in the embodiments of this application;

[0065] Figure 5 This is a schematic diagram of the structure of the test cases provided in the embodiments of this application;

[0066] Figure 6 A schematic diagram of the structure of a large model-driven time-series database user-defined function testing system provided in an embodiment of this application;

[0067] Figure 7 This is a schematic diagram of the structure of the report generation module provided in an embodiment of this application;

[0068] Figure 8A schematic diagram of the structure of a large model-driven time-series database user-defined function test device provided in an embodiment of this application.

[0069] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation

[0070] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0071] The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein.

[0072] In this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0073] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0074] Figure 1 This is a flowchart illustrating the large model-driven time-series database user-defined function testing method provided in this application embodiment. For example... Figure 1 As shown in this embodiment, the method for testing user-defined functions in a large model-driven time series database includes:

[0075] S1: Receive and decode the code package of the user-defined function to obtain the source code of the user-defined function.

[0076] Specifically, after receiving the code package of user-defined functions, a decoding tool is used to convert it into an analyzable source code format.

[0077] S2: Perform pre-validation on the source code of the user-defined function and determine whether the pre-validation passes.

[0078] Specifically, the source code of the user-defined function undergoes an integrity check to verify whether it contains the necessary components for execution, including function definition, parameter declaration, return value declaration, and function body logic. A general syntax check is also performed on the source code to verify whether the user-defined function conforms to preset syntax rules, including language specifications, keyword usage, operator precedence, and statement structure. If both the integrity and syntax checks pass, the user-defined function is deemed executable, and thus passes pre-validation. If either check fails, the validation is considered failed, the process terminates, and an error report is returned.

[0079] This step can determine whether the user-defined function conforms to the basic syntax and specifications of time-series databases.

[0080] S3: If the pre-verification passes, a retrieval enhancement generation strategy is used to extract domain knowledge related to the user-defined function from a pre-built time-series database knowledge base, and static analysis instructions are constructed based on the domain knowledge.

[0081] Figure 2 This is a schematic diagram of the process for constructing a time-series database knowledge base according to an embodiment of this application. Specifically, constructing a time-series database knowledge base includes the following steps:

[0082] S311: Collect text and code related to the user-defined function from the official repository, technical manual, API documentation, and open-source community of the target time series database to form an original knowledge set.

[0083] The target time-series database could be, for example, Apache IoTDB.

[0084] Specifically, textual materials and code related to user-defined functions are collected from the official repositories, technical manuals, API documentation, and open-source communities of the target time-series database to form an initial knowledge set. .

[0085] S312: The text and code in the original knowledge set are segmented to obtain the original knowledge fragment set.

[0086] Specifically, text parsing tools are used to cut long documents into sections, functional points, etc., and code is sliced ​​into functions or logical modules to form semantically independent knowledge fragments, thus constituting the original knowledge fragment set. .

[0087] S313: Generate a corresponding index summary for each original knowledge fragment in the original knowledge fragment set.

[0088] Specifically, if the original knowledge fragment is a text fragment, the core semantic points of the text fragment are extracted using natural language processing technology to generate an index summary of the text fragment; if the original knowledge fragment is a code fragment, the logical structure of the code fragment is parsed to generate an index summary of the code fragment.

[0089] Understandably, to improve the accuracy and efficiency of subsequent searches, natural language processing technology is used to process each knowledge fragment. The data is summarized and refined to generate an index summary that highly encapsulates its core semantics. For text documents, extract the core points from the text snippets and summarize the content. For code snippets, extract information such as method names and functions as core points based on the code logic. This step uses a summary function. This effectively eliminates the interference of redundant code details and descriptive text in semantic matching.

[0090] S314: The index digest and the corresponding original knowledge fragment identifier are combined to form an ordered pair of index knowledge, and the ordered pair of index knowledge is stored in the index knowledge base.

[0091] Specifically, the index digest generated in step S313 Its corresponding unique content identifier Forming an ordered pair Store in the index knowledge base ;

[0092] S315: Form an ordered content pair by combining each original knowledge fragment with its corresponding original knowledge fragment identifier, and store the ordered content pair in the content knowledge base.

[0093] Specifically, the original knowledge fragments obtained in step S312 Its corresponding unique content identifier Forming an ordered pair Store in content knowledge base This achieves the separation of index and content.

[0094] S316: Establish a mapping relationship between the index knowledge base and the content knowledge base, and form a complete time-series database knowledge base based on the index knowledge base, the content knowledge base, and the mapping relationship.

[0095] Understandably, in a stored procedure, a mapping function is created and saved. This function is passed through To achieve precise mapping, i.e. This association ensures that the corresponding complete content can be accurately located through the index. The knowledge base built through the above steps will serve as the knowledge source for static analysis, providing crucial support for subsequent user-defined function analysis and test case generation.

[0096] Figure 3 This is a flowchart illustrating the process of retrieving a time-series database knowledge base according to an embodiment of this application. Taking a user-defined function test scenario of the time-series database Apache IoTDB as an example, a retrieval enhancement generation strategy is used to extract domain knowledge related to the user-defined function from a pre-built time-series database knowledge base, including:

[0097] S321: Analyze the source code of the user-defined function to determine the query requirements of the user-defined function.

[0098] Specifically, the code of user-defined functions is analyzed to determine the domain knowledge to be queried based on the code context. For example, when a user-defined function code contains an internally defined utility function, a query is generated: "The implementation method or related description of a certain utility function in IoTDB".

[0099] S322: Convert the query requirement into a query vector, and perform similarity matching between the query vector and the index summary in the index knowledge base.

[0100] Specifically, the above query requirements are transformed into query vectors, and similarity matching is performed in the index knowledge base. The index knowledge base stores vectorized indexes of knowledge such as the IoTDB official documentation, user-defined function construction processes, and IoTDB open-source code. Through cosine similarity calculation, indexes related to "a certain utility function" are quickly matched.

[0101] S323: Determine the target index summary that is most relevant to the query vector based on the similarity matching results.

[0102] The target index summary is the index summary stored in the index knowledge base that has the highest cosine similarity to the query vector.

[0103] S324: Determine the target original knowledge fragment in the content knowledge base corresponding to the target index summary based on the mapping relationship between the index knowledge base and the content knowledge base, and use the target original fragment as the domain knowledge query result.

[0104] Specifically, based on the defined target index summary and the mapping information stored in the time-series database knowledge base, the corresponding complete knowledge content is precisely located in the content knowledge base. This content may include specific code examples and parameter descriptions. For example, for "a certain utility function," there will be code snippets, the actual function of the utility function, and information on its various parameters.

[0105] Furthermore, the complete knowledge content obtained from the content knowledge base will be integrated as static analysis instructions to enhance the understanding of time-series database-specific syntax and specifications during subsequent static analysis, thereby improving the accuracy of the analysis.

[0106] S4: Input the source code of the user-defined function and the static analysis instructions into the first language model for analysis and processing to obtain the static analysis results.

[0107] The first major language model could be, for example, Qwen2.5-Coder.

[0108] Specifically, the first language model performs static analysis on the user-defined function code based on the static analysis instructions obtained from the above steps. The analysis includes whether the function has CWE vulnerabilities, whether there are logical errors, the function type, and the function's functionality. Finally, it outputs a static analysis report containing information such as the user-defined function category, the user-defined function's functionality, and potential vulnerabilities in the user-defined function.

[0109] S5: Determine the temporal characteristics of the test data based on the static analysis results, and generate test data that conforms to the temporal characteristics using the fine-tuned second language model.

[0110] Figure 4 This is a schematic diagram of the test data generation process provided in the embodiments of this application, which specifically includes the following steps:

[0111] S51: Extract the functional category of the user-defined function from the static analysis results.

[0112] For example, referring to the IoTDB's classification method, user-defined function categories include: data quality, data profiling, anomaly detection, frequency domain analysis, data matching, data repair, sequence discovery, and machine learning.

[0113] S52: Map the functional category of the user-defined function to the temporal characteristics of the test data through a mapping function.

[0114] The mapping function is instantiated using a predefined rule base.

[0115] Specifically, the mapping function is expressed as .in, It is a collection of UDF functional categories. It is a set of time-series data features.

[0116] Understandably, this mapping function It is instantiated from a predefined rule base. By querying this mapping, the characteristics of time series data can be determined. This generates precise test data generation instructions. The core content of the rule base is shown in Table 1:

[0117] Table 1. Feature Mapping Table for Time Series Data

[0118]

[0119] This formal mapping enables the automatic and precise determination of test data generation requirements based on the inherent functionality of user-defined functions, laying the foundation for generating high-quality test data for driving large language models in the future.

[0120] S53: Construct prompts for the second language model based on the temporal features, and use the second language model to generate test data that conforms to the temporal features based on the prompts.

[0121] The second largest language model was obtained by fine-tuning a dataset, which can be derived from the official test cases of IoTDB and its derivatives.

[0122] Specifically, based on the characteristics of time series data Build precise prompts and instructions. ,in This is the instruction generation function, ensuring that instructions can be accurately parsed by the large language model. For example, "Generate time series data with a length of 1000 time points, in which 5 null values ​​are randomly inserted." The large language model generates the corresponding time series data based on the instruction to test its computational correctness and stability.

[0123] S6: Construct test cases based on the test data, and use the test cases to test the user-defined function according to the preset execution strategy to obtain dynamic test results.

[0124] S61: Generate an execution script based on the running requirements of the user-defined function.

[0125] Understandably, scripts required for the build environment are generated based on the execution requirements of user-defined functions. If no special requirements are specified, the default execution script is used.

[0126] S62: Construct test cases based on the test data, the execution script, test scenario requirements, expected output, and test case metadata.

[0127] Specifically, through assembling functions The generated data With test environment configuration The system assembles information such as execution scripts to form complete and defined test cases. For the assembled test cases The content undergoes final verification and formatting, and a complete test case file is output.

[0128] like Figure 5 As shown, it defines a structured test case composition standard specifically for testing user-defined functions in time-series databases. This test case is not a simple data set, but a complete execution unit containing five key elements:

[0129] Use case metadata: contains the unique identifier (ID) of the use case, the associated target user-defined function category, and descriptive information.

[0130] Time-series data: Timestamped test data simulating real-world scenarios, used as input for user-defined functions. Its construction must reflect time-series characteristics, such as periodicity, trends, or inherent outliers.

[0131] Expected output: The correct result or behavior that a user-defined function should return after processing the provided time-series data.

[0132] Environment Requirements: Specify the specific time series database environment configuration required to execute this test case, such as the specific version of the time series database, the required memory, or the time series.

[0133] Execution script: This encapsulates an automated script that automatically registers user-defined functions, loads data, executes tests, and captures results in the test environment.

[0134] This structure ensures that test cases can fully and repeatably verify the behavior of user-defined functions in specific time-series scenarios.

[0135] S63: Register the user-defined function in the test environment of the time series database, load and execute the test cases in the test environment according to the preset execution strategy, and obtain dynamic test results.

[0136] Understandably, the preset execution strategies include normal testing, boundary testing, stress testing, and exception testing. The execution strategy is dynamically adjusted according to the testing objectives. For example, when performing exception testing, a large number of exception test cases will be added.

[0137] Specifically, a time-series database instance is configured in the test environment, user-defined functions are registered, and the generated test data is written to the environment according to the execution strategy. Tests are run according to the preset number of executions, execution frequency, and execution time, while resource consumption and execution status are monitored. Finally, a dynamic test report is output, including test frequency, number of tests, test pass rate, and failed cases.

[0138] Furthermore, the method also includes: performing correlation analysis on the static analysis results and dynamic test results to obtain correlation analysis results; and generating a structured report for the user-defined function according to the correlation analysis results and preset quality assessment dimensions.

[0139] Specifically, the system receives and summarizes static analysis reports and dynamic test results, comparing and correlating the two sets of information to correlate code defects discovered in static analysis with crashes that occurred during testing. Based on this, a structured report is generated that includes test coverage, vulnerability information, and optimization suggestions.

[0140] In an optional embodiment, the generated structured report includes: time-series scenario coverage analysis, which evaluates the coverage of test cases for various time-series data processing scenarios, including the completeness of time window calculation, the correctness of out-of-order data processing, and the sufficiency of verification for boundary conditions such as data interruption; time-series specific vulnerability analysis, in which the report not only lists general programming vulnerabilities but also focuses on revealing specific issues related to time-series data processing, such as timestamp overflow, window boundary calculation errors, and streaming state management defects; and performance and optimization suggestions, which provide performance evaluations for time-series database environments, including query response time analysis and memory usage efficiency evaluation, and provide optimization guidance with time-series characteristics, such as adopting a more efficient window sliding strategy for high-frequency data stream processing.

[0141] Through the above mechanism, the generated comprehensive test report not only includes a quality assessment of user-defined functions, but also provides specific suggestions for optimizing user-defined functions based on the characteristics of time-series databases, providing a comprehensive basis for improving the reliability, performance, and standardization of user-defined functions in time-series data processing scenarios.

[0142] This application provides a large-model-driven testing method for user-defined functions in time-series databases. The method receives and decodes the code package of a user-defined function to obtain its source code. It then pre-verifies the source code and determines whether the pre-verification passes. If the pre-verification passes, it uses a retrieval-enhanced generation strategy to extract domain knowledge related to the user-defined function from a pre-built time-series database knowledge base and constructs static analysis instructions based on this domain knowledge. The source code of the user-defined function and the static analysis instructions are input into a first large-scale language model for analysis, yielding static analysis results. Based on the static analysis results, the method determines the temporal characteristics of the test data and generates test data conforming to these characteristics using a fine-tuned second large-scale language model. Finally, it constructs test cases based on the test data and uses these test cases to test the user-defined function according to a preset execution strategy, obtaining dynamic test results. This method improves the efficiency of test case generation and the coverage of time-series scenarios, exhibiting higher testing accuracy and adaptability in complex scenarios.

[0143] Figure 6 This is a schematic diagram of the structure of a large model-driven time-series database user-defined function testing system provided in an embodiment of this application. Figure 6 As shown, the large model-driven time series database user-defined function testing system 600 provided in this embodiment includes:

[0144] The preliminary analysis unit 601 is used to receive and decode the code package of the user-defined function to obtain the source code of the user-defined function;

[0145] The preliminary analysis unit 601 is also used to pre-verify the source code of the user-defined function and determine whether the pre-verification passes.

[0146] The knowledge retrieval unit 602 is used to extract domain knowledge related to the user-defined function from a pre-built time-series database knowledge base by adopting a retrieval enhancement generation strategy when the pre-verification is passed, and to construct static analysis instructions based on the domain knowledge.

[0147] The static analysis unit 603 is used to input the source code of the user-defined function and the static analysis instructions into the first language model for analysis and processing to obtain static analysis results;

[0148] Test management unit 604 is used to determine the timing characteristics of test data based on static analysis results;

[0149] The test data generation unit 605 is used to generate test data that conforms to the time sequence characteristics using the fine-tuned second language model;

[0150] The test execution unit 606 constructs test cases based on the test data and uses the test cases to test the user-defined function according to a preset execution strategy to obtain dynamic test results.

[0151] The large model-driven time-series database user-defined function testing system provided in this application embodiment also includes: a report generation module. For example... Figure 7 As shown, the report generation module provided in this embodiment includes:

[0152] The data aggregation unit 701 is used to perform correlation analysis on the static analysis results and dynamic test results to obtain correlation analysis results;

[0153] The report generation unit 702 is used to generate a structured report for the user-defined function according to the association analysis results and a preset quality assessment dimension.

[0154] The large model-driven time series database user-defined function testing system provided in this embodiment can execute the large model-driven time series database user-defined function testing method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.

[0155] Figure 8 This is a schematic diagram of the structure of a test device for a large model-driven time-series database user-defined function, provided in an embodiment of this application. Figure 8 As shown in the embodiment of this application, the large model-driven time series database user-defined function test device 800 includes: a receiver 801, a transmitter 802, a processor 803, and a memory 804.

[0156] Receiver 801 is used to receive instructions and data;

[0157] Transmitter 802 is used to send commands and data;

[0158] Memory 804 is used to store instructions executed by the computer;

[0159] Processor 803 is used to execute computer execution instructions stored in memory 804 to implement the various steps of the large model-driven time-series database user-defined function testing method in the above embodiments. For details, please refer to the relevant descriptions in the aforementioned embodiments of the large model-driven time-series database user-defined function testing method.

[0160] Alternatively, the memory 804 described above can be either standalone or integrated with the processor 803.

[0161] When the memory 804 is set up independently, the electronic device also includes a bus for connecting the memory 804 and the processor 803.

[0162] This application embodiment also provides a computer storage medium storing computer execution instructions. When the processor executes the computer execution instructions, it implements the large model-driven time series database user-defined function testing method executed by the large model-driven time series database user-defined function testing device described above.

[0163] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described large model-driven time-series database user-defined function testing method.

[0164] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0165] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.

[0166] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A method for testing user-defined functions in a large model-driven time-series database, characterized in that, The method includes: Receive and decode the code package of the user-defined function to obtain the source code of the user-defined function; The source code of the user-defined function is pre-validated, and it is determined whether the pre-validation passes. If the pre-verification passes, a retrieval-enhanced generation strategy is used to extract domain knowledge related to the user-defined function from a pre-built time-series database knowledge base, and static analysis instructions are constructed based on the domain knowledge. The source code of the user-defined function and the static analysis instructions are input into the first language model for analysis and processing to obtain the static analysis results; The temporal characteristics of the test data are determined based on the static analysis results, and test data conforming to the temporal characteristics are generated using the fine-tuned second language model. Test cases are constructed based on the test data, and the user-defined function is tested using the test cases according to a preset execution strategy to obtain dynamic test results.

2. The method according to claim 1, characterized in that, The pre-verification includes: Perform an integrity check on the source code of the user-defined function to verify whether the user-defined function contains the components necessary for its operation; Perform a general syntax check on the source code of the user-defined function to verify whether the user-defined function conforms to the preset syntax rules; When both the integrity check and the syntax check pass, the user-defined function is determined to be executable.

3. The method according to claim 1, characterized in that, Constructing a time-series database knowledge base, including: Collect text and code related to the user-defined function from the official repository, technical manuals, API documentation, and open-source community of the target time series database to form an original knowledge set; The text and code in the original knowledge set are segmented into fragments to obtain the original knowledge fragment set; Generate a corresponding index summary for each original knowledge fragment in the original knowledge fragment set; The index summary and the corresponding original knowledge fragment identifier are combined to form an ordered pair of index knowledge, and the ordered pair of index knowledge is stored in the index knowledge base; Each original knowledge fragment and its corresponding original knowledge fragment identifier are combined into an ordered content pair, and the ordered content pair is stored in the content knowledge base. Establish a mapping relationship between the index knowledge base and the content knowledge base, and form a complete time-series database knowledge base based on the index knowledge base, the content knowledge base, and the mapping relationship.

4. The method according to claim 3, characterized in that, The step of generating a corresponding index summary for each original knowledge fragment in the original knowledge fragment set includes: If the original knowledge fragment is a text fragment, the core semantic points of the text fragment are extracted using natural language processing techniques to generate an index summary of the text fragment; If the original knowledge fragment is a code fragment, the logical structure of the code fragment is parsed to generate an index summary of the code fragment.

5. The method according to claim 3, characterized in that, The method of extracting domain knowledge related to the user-defined function from a pre-built time-series database knowledge base using a retrieval enhancement generation strategy includes: The source code of the user-defined function is analyzed to determine the query requirements of the user-defined function; The query requirement is converted into a query vector, and the query vector is matched with the index summary in the index knowledge base based on similarity. The target index summary most relevant to the query vector is determined based on the similarity matching results; Based on the mapping relationship between the index knowledge base and the content knowledge base, the target original knowledge fragment corresponding to the target index summary in the content knowledge base is determined, and the target original fragment is used as the domain knowledge query result.

6. The method according to claim 1, characterized in that, The step of determining the temporal characteristics of the test data based on the static analysis results, and generating test data that conforms to the temporal characteristics using the fine-tuned second language model, includes: Extract the functional category of the user-defined function from the static analysis results; The user-defined function's functional category is mapped to the temporal features of the test data through a mapping function, which is instantiated from a predefined rule base. Based on the temporal features, a prompt instruction for the second language model is constructed, and test data conforming to the temporal features is generated using the second language model based on the prompt instruction.

7. The method according to claim 1, characterized in that, The step of constructing test cases based on the test data and using the test cases to test the user-defined function according to a preset execution strategy to obtain dynamic test results includes: An execution script is generated based on the execution requirements of the user-defined function; Test cases are constructed based on the test data, the execution script, test scenario requirements, expected output, and test case metadata. Register the user-defined function in the test environment of the time-series database, load and execute the test cases in the test environment according to the preset execution strategy, and obtain dynamic test results.

8. The method according to claim 1, characterized in that, The method further includes: A correlation analysis is performed on the static analysis results and the dynamic test results to obtain the correlation analysis results; Based on the correlation analysis results, a structured report is generated for the user-defined function according to the preset quality assessment dimensions.

9. A large-model-driven time-series database user-defined function testing system, characterized in that, The system includes: A preliminary analysis unit is used to receive and decode the code package of a user-defined function to obtain the source code of the user-defined function; The preliminary analysis unit is also used to pre-verify the source code of the user-defined function and determine whether the pre-verification passes. The knowledge retrieval unit is used to extract domain knowledge related to the user-defined function from a pre-built time-series database knowledge base by employing a retrieval enhancement generation strategy, provided that the pre-verification passes, and to construct static analysis instructions based on the domain knowledge. The static analysis unit is used to input the source code of the user-defined function and the static analysis instructions into the first language model for analysis and processing to obtain static analysis results; The test management unit is used to determine the timing characteristics of test data based on static analysis results. The test data generation unit is used to generate test data that conforms to the time sequence characteristics using the fine-tuned second language model; The test execution unit constructs test cases based on the test data and uses the test cases to test the user-defined function according to a preset execution strategy to obtain dynamic test results.

10. A large-model-driven time-series database user-defined function testing device, characterized in that, The device includes: Memory; processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the large model-driven time-series database user-defined function testing method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Large language model test case adaptive generation method based on domain knowledge enhancement and closed-loop feedback

    CN121210326A

  • Static analysis tool test case generation method based on program slicing technology

    CN121349891A