Large model algorithm capability data set construction method and apparatus, and electronic device
By constructing a large model algorithm capability data set, using multi-platform data acquisition, cleaning, verification, annotation and standardization, the limitations and quality problems of data sources in the existing technology are solved, and a high-quality and standardized data set and evaluation framework are realized, which improves the scientificity and efficiency of evaluation.
Patent Information
- Application Number
- CN202510069082.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
The existing large-model algorithm capability evaluation technology and data sets have problems such as limitations in data sources, insufficient coverage, uneven data quality, lack of algorithm-related meta information, insufficient standardization of data sets, and lack of evaluation frameworks, which affect the scientificity, comprehensiveness and universality of the evaluation.
A method for building a large-model algorithm capability data set is proposed, including writing network crawlers for multi-platform data acquisition, data cleaning, correctness verification, algorithm annotation, data standardization and evaluation framework modules to build high-quality and standardized data sets and provide evaluation frameworks.
Through multi-platform acquisition, we ensure a wide data coverage and high quality, and add detailed algorithm tags, unify the data format and verification accuracy, make the evaluation results more accurate, and provide interactive tools, which greatly reduces user usage costs.
Smart Images

Figure CN119988907A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of large model technology, and in particular relates to a method, device and electronic device for constructing a large model algorithm capability data set. Background Art
[0002] With the rapid development of artificial intelligence, especially large-scale language models (Large Language Models, referred to as "big models"), these models are widely used in code generation, algorithm design and other fields. However, the existing large model algorithm capability evaluation technology and data sets still have the following problems: limited data sources, insufficient coverage, uneven data quality, lack of algorithm-related meta-information, insufficient data set standardization, and lack of evaluation framework. The above problems seriously affect the scientificity, comprehensiveness and versatility of large model algorithm capability evaluation.
[0003] In order to solve the above problems, it is necessary to propose a method, device and electronic device for constructing a large model algorithm capability data set of the present application. Summary of the invention
[0004] In order to address the deficiencies of the prior art, the present application provides a method, device and electronic device for constructing a large model algorithm capability dataset. The method can solve the problems of limited data sources, insufficient coverage, uneven data quality, lack of algorithm-related meta-information, insufficient dataset standardization and lack of evaluation framework for large language models in the prior art.
[0005] The technical effects to be achieved by this application are achieved through the following solutions: In a first aspect, the present application provides a method for constructing a large model algorithm capability dataset, comprising: Write a web crawler and use the data collection module to crawl data from multiple platforms to obtain raw data; Using a data cleaning module to clean the original data to obtain target cleaned data; Performing correctness verification on the target cleansing data using a correctness verification module to obtain verification data; Performing algorithm annotation on the verification data using an algorithm annotation module to obtain annotated data; Using a data standardization module to perform data standardization on the labeled data to obtain standardized data; The standardized data is evaluated using the evaluation framework module to obtain evaluation results.
[0006] In some embodiments, the method of using a data acquisition module to capture data from multiple platforms includes: The data collection module is used to collect formatted data from multiple platforms based on an API interface. The formatted data includes topic descriptions, code snippets, and test cases.
[0007] In some embodiments, the step of using a data cleaning module to clean the raw data includes: Using a data cleaning module, detecting duplicate codes in the original data based on a MinHash algorithm; Setting a Jaccard similarity threshold, and removing redundant data in the original data based on the Jaccard similarity threshold; The code structure is analyzed using a syntax tree parsing tool, and invalid comments in the original data are cleaned up.
[0008] In some embodiments, the using a correctness verification module to perform correctness verification on the target cleansing data to obtain verification data includes: Using the correctness verification module, automatically generating unit test cases, the unit test cases including input and output comparisons; The unit test case is run to perform a correctness verification test on the target clean data, the code that fails the test in the target clean data is deleted, and the code that passes the test in the target clean data is saved as verification data.
[0009] In some embodiments, the algorithm annotation module is used to perform algorithm annotation on the verification data to obtain the annotation data, including: Using the algorithm annotation module, the verification data is parsed based on the pre-trained BERT model to generate preliminary algorithm labels, wherein the preliminary algorithm labels are manually annotated and include the subject content of the problem, the solution, and possible application scenarios; Screening the preliminary algorithm labels to obtain filtering algorithm labels; Performing usability analysis on the filtering algorithm label to obtain an accurate algorithm label; Classifying the precise algorithm labels; Dynamically adjust the classification of the precise algorithm labels according to the distribution of actual problems and the update of the original data; The precise algorithm labels after classification are manually reviewed to obtain labeled data.
[0010] In some embodiments, using a data standardization module to perform data standardization on the labeled data to obtain standardized data includes: Using a data standardization module, using a standardization tool to convert the labeled data into the standardized data in a target format, wherein the target format includes a Python 3 format; After obtaining the standardized data, the method further includes: The standardized data is cleaned of redundant line breaks, comments, and non-standard naming.
[0011] In some embodiments, the use of the evaluation framework module to evaluate the standardized data to obtain an evaluation result includes: An interactive tool is constructed to evaluate the standardized data using the evaluation framework module, support the comparison of the input model-generated code with the standard answer, and obtain evaluation results, which include three evaluation indicators: output accuracy, efficiency, and coverage.
[0012] In some embodiments, the classified precise algorithm labels include three categories: algorithm design method, data structure and specific application.
[0013] In the second aspect, the present application provides a large model algorithm capability dataset construction device, which includes: a data acquisition module, a data cleaning module, a correctness verification module, an algorithm labeling module, a data standardization module and an evaluation framework module. The large model algorithm capability dataset construction device is used to implement any of the aforementioned large model algorithm capability dataset construction methods.
[0014] In a third aspect, the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the aforementioned methods for constructing a large model algorithm capability dataset when executing the computer program.
[0015] Through the large model algorithm capability data set construction method, device and electronic device of the embodiment of the present application, the method ensures wide data coverage and high quality through multi-platform collection; adds detailed algorithm labels, which can deeply analyze the advantages and disadvantages of the model in different algorithm types; unifies data format and verifies correctness to make evaluation results more accurate; provides interactive tools to greatly reduce user usage costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present application or the existing technical solutions, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 A flowchart of a method for constructing a large model algorithm capability data set in one embodiment of the present application; Figure 2A structural diagram of a large model algorithm capability data set construction device in one embodiment of the present application; Figure 3 It is a schematic block diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.
[0019] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of the present application should be understood by people with ordinary skills in the field to which the present application belongs. The "first", "second" and similar words used in one or more embodiments of the present application do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0020] In related technologies, large model algorithm capability evaluation technology and data sets still have the following problems: 1. Limited data sources and insufficient coverage: Current evaluation data sets mainly come from a single platform (such as LeetCode, HackerRank, etc.). Although these platforms contain many programming questions, their content is limited to certain specific fields or question types, making it difficult to comprehensively evaluate the model's generalization ability and performance in different algorithm tasks.
[0021] 2. Data quality varies: Codes and questions collected from public data sources often contain erroneous codes, incomplete solutions or redundant data, which affects the accuracy and scientificity of the evaluation. In addition, some questions on public platforms contain a large number of duplicates and similar variants, which have not been strictly deduplicated.
[0022] 3. Lack of algorithm-related metadata: Existing datasets lack metadata related to questions and algorithms (such as classification labels, dynamic programming, greedy algorithms, etc.), making it impossible to implement refined capability evaluation for specific algorithm types, resulting in insufficient in-depth analysis of the pros and cons of the model.
[0023] 4. Insufficient data set standardization: There are inconsistencies in data formats and code implementation languages from different sources. For example, many codes are still in Python 2 format and fail to adapt to modern language versions, affecting the universality of the evaluation.
[0024] 5. Lack of evaluation framework: There is a lack of a supporting evaluation framework. Users need to develop their own testing tools to analyze model performance, which increases usage costs and reduces efficiency.
[0025] The above problems seriously affect the scientificity, comprehensiveness and universality of large model algorithm capability evaluation.
[0026] This application aims to propose an efficient and standardized method for constructing evaluation datasets, addressing issues such as data source, quality, metadata annotation, and insufficient standardization, while providing a supporting evaluation framework for convenient use.
[0027] Various non-limiting implementations of the present application are described in detail below in conjunction with the accompanying drawings.
[0028] First, refer to Figure 1 , the method for constructing a large model algorithm capability data set of the present application is described in detail. The present application provides a method for constructing a large model algorithm capability data set, including: S1: Write a web crawler and use the data collection module to crawl data from multiple platforms to obtain raw data; S2: Using a data cleaning module to clean the original data to obtain target cleaned data; S3: Performing correctness verification on the target cleansing data using a correctness verification module to obtain verification data; S4: using an algorithm annotation module to perform algorithm annotation on the verification data to obtain annotated data; S5: using a data standardization module to perform data standardization on the labeled data to obtain standardized data; S6: Use the evaluation framework module to evaluate the standardized data to obtain evaluation results.
[0029] The large-model algorithm capability dataset construction method of this application ensures wide data coverage and high quality through multi-platform collection; adds detailed algorithm labels to deeply analyze the advantages and disadvantages of the model in different algorithm types; unifies data formats and verifies correctness to make evaluation results more accurate; and provides interactive tools to significantly reduce user costs.
[0030] In some embodiments, the method of using a data acquisition module to capture data from multiple platforms includes: The data collection module is used to collect formatted data from multiple platforms based on an API interface. The formatted data includes topic descriptions, code snippets, and test cases.
[0031] For example, by crawling data from multiple platforms (such as Codeforces, LeetCode, HackerRank, etc.), a diverse set of questions and codes can be obtained, covering a wide range of algorithm types; this can achieve wide coverage of data source collection.
[0032] In some embodiments, the step of using a data cleaning module to clean the raw data includes: Using a data cleaning module, detecting duplicate codes in the original data based on a MinHash algorithm; A Jaccard similarity threshold is set, and redundant data in the original data is removed based on the Jaccard similarity threshold; illustratively, the Jaccard similarity threshold may be 0.85, which is exemplary here, and other values known to those skilled in the art may also be applied here, and may be adjusted according to actual circumstances, without limitation thereto.
[0033] The code structure is analyzed using a syntax tree parsing tool, and invalid comments in the original data are cleaned up.
[0034] Exemplarily, the original data may include a training set (eg, 25,443 questions) and a test set (eg, 1,000 questions), that is, used as a code generation data set.
[0035] Each question in the generated data set is matched with as diverse answers as possible. The number of answers can be as high as 1.55 million, for example, to ensure that the model is not prone to overfitting during training and the effectiveness of the evaluation results, thus achieving a larger data scale and higher data quality.
[0036] In some embodiments, the using a correctness verification module to perform correctness verification on the target cleansing data includes: Using the correctness verification module, automatically generating unit test cases, the unit test cases including input and output comparisons; The unit test case is run to perform a correctness verification test on the target clean data, the code that fails the test in the target clean data is deleted, and the code that passes the test in the target clean data is saved as verification data.
[0037] This application has a larger number of test cases / questions, which can more stably alleviate the false positive phenomenon in the test and ensure the validity of the test results.
[0038] In some embodiments, the algorithm annotation module is used to perform algorithm annotation on the verification data to obtain the annotation data, including: Using the algorithm annotation module, the verification data is parsed based on the pre-trained BERT model to generate preliminary algorithm labels, wherein the preliminary algorithm labels are manually annotated and include the subject content of the problem, the solution, and possible application scenarios; Screening the preliminary algorithm labels to obtain filtering algorithm labels; Performing usability analysis on the filtering algorithm label to obtain an accurate algorithm label; Classifying the precise algorithm labels; Dynamically adjust the classification of the precise algorithm labels according to the distribution of actual problems and the update of the original data; The precise algorithm labels after classification are manually reviewed to obtain labeled data.
[0039] Exemplarily, the pre-trained BERT model is used to parse the question text and generate preliminary algorithm labels. The label source and preliminary screening may include: extracting multiple initial labels from the Source Programming Problems dataset, which can be manually annotated and cover the subject content of the problem, solution strategy, and possible application scenarios.
[0040] Criteria for performing an initial screening of preliminary algorithmic labels include: 1) Content relevance: The label should accurately describe the core attributes of the problem, such as mathematics, geometry, graphs, strings, etc.
[0041] 2) Solution feasibility: Prioritize labels that are directly related to algorithm design, such as dynamic programming, binary search, divide and conquer strategy, etc.
[0042] Simplify and categorize the filter algorithm labels to include: After usability analysis, the filter algorithm labels were further streamlined into precise algorithm labels to ensure the following standards: 1) Comprehensive coverage: covers mainstream programming techniques and algorithm design methods.
[0043] 2) Balance between specificity and versatility: It has the ability to conduct detailed analysis while maintaining an overview.
[0044] 3) High label quality: reflects the core characteristics of classic problems and solutions, avoiding duplication and ambiguity.
[0045] The precise algorithm labels are classified, and the classified label system mainly includes the following aspects: 1) Algorithm design methods: such as dynamic programming, divide-and-conquer method, and greedy algorithm.
[0046] 2) Data structures: such as trees, graphs, and matrices.
[0047] 3) Specific applications: such as path planning, flow networks, combinatorial optimization, etc.
[0048] The distribution analysis for each tag includes: 1) Statistical analysis shows that label distribution has a long-tail characteristic, and the vast majority of questions are concentrated in a small number of mainstream categories (such as mathematics and graph algorithms).
[0049] 2) Distribution maps can be used to optimize labeling strategies, break down high-frequency categories into more fine-grained categories, and integrate and classify low-frequency categories.
[0050] The quality assessment and optimization for each tag include: 1) Evaluation indicators: Coverage: whether the problem and solution can be fully described; Classification rationality: whether the division between labels is clear and independent of each other.
[0051] 2) Optimization strategy: Dynamic adjustment: Regularly optimize category division based on problem distribution and updated data; Expert feedback: Improve the labeling system based on the opinions of field experts.
[0052] Manually review and improve the labeling system, covering categories such as dynamic programming, greedy algorithms, and search algorithms.
[0053] Form a labeling system (such as algorithm labels and skill labels) and obtain labeled data.
[0054] Each topic in the dataset of this application contains fine-grained labels such as task topic, algorithm, skill, and difficulty, which provides a more accurate reference for the training and evaluation of the code generation model.
[0055] In some embodiments, using a data standardization module to perform data standardization on the labeled data to obtain standardized data includes: Using a data standardization module, using a standardization tool (such as a 2to3 tool) to convert the annotated data in Python 2 into the standardized data in a target format, wherein the target format includes a Python 3 format; After obtaining the standardized data, the method further includes: The standardized data is cleaned of redundant line breaks, comments, and non-standard naming.
[0056] At present, the model python evaluation frameworks Python2 and Python3 are mixed together, which may easily lead to inaccurate evaluation results. This application unifies the data format to improve the accuracy of the evaluation results.
[0057] In some embodiments, the use of the evaluation framework module to evaluate the standardized data to obtain an evaluation result includes: An interactive tool is constructed to evaluate the standardized data using the evaluation framework module, support the comparison of the input model-generated code with the standard answer, and obtain evaluation results, which include three evaluation indicators: output accuracy, efficiency, and coverage.
[0058] In the evaluation framework of this application, by developing a variety of code processing and code execution adaptations, the output of various types of code writing of the model is supported to ensure the stability of the model evaluation effect.
[0059] In some embodiments, the classified precise algorithm labels include three categories: algorithm design method, data structure and specific application.
[0060] Compared with the prior art, the large model algorithm capability dataset construction method of the present application has the following advantages: 1. Diversity of data sets: Multi-platform collection ensures wide data coverage and high quality.
[0061] 2. Fine-grained algorithm evaluation: Adding detailed algorithm labels can deeply analyze the advantages and disadvantages of the model in different algorithm types.
[0062] 3. High-quality standardized data: Unify data formats and verify correctness to make evaluation results more accurate.
[0063] 4. Convenient evaluation framework: Provides interactive tools to significantly reduce user costs.
[0064] like Figure 2 As shown, the present application provides a large model algorithm capability dataset construction device, and the large model algorithm capability dataset construction device includes: a data acquisition module, a data cleaning module, a correctness verification module, an algorithm labeling module, a data standardization module and an evaluation framework module. The large model algorithm capability dataset construction device is used to implement any of the aforementioned large model algorithm capability dataset construction methods.
[0065] Exemplarily, the data collection module is used to implement a multi-platform data crawling tool, which is responsible for collecting questions and codes. By crawling data from multiple platforms (such as Codeforces, LeetCode, HackerRank, etc.), a diverse set of questions and codes is obtained, covering a wide range of algorithm types.
[0066] Exemplarily, the data cleaning module is used to implement functions including syntax analysis, deduplication algorithm and invalid data screening. Based on syntax tree analysis and similarity detection algorithms (such as MinHash and Jaccard similarity), duplicate and redundant data are deduplicated to ensure the uniqueness and high quality of the data.
[0067] Exemplarily, the correctness verification module is used to verify the correctness and functionality of the code based on the unit test framework. An automated unit test case is built for each topic, and the correctness and functionality of the code are verified by running the test, and the error code is deleted.
[0068] For example, the algorithm annotation module is used to combine the NLP model and the rule base to add labels to the questions. Using natural language processing (NLP) technology and human assistance, algorithm labels (such as "dynamic programming", "graph algorithm", etc.) are automatically generated for the questions to achieve fine-grained evaluation.
[0069] For example, the data standardization module is responsible for language version unification and redundant information cleaning. All codes are converted to Python 3 version and useless comments are cleaned up to form a unified high-quality data format.
[0070] Exemplarily, the evaluation framework module is used to provide a user-friendly large model performance evaluation tool. It provides an evaluation framework that matches the data set and supports multi-dimensional model capability evaluation, including correctness, efficiency, generalization ability, and algorithm type performance.
[0071] The large model algorithm capability dataset construction device of the present application can achieve all the technical effects achieved by the above-mentioned large model algorithm capability dataset construction method, which will not be repeated here.
[0072] It should be noted that the method of one or more embodiments of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of one or more embodiments of the present application, and the multiple devices will interact with each other to complete the described method.
[0073] It should be noted that the above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0074] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present application also discloses an electronic device; Specifically, Figure 3 The hardware structure diagram of an electronic device for constructing a large model algorithm capability data set provided in this embodiment is shown, and the device may include: a processor 410, a memory 420, an input / output interface 430, a communication interface 440, and a bus 450. The processor 410, the memory 420, the input / output interface 430, and the communication interface 440 are connected to each other in communication within the device through the bus 450.
[0075] The processor 410 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0076] The memory 420 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 420 may store an operating system and other application programs. When the technical solution provided in the embodiment of the present application is implemented by software or firmware, the relevant program code is stored in the memory 420 and is called and executed by the processor 410.
[0077] The input / output interface 430 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0078] The communication interface 440 is used to connect a communication module (not shown in the figure) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired manner (for example, USB, network cable, etc.) or a wireless manner (for example, mobile network, Wi-Fi, Bluetooth, etc.).
[0079] The bus 450 includes a path that transmits information between the various components of the device (eg, the processor 410 , the memory 420 , the input / output interface 430 , and the communication interface 440 ).
[0080] It should be noted that, although the above device only shows the processor 410, the memory 420, the input / output interface 430, the communication interface 440 and the bus 450, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiment of the present application, and does not necessarily include all the components shown in the figure.
[0081] The electronic device of the above-mentioned embodiment is used to implement the corresponding large model algorithm capability data set construction method in any of the above-mentioned embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0082] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, one or more embodiments of the present application also provide a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the large model algorithm capability dataset construction method as described in any of the above embodiments.
[0083] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0084] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the large model algorithm capability data set construction method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0085] A person skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. In line with the concept of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of different aspects of one or more embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.
[0086] In addition, to simplify the description and discussion, and in order not to make one or more embodiments of the present application difficult to understand, the well-known power / ground connections to the integrated circuit (IC) chip and other components may or may not be shown in the provided figures. In addition, the device may be shown in the form of a block diagram to avoid making one or more embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which one or more embodiments of the present application will be implemented (that is, these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present application, it is obvious to those skilled in the art that one or more embodiments of the present application can be implemented without these specific details or with changes in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0087] Although the present application has been described in conjunction with specific embodiments of the present application, many alternatives, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the discussed embodiments.
[0088] One or more embodiments of the present application are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present application should be included in the scope of protection of the present application.
Claims
1. A method for constructing a large model algorithm capability dataset, characterized in that: include: Write a web crawler and use the data collection module to crawl data from multiple platforms to obtain raw data; Using a data cleaning module to clean the original data to obtain target cleaned data; Performing correctness verification on the target cleansing data using a correctness verification module to obtain verification data; Performing algorithm annotation on the verification data using an algorithm annotation module to obtain annotated data; Using a data standardization module to perform data standardization on the labeled data to obtain standardized data; The standardized data is evaluated using the evaluation framework module to obtain evaluation results.
2. The method for constructing a large model algorithm capability dataset according to claim 1, characterized in that: The data acquisition module is used to capture data from multiple platforms, including: The data collection module is used to collect formatted data from multiple platforms based on an API interface. The formatted data includes topic descriptions, code snippets, and test cases.
3. The method for constructing a large model algorithm capability dataset according to claim 1 or 2, characterized in that: The using a data cleaning module to clean the original data includes: Using a data cleaning module, detecting duplicate codes in the original data based on a MinHash algorithm; Setting a Jaccard similarity threshold, and removing redundant data in the original data based on the Jaccard similarity threshold; The code structure is analyzed using a syntax tree parsing tool, and invalid comments in the original data are cleaned up.
4. The method for constructing a large model algorithm capability dataset according to claim 3, characterized in that: The using a correctness verification module to perform correctness verification on the target cleansing data to obtain verification data includes: Using the correctness verification module, automatically generating unit test cases, the unit test cases including input and output comparisons; The unit test case is run to perform a correctness verification test on the target clean data, the code that fails the test in the target clean data is deleted, and the code that passes the test in the target clean data is saved as verification data.
5. The method for constructing a large model algorithm capability dataset according to claim 1, characterized in that: The method of using an algorithm annotation module to perform algorithm annotation on the verification data to obtain annotated data includes: Using the algorithm annotation module, the verification data is parsed based on the pre-trained BERT model to generate preliminary algorithm labels, wherein the preliminary algorithm labels are manually annotated and include the subject content of the problem, the solution, and possible application scenarios; Screening the preliminary algorithm labels to obtain filtering algorithm labels; Performing usability analysis on the filtering algorithm label to obtain an accurate algorithm label; Classifying the precise algorithm labels; Dynamically adjust the classification of the precise algorithm labels according to the distribution of actual problems and the update of the original data; The precise algorithm labels after classification are manually reviewed to obtain labeled data.
6. The method for constructing a large model algorithm capability dataset according to claim 1, characterized in that: The data standardization module is used to perform data standardization on the labeled data to obtain standardized data, including: Using a data standardization module, using a standardization tool to convert the labeled data into the standardized data in a target format, wherein the target format includes a Python 3 format; After obtaining the standardized data, the method further includes: The standardized data is cleaned of redundant line breaks, comments, and non-standard naming.
7. The method for constructing a large model algorithm capability dataset according to claim 1, characterized in that: The step of evaluating the standardized data using the evaluation framework module to obtain an evaluation result includes: An interactive tool is constructed to evaluate the standardized data using the evaluation framework module, support the comparison of the input model-generated code with the standard answer, and obtain evaluation results, which include three evaluation indicators: output accuracy, efficiency, and coverage.
8. The method for constructing a large model algorithm capability dataset according to claim 5, characterized in that: The precise algorithm labels after classification include three categories: algorithm design method, data structure and specific application.
9. A device for constructing a large model algorithm capability data set, characterized in that: The large model algorithm capability dataset construction device includes: a data acquisition module, a data cleaning module, a correctness verification module, an algorithm annotation module, a data standardization module and an evaluation framework module. The large model algorithm capability dataset construction device is used to implement the large model algorithm capability dataset construction method as described in any one of claims 1 to 8.
10. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method for constructing a large model algorithm capability dataset as described in any one of claims 1 to 8 when executing the computer program.
Citation Information
Cited By
Method and device for making COT data set
CN121119125A