Method and device for training Schema generation model based on computing power of intelligent computing center

By automatically building training data sets using large language models on the intelligent computing center and training the Schema generation model, the problems of low degree of automation and low quality of Schema generation in the existing technology are solved, and efficient and accurate Schema generation is achieved.

CN119917162APending Publication Date: 2025-05-02DATACANVAS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411986793.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-02

Smart Images

  • Figure CN119917162A_ABST
    Figure CN119917162A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for training a Schema generation model based on computing power of an intelligent computing center, and the method comprises the steps: obtaining a training data set, which comprises the steps: determining a prompt instruction of a sample to-be-processed task type; generating a development task based on the first large language model; generating a development demand based on the development task and the model; generating and processing code to remove duplicated and low-quality code based on the development requirements and the model; generating Schema based on the processed code and the model; the processed code and Schema serve as a training data set to train a second large language model, and the trained second large language model can output Schema according to the processed code of the input to-be-processed task type. Therefore, automatic construction of the training data set is realized, the efficiency of the training model is improved, a high-quality Schema generation model can be obtained, Schema can be generated based on the model, and efficient and accurate Schema generation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the technical field of computing power infrastructure, and in particular, to a method and device for generating a model based on an intelligent computing center computing power training schema. Background Art

[0002] With the development of artificial intelligence technology and computing power technology, the concept of intelligent computing center has emerged. "Intelligent computing center" refers to the use of large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, mainly for artificial intelligence applications (such as artificial intelligence deep learning model development, model training and model reasoning and other scenarios) to provide the required computing power, data and algorithms. Intelligent computing center covers facilities, hardware, software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.

[0003] Currently, the existing Schema generation methods mainly fall into two categories: manual writing and maintenance, and the use of semi-automatic tools.

[0004] Manual writing and maintenance usually involves developers manually creating and updating documents to ensure that the Schema accurately reflects the design and functionality of the system. Although this approach may seem straightforward in the early stages, it has the following problems: inefficiency, error-proneness, and difficulty in ensuring timely updates and consistency of documents. For example, when the code is modified, developers may forget or delay updating the document, resulting in inconsistencies between the document and the code.

[0005] Semi-automatic tools, such as Swagger / OpenAPI generators, API (Application Programming Interface) document automation tools, and code comment extraction tools, can automatically generate Schema based on comments or specific formats in the code, thereby reducing the workload of manual maintenance. However, they also have the following problems: these tools can usually only generate simple API documents based on code comments, which is difficult to express for complex business logic and data structures, and lack an effective quality assessment mechanism, making it difficult to ensure the quality of the Schema. For example, Swagger / OpenAPI generators can usually only extract basic interface information, such as interface path, parameter type, etc., but cannot automatically extract information such as parameter meanings and constraints.

[0006] In summary, the existing Schema generation methods have technical problems such as low automation, low generation efficiency, and low quality of generated Schema. Summary of the invention

[0007] The embodiments of the present invention provide a method and device for training a Schema generation model based on the computing power of an intelligent computing center, so as to solve the technical problems that the existing Schema generation method has low automation, low generation efficiency, and low quality of the generated Schema.

[0008] In order to solve the above-mentioned technical problems, the present invention is achieved as follows:

[0009] In a first aspect, an embodiment of the present application provides a method for training a Schema generation model based on computing power of an intelligent computing center, the method comprising:

[0010] Step S1: obtaining a training data set, which includes: step S11: determining a prompt instruction of a sample task type to be processed; step S12: generating a development task corresponding to the sample task type to be processed based on the prompt instruction and the first large language model; step S13: generating a development requirement corresponding to the sample task type to be processed based on the development task and the first large language model; step S14: generating a code corresponding to the sample task type to be processed based on the development requirement and the first large language model; step S15: processing the code to remove duplicate and low-quality code to obtain a processed code; step S16: generating a Schema corresponding to the sample task type to be processed based on the processed code and the first large language model; step S17: using the processed code and the Schema together as a training data set;

[0011] Step S2: training a second largest language model based on the training data set, wherein the trained second largest language model can output a corresponding target Schema according to the processed code of the input task type to be processed.

[0012] Optionally, the step S15 includes:

[0013] Step S151: Process the code to remove duplicate and low-quality code to obtain processed code, wherein the processed code includes at least one of the following: function, class, variable; the processing includes at least one of the following: cleaning, formatting, syntax checking.

[0014] Optionally, after step S2, the method further includes:

[0015] Step S3: obtaining the target Schema based on the trained second largest language model;

[0016] Step S4: performing quality assessment on the target Schema to obtain a quality assessment result;

[0017] Step S5: Based on the quality assessment result, the parameters of the second largest language model are adjusted to obtain an adjusted second largest language model.

[0018] Optionally, there are multiple target Schemas, and step S4 includes:

[0019] Step S41: Obtain a first evaluation dimension, wherein the first evaluation dimension includes at least one of the following: necessary field completeness, nested structure completeness, data type definition accuracy, constraint condition rationality, parameter description quality, and example value coverage;

[0020] Step S42: Based on the first evaluation dimension, detect the quality of the target Schema, and score the quality of the target Schema to obtain a sub-score corresponding to each first evaluation dimension;

[0021] Step S43: Calculate the average value of the sub-scores corresponding to each first evaluation dimension, and use the average value as the target score of the target Schema;

[0022] Step S44: According to the target score of each of the target Schemas, the second language model is quality evaluated from a second evaluation dimension to obtain a quality evaluation result, wherein the second evaluation dimension includes: average score, number of Chinese characters, stability, robustness, and failure rate.

[0023] In a second aspect, an embodiment of the present application provides a method for generating a Schema based on computing power of an intelligent computing center, the method comprising:

[0024] Step Sa: Obtain the processed code of the task type to be processed;

[0025] Step Sb: inputting the processed code into the trained second language model to obtain a target Schema corresponding to the task type to be processed;

[0026] The second large language model is a large language model trained based on the method described in the first aspect.

[0027] In a third aspect, an embodiment of the present application provides a device for training a Schema generation model based on the computing power of an intelligent computing center, the device comprising:

[0028] The first acquisition module is used to acquire a training data set, which includes: determining a prompt instruction of a sample task type to be processed; generating a development task corresponding to the sample task type to be processed based on the prompt instruction and the first large language model; generating a development requirement corresponding to the sample task type to be processed based on the development task and the first large language model; generating a code corresponding to the sample task type to be processed based on the development requirement and the first large language model; processing the code to remove duplicate and low-quality code to obtain a processed code; generating a Schema corresponding to the sample task type to be processed based on the processed code and the first large language model; and using the processed code and the Schema together as a training data set;

[0029] A training module is used to train a second language model based on the training data set, and the trained second language model can output a corresponding target Schema according to the processed code of the input task type to be processed.

[0030] In a fourth aspect, an embodiment of the present application provides a device for generating a Schema based on computing power of an intelligent computing center, the device comprising:

[0031] The second acquisition module is used to obtain the processed code of the task type to be processed;

[0032] A Schema generation module, used for inputting the processed code into the trained second language model to obtain a target Schema corresponding to the task type to be processed;

[0033] The second large language model is a large language model trained based on the method described in the first aspect.

[0034] In a fifth aspect, an embodiment of the present invention provides an electronic device, comprising: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the method for training a Schema generation model based on the computing power of an intelligent computing center are implemented as described in the first aspect, or, when the program is executed by the processor, the steps of the method for generating a Schema based on the computing power of an intelligent computing center are implemented as described in the second aspect.

[0035] In a sixth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, the steps of the method for training a Schema generation model based on the computing power of an intelligent computing center are implemented as described in the first aspect, or, when the program is executed by the processor, the steps of the method for generating a Schema based on the computing power of an intelligent computing center are implemented as described in the second aspect.

[0036] In the seventh aspect, an embodiment of the present invention provides a computer program product, comprising computer instructions, which, when executed by the processor, implement the steps of the method for training a Schema generation model based on the computing power of an intelligent computing center as described in the first aspect, or, when executed by the processor, implement the steps of the method for generating a Schema based on the computing power of an intelligent computing center as described in the second aspect.

[0037] The method provided in the embodiment of the present application realizes the automatic construction of the training data set based on the powerful computing power and large language model of the intelligent computing center, and greatly improves the training efficiency of the Schema generation model. Specifically, the high-performance computing resources that the intelligent computing center can provide make the training process of complex models faster and more efficient, and can process massive data and perform high-frequency model iterations, thereby accelerating the convergence speed of the model; and by automatically constructing the training data set, the necessity of human intervention is reduced, and the probability of human error is significantly reduced, which not only improves the accuracy and consistency of the data, but also ensures that the quality of the generated Schema is higher, which can meet the actual application needs (such as automatic generation of UI interfaces, automatic development of interfaces, automatic generation of test cases, etc.); with the computing power of the intelligent computing center, the model can be tested and optimized multiple times in a short period of time, and quickly adapt to various development needs.

[0038] In summary, the powerful computing power of the intelligent computing center has improved the efficiency and quality of schema generation model training, and also achieved efficient, accurate and automated schema generation during the schema generation model usage phase. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0040] Figure 1 A flowchart of a method for generating a model based on the intelligent computing center computing power training Schema provided in an embodiment of the present invention;

[0041] Figure 2 A flowchart of a method for generating a model based on the intelligent computing center computing power training Schema provided in an embodiment of the present invention;

[0042] Figure 3 A flowchart of a method for generating a model based on the intelligent computing center computing power training Schema provided in an embodiment of the present invention;

[0043] Figure 4 A flowchart of a method for generating a Schema based on the computing power of an intelligent computing center provided in an embodiment of the present invention;

[0044] Figure 5 A schematic diagram of a method for constructing a training data set for a Schema generation model provided in an embodiment of the present application;

[0045] Figure 6 A structural block diagram of a device for generating a model for training a Schema based on the computing power of an intelligent computing center provided in an embodiment of the present application;

[0046] Figure 7 A structural block diagram of a device for generating a Schema based on the computing power of an intelligent computing center provided in an embodiment of the present application;

[0047] Figure 8 Schematic diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0048] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0049] The technical terms involved in the present invention are briefly described below.

[0050] The "computing power" mentioned in the present invention is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to execute certain computing requirements. It is the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity. It mainly provides services to the society through computing power infrastructure.

[0051] The "computing power" (CP) described in the present invention is the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, super computing power and intelligent computing power. The commonly used unit of measurement is the number of floating point operations performed per second (FLOPS: Floating Point Operations Per Second, 1EFLOPS = 10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1EFLOPS is about the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream notebooks. The calculation formula is: CP = CP general + CP intelligent + CP super.

[0052] The "Network Power" (NP) described in the present invention is a manifestation of the data transmission capability of computing facilities, including comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, etc. The carrying capacity involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities. In the embodiment of the present invention, the carrying capacity uses the memory bandwidth.

[0053] The "Storage Power" (SP) described in the present invention is the comprehensive ability of a data center in terms of data storage capacity, performance, safety and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and server built-in storage devices. The commonly used unit of measurement for storage capacity is exabyte (EB, 1EB = 2^60bytes), and the commonly used unit of measurement for performance is the number of reads and writes per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB). The disaster recovery ratio is an important manifestation of safety and reliability.

[0054] The "computing power infrastructure" described in the present invention is a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity. It can realize centralized computing, storage, transmission and application of information, and presents characteristics such as diversity and ubiquity, intelligence and agility, security and reliability, and green and low-carbon.

[0055] The “computing power” mentioned in the present invention includes general computing power, intelligent computing power and super computing power.

[0056] The “general computing power” mentioned in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0057] The "intelligent computing power" described in the present invention is a computing platform for large-scale deployment of special chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit) for various innovative applications of artificial intelligence, such as natural language processing and machine vision.

[0058] The "super computing power" mentioned in the present invention is mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, and gene analysis.

[0059] The "intelligent computing center" described in the present invention refers to a facility that provides the required computing power, data and algorithms for artificial intelligence applications (such as artificial intelligence deep learning model development, model training and model reasoning scenarios) by using large-scale heterogeneous computing power resources, including general computing power (CPU: Central Processing Unit) and intelligent computing power (GPU: Graphics Processing Unit, FPGA: Field Programmable Gate Array, ASIC: Application Specific Integrated Circuit, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.

[0060] The "computing resources" mentioned in the present invention refer to the technologies and facilities with information calculation, transmission, storage and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPU (Central Processing Unit), GPU (Graphics Processing Unit), network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, as well as supporting and guarantee resources such as wind, fire, water and electricity.

[0061] The "Large Language Model (LLM)" described in the present invention is an artificial intelligence model based on deep learning, which is designed to understand and generate natural language. These models are usually trained with large amounts of text data to learn the structure, grammar, semantics and context of the language, and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.

[0062] The "Schema" mentioned in the present invention mainly refers to the code input and output structure. Schema is a structured definition used to describe the organization, format and constraints of data. It provides a standardized way to define data models, including field names, data types, relationships, constraints, etc., to ensure the consistency, integrity and validity of data. Schema is widely used in database design, API specifications, data exchange formats and other fields to help developers understand and manage the structure of data and achieve compatibility and interoperability between different systems.

[0063] Please refer to Figure 1 , Figure 1 A method for generating a model based on the intelligent computing center computing power training Schema according to an embodiment of the present application is shown, such as Figure 1 As shown, the method includes:

[0064] Step S1: Obtain training data set;

[0065] Step S2: training a second language model based on the training data set, where the trained second language model can output a corresponding target Schema according to the processed code of the input task type to be processed.

[0066] In step S1, the acquisition of the training data set usually includes the following steps: data collection, data labeling, data cleaning and data formatting, wherein data collection is to collect codes related to the type of task to be processed, which may include open source projects, documents, API definitions, database schemas, etc., and the data source may be code hosting platforms such as GitHub, API documentation websites, technical forums, etc.; data labeling is to label the collected data so that the model can learn how to map the code to the corresponding schema, which involves a detailed description of the function, structure, input and output of the code; data cleaning is to clean the data, remove redundant, erroneous or irrelevant information, and ensure the quality and consistency of the data; data formatting is to organize the data into a format suitable for model training, usually including the pairing of input (processed code) and output (target schema).

[0067] In one possible implementation, Figure 2 As shown, step S1 includes:

[0068] Step S11: Determine the prompt instruction of the type of the sample task to be processed;

[0069] Step S12: generating a development task corresponding to the type of the sample to be processed task based on the prompt instruction and the first language model;

[0070] Step S13: Based on the development task and the first language model, generate development requirements corresponding to the type of the sample task to be processed;

[0071] Step S14: Based on the development requirements and the first language model, generate code corresponding to the type of sample task to be processed;

[0072] Step S15: Process the code to remove duplicate and low-quality codes to obtain processed codes;

[0073] Step S16: Based on the processed code and the first language model, generate a Schema corresponding to the type of task to be processed in the sample;

[0074] Step S17: The processed code and Schema are used as a training data set.

[0075] It should be noted that in step S11, it is first necessary to define clear prompts to guide the first language model (such as GPT-4.0, DeepSeek-V3, etc.) to understand the specific content of the task type to be processed. Prompts usually include background information of the task, expected functions, input and output requirements, etc. Through precise prompts, the first language model can better understand the context of the task and generate more relevant outputs.

[0076] In step S12, the first language model can be used to generate specific development tasks according to the prompt instructions. The development tasks provide the context and environment settings of the tasks, helping developers to better understand the background and requirements of the tasks.

[0077] In step S13, development requirements corresponding to the sample task type to be processed can be generated based on the development task and the first large language model, that is, after the development task is clarified, the first large language model will generate specific development requirements based on these scenarios. Development requirements can include functional requirements, non-functional requirements, performance indicators, etc. In this way, the large language model can convert abstract scenarios into specific and executable requirements documents to ensure that the development team has a clear direction during implementation.

[0078] In step S14, code corresponding to the sample task type to be processed can be further generated based on the development requirements and the first large language model. The large language model will use its understanding of programming languages ​​and logical reasoning capabilities to generate code snippets or modules that meet the requirements. This process can greatly reduce the coding workload of developers and improve the consistency and quality of the code.

[0079] In step S15, after the code is generated, the code needs to be processed to remove duplicate and low-quality code, and the processed code can provide necessary data support for subsequent Schema generation.

[0080] In step S16, a Schema corresponding to the sample task type to be processed can be generated based on the processed code and the first large language model, that is, the large language model will generate a Schema corresponding to the sample task type to be processed using the processed code. This process maps the logic and structure of the code to the Schema, ensuring that the Schema can accurately reflect the function and data model of the code.

[0081] In step S17, the processed code and Schema are used as a new training data set for subsequent training and optimization of the second largest language model. This process not only provides more learning and training materials for the model, but also improves the performance and accuracy of the model in subsequent iterations.

[0082] And in a possible implementation, step S15 further includes: step S151: processing the code to remove duplicate and low-quality code to obtain processed code, wherein the processed code includes at least one of the following: functions, classes, variables; processing includes at least one of the following: cleaning, formatting, syntax checking.

[0083] It should be noted that Figure 2 The methods shown can all be executed relying on the computing resources of an intelligent computing center. An intelligent computing center is usually equipped with powerful computing resources, including high-performance GPUs, GPUs, and large-scale storage systems, which can support complex machine learning tasks and large-scale data processing, thereby improving the efficiency of acquiring training data sets and the quality of training data sets.

[0084] And the same large language model is used to generate the key information required in each step, ensuring the consistency between development tasks, requirements, codes and schemas, and reducing the risk of information inconsistency; as more tasks are processed and data sets are accumulated, the performance of the model can be continuously improved to adapt to new development requirements and technological changes; everything from prompt instructions to development tasks, development requirements, to code, and then to schema can be automatically generated, achieving full automation of the entire process, significantly reducing the manual workload of developers and improving development efficiency.

[0085] In step S2, after obtaining and processing the training data set, the second language model (Schema generation model) can be trained. First, a suitable model architecture can be selected, such as llama, qwen, etc. as the basis of the second language model; then the model is trained using the prepared training data set. The model will gradually adjust its parameters by learning the relationship between the input processed code and the target Schema to improve the accuracy and rationality of the output; during the training process, the model's hyperparameters (such as learning rate, batch size, etc.) can be adjusted to optimize the model's performance; the trained model can also be evaluated using a validation set to check its performance on unseen data to ensure the model's generalization ability.

[0086] It should be noted that there are obvious differences in function and application between the first language model and the second language model. The first language model is a general model suitable for a wider range of application scenarios (such as GPT-4.0, DeepSeek-V3, etc.). In the embodiment of the present application, it is mainly used to generate development tasks, development requirements, code and schema, and then build a training data set for the second language model based on the code and schema.

[0087] The second largest language model is a specially trained model (the model architecture can be llama, qwen), which is suitable for specific application scenarios. Its training goal is to focus on generating Schema from specific processed code.

[0088] In a possible implementation, after step S2, the method further includes: step S3: obtaining a target Schema based on the trained second largest language model; step S4: performing quality assessment on the target Schema to obtain a quality assessment result; step S5: adjusting parameters of the second largest language model based on the quality assessment result to obtain an adjusted second largest language model.

[0089] It should be noted that the trained second largest language model can be used to generate a target schema, and the target schema can be quality evaluated to obtain a quality evaluation result. The quality evaluation result can be evaluated from the following aspects:

[0090] Consistency check: Ensure that the data types, constraints, and relationships in the Schema meet expectations and are consistent with development requirements; Completeness check: Confirm that all necessary entities and attributes are included and that no important information is omitted; Performance evaluation: Evaluate whether the design of the Schema can support the expected query performance and data processing capabilities; Scalability: Check whether the Schema has good scalability to adapt to possible future changes in requirements.

[0091] Through these assessments, the team can get quality assessment results and identify potential problems and improvement points in the schema.

[0092] Afterwards, the parameters of the second largest language model can be adjusted based on the quality assessment results to obtain the adjusted second largest language model. That is, the quality assessment results are used as feedback to identify the deficiencies of the model when generating the Schema, and based on the evaluation results, the model's hyperparameters, training strategies or data sets are adjusted to improve the model's generation capabilities. For example, data in certain specific areas can be added, or the model's learning rate, regularization and other parameters can be adjusted. If the adjustment range is large, the model can also be retrained to ensure that the model can adapt to the new parameter settings and improve the generation quality. Through this adjustment process, the model can continue to learn and adapt, so that it can perform better in subsequent Schema generation, further improving the quality and accuracy of Schema generation.

[0093] In one possible implementation, there are multiple target schemas, such as Figure 3 As shown, step S4 includes:

[0094] Step S41: obtaining a first evaluation dimension;

[0095] The first evaluation dimension includes at least one of the following: necessary field completeness, nested structure completeness, data type definition accuracy, constraint rationality, parameter description quality, and example value coverage;

[0096] Step S42: Based on the first evaluation dimension, detect the quality of the target schema, and score the quality of the target schema to obtain a sub-score corresponding to each first evaluation dimension;

[0097] Step S43: Calculate the average value of the sub-scores corresponding to each first evaluation dimension, and use the average value as the target score of the target Schema;

[0098] Step S44: performing a quality assessment on the second largest language model from a second assessment dimension according to the target score of each target Schema to obtain a quality assessment result;

[0099] Among them, the second evaluation dimension includes: average score, number of Chinese, stability, robustness, and failure rate.

[0100] The first evaluation dimension in step S41 is now described.

[0101] Necessary field completeness: By comparing the generated schema with the manually annotated schema, the coverage of necessary fields is calculated. For example, if the manually annotated schema contains fields A, B, and C, and the generated schema only contains fields A and B, the necessary field completeness is 66.7%.

[0102] Nested structure integrity: Check the correctness of the nested structure by comparing the generated Schema with the manually annotated Schema. For example, if field A in the manually annotated Schema contains subfields A1 and A2, while field A in the generated Schema only contains subfield A1, then the nested structure integrity is not complete.

[0103] Data type definition accuracy: The matching rate is calculated by comparing the generated Schema with the data types actually used in the code. For example, if the type of variable x in the code is int, but the generated Schema defines the type of x as string, the data type definition is inaccurate.

[0104] Reasonableness of constraints: By analyzing the comments in the code, variable naming and other information, determine whether the constraints in the generated Schema are reasonable. For example, if the code comment states that the value range of variable x is 0-100, but the generated Schema defines the value range of x as 0-50, then the constraints are unreasonable.

[0105] Parameter description quality: Use natural language processing techniques (such as BLEU, ROUGE, etc.) to evaluate parameter descriptions, such as fluency and clarity.

[0106] Example value coverage: Checks whether the generated Schema contains valid example values ​​and whether the example values ​​cover all possible value scenarios.

[0107] In step S42, the quality of the target schema can be detected based on the first evaluation dimension, and the quality of the target schema can be scored to obtain the sub-scores corresponding to each first evaluation dimension, and the scoring situation can be shown in the following Table 1. Table 1 is a sub-score table corresponding to the target schema under the first evaluation dimension.

[0108] Table 1

[0109] Evaluation Dimensions Points describe Required Field Completeness 0-20 The more required fields are missing, the lower the score Nested structural integrity 0-15 If the nesting structure is wrong or incomplete, no points will be awarded. Data type definition accuracy 0-20 The more data type definition errors, the lower the score. Reasonableness of constraints 0-15 The lower the score, the more the constraints are inconsistent with the code or unreasonable Parameters describing quality 0-15 The lower the score, the less clear, fluent or missing the parameter description is. Example Value Coverage 0-15 The more missing the sample value or the less coverage, the lower the score

[0110] In step S43, the average value of the sub-scores corresponding to each first evaluation dimension may be calculated, and the average value may be used as the target score of the target Schema.

[0111] In step S44, the second largest language model may be quality evaluated from a second evaluation dimension according to the target score of each target Schema to obtain a quality evaluation result, wherein the second evaluation dimension includes: average score, number of Chinese characters, stability, robustness, and failure rate.

[0112] That is to say, the quality and stability of the overall batch generation of a type of model can be evaluated from the following dimensions: Average score: the arithmetic mean of all Schema scores; Median: the middle value of all Schema scores after sorting; Stability: the degree of fluctuation of the average score and median of Schemas generated in different batches; Robustness: the quality of Schema generation by the model when there is noise or errors in the input code; Failure rate: the rate of ungenerated null values, that is, the proportion of valid Schemas that the model fails to generate.

[0113] The entire evaluation process achieved the following results: by defining clear evaluation dimensions, a comprehensive and systematic quality evaluation of the target schema was ensured; by calculating sub-scores and target scores, the quality evaluation was quantified to facilitate comparison and analysis between different schemas; by evaluating the quality of the second largest language model, a feedback mechanism for the model was formed, which was helpful for subsequent model optimization and adjustment; by evaluating the model's stability, robustness and other dimensions, the model's shortcomings could be identified, thus providing direction for model improvement; the comprehensive score and detailed evaluation results could provide an important basis for the development team, helping the team make more informed decisions during the development process.

[0114] In summary, through the refined evaluation dimensions and quantitative evaluation results, not only the quality of the target schema is improved, but also important feedback and guidance are provided for the improvement of the second largest language model, which promotes the optimization of the entire automated development process.

[0115] Figure 4 A method for generating a Schema based on the computing power of an intelligent computing center according to an embodiment of the present application is shown. Figure 4 As shown, the method includes:

[0116] Step Sa: Obtain the processed code of the task type to be processed;

[0117] Step Sb: Input the processed code into the trained second language model to obtain the target Schema corresponding to the task type to be processed;

[0118] Among them, the second largest language model is based on Figure 1 , Figure 2 , Figure 3 Large language models trained using the method shown.

[0119] It should be noted that based on Figures 1 to 3 The training process of the second largest language model is shown. After the model is trained, the trained second largest language model can be used to generate the target Schema, and the generated Schema can be further applied to scenarios such as automatic generation of UI interfaces, automatic interface development, and automatic generation of test cases, which significantly improves software development efficiency.

[0120] It should be noted that the generated target Schema can also be recorded in a database or file system to ensure its traceability and manageability, and record metadata such as the timestamp of the generated Schema, relevant task information, and the version of the generated model for subsequent analysis.

[0121] In a possible implementation, the aforementioned (such as Figure 3 The target schema is scored based on the evaluation dimensions (such as the completeness of necessary fields, accuracy of data type definitions, etc.) shown in the figure, and the evaluation results are recorded for subsequent model feedback and optimization. The evaluation results can be associated with the generated target schema to form a feedback mechanism. The feedback information (such as evaluation scores, identified problems, etc.) is sorted and input into the training data set of the model; the second largest language model is retrained regularly or as needed, and the performance of the model is improved by using the newly collected target schema and feedback information. The improvement effect of the retrained model in generating schema is evaluated to ensure the continuous optimization of the model; this process is closed-loop. Each time a new type of task to be processed is obtained, steps Sa and Sb are repeated, and the generated target schema is recorded and evaluated. Over time, the model will continue to accumulate more training data and feedback, thereby improving its accuracy and efficiency in schema generation.

[0122] The above introduces the training method of the Schema generation model based on the computing power of the intelligent computing center, and the method of generating the Schema based on the computing power of the intelligent computing center and the Schema generation model from the perspective of methods. Now, from the perspective of virtual modules, the overall process of Schema generation is introduced. Specifically, it can be divided into the following modules:

[0123] Training data construction module: responsible for extracting code from the code base and converting it into an input format acceptable to the model, including the code's grammatical structure, semantic information, comments, etc. This module can also automatically generate development tasks and development requirements through prompt engineering technology to expand training data.

[0124] Specifically, in the training data construction module, the construction process of the training set can refer to Figure 5 .

[0125] Taking the computer vision (CV) development task as an example, 218 development task scenarios and development requirements can be automatically generated, and development code can be generated, which are then automatically parsed into 1992 functions. After deduplication, 1187 functions are cleaned and the schema is finally generated.

[0126] The specific steps are as follows:

[0127] 1. Scenario and requirement generation: Use Prompt Engineering technology to input specific prompts into the SOTA large language model (such as GPT-4.0), such as "generate an image classification development task", and the model will automatically generate corresponding development tasks and development requirements.

[0128] 2. Code generation: Based on the generated development requirements, a large language model is used to automatically generate the corresponding code.

[0129] 3. Code parsing and cleaning: Use code parsing tools (such as ANTLR, Tree-sitter, etc.) to parse the generated code, extract information such as functions, classes, variables, etc., and clean and format it.

[0130] 4.Schema generation: Generate a schema based on the extracted code using a large language model.

[0131] 5. Data integration: Integrate the generated functions and Data-Schema into a training set.

[0132] Model training optimization module: responsible for selecting appropriate pre-trained models (such as Llama, Qwen, etc.) and fine-tuning them using the constructed training data to improve the accuracy and completeness of the model-generated schema. This module can also select the optimal model and optimization strategy through multi-model evaluation and performance comparison.

[0133] Specifically, the optimization scheme based on the model training optimization module is as follows:

[0134] 1. Multi-model evaluation: Try pre-trained models of different types and with different model parameters (e.g., llama, qwen, etc.), fine-tune them using the same training data, and then compare their performance on Schema generation quality.

[0135] 2. Hyperparameter adjustment: Adjust the hyperparameters of model training, such as learning rate, batch size, number of training rounds, etc., to find the optimal parameter combination.

[0136] 3. Prompt engineering optimization: Optimize the prompt of the input model, such as adding more detailed instructions and examples, to improve the accuracy and completeness of the model-generated schema.

[0137] 4. Data augmentation: Use data augmentation techniques (such as code mutation, back translation, etc.) to expand training data to improve the generalization ability of the model.

[0138] 5. Performance comparison: Comprehensively consider the model's inference speed (such as token rate, generation delay, etc.) and schema generation quality, and select the most appropriate model for practical application.

[0139] Quality Assessment Module: Responsible for multi-dimensional assessment of the generated Schema and generating an assessment report. The assessment dimensions include completeness, accuracy, and availability.

[0140] In the quality assessment module, the quality of a single generated Schema is assessed using the following method:

[0141] 1. Necessary field completeness: By comparing the generated schema with the manually annotated schema, the coverage of necessary fields is calculated. For example, if the manually annotated schema contains fields A, B, and C, and the generated schema only contains fields A and B, the necessary field completeness is 66.7%.

[0142] 2. Nested structure integrity: Check the correctness of the nested structure by comparing the generated Schema with the manually annotated Schema. For example, if field A in the manually annotated Schema contains subfields A1 and A2, while field A in the generated Schema only contains subfield A1, then the nested structure integrity is not complete.

[0143] 3. Data type definition accuracy: Calculate the matching rate by comparing the generated Schema with the data types actually used in the code. For example, if the type of variable x in the code is int, but the generated Schema defines the type of x as string, the data type definition is inaccurate.

[0144] 4. Reasonableness of constraints: By analyzing the comments in the code, variable naming and other information, determine whether the constraints in the generated Schema are reasonable. For example, if the code comment states that the value range of variable x is 0-100, but the generated Schema defines the value range of x as 0-50, then the constraints are unreasonable.

[0145] 5. Parameter description quality: Use natural language processing techniques (such as BLEU, ROUGE, etc.) to evaluate parameter descriptions, such as fluency and clarity.

[0146] 6. Example value coverage: Check whether the generated Schema contains valid example values ​​and whether the example values ​​cover all possible value scenarios.

[0147] As shown in Table 1 above, the scores corresponding to the target Schema under different evaluation dimensions are shown.

[0148] The quality and stability of the overall batch generation of a type of model are evaluated from the following dimensions: average score: the arithmetic mean of all Schema scores; median: the middle value of all Schema scores after sorting; stability: the degree of fluctuation of the average score and median of Schemas generated in different batches; robustness: the quality of Schema generation by the model when there is noise or error in the input code; failure rate: the rate of ungenerated null values, that is, the proportion of valid Schemas that the model fails to generate.

[0149] Schema generation module: responsible for receiving code as input and generating Schema using the trained model.

[0150] It should be noted that the Schema generation module is responsible for receiving the code as input and generating the Schema using the trained model. The steps are as follows:

[0151] 1. Generation process: The model generates the Schema based on the input information and requirements.

[0152] 2. Multiple verification mechanism: Multiple verifications are performed on the generated Schema, including: syntax check (whether it complies with the JSONSchema syntax), semantic check (check whether the generated Schema complies with the semantics of the code), and consistency check (consistent with other Schemas in the code base).

[0153] 3. Post-processing: Post-process the schema generated by the model, such as removing redundant information, formatting, and adding additional information to the schema according to the user's specific needs.

[0154] 4. Error handling mechanism: Output errors are packaged and detailed error information is provided, such as error type and error location, to facilitate debugging and troubleshooting. For situations where a valid Schema cannot be generated, a default Schema or prompt information is provided.

[0155] Continuous optimization module: responsible for adjusting model parameters and training strategies based on evaluation results to continuously improve the quality of Schema generation.

[0156] Specifically, the continuous optimization process based on the continuous optimization module can be divided into the following steps:

[0157] 1. Version management system: The different versions of the model are managed through the version control system, which facilitates the backtracking and comparison of the effects and performance of different versions.

[0158] 2. Automated deployment: Use automated deployment tools (such as Jenkins, Docker, Kubernetes, etc.) to automatically deploy the new version of the model to the production environment.

[0159] 3. Performance monitoring: Monitor various performance indicators of the system, including generation speed, resource consumption, error rate, etc.; use monitoring tools (such as Prometheus, Grafana, etc.) to visualize and analyze performance data.

[0160] 4. Quality feedback: Collect user feedback on the generated schema, such as through user ratings, error reports, etc. Analyze user feedback to find out the deficiencies in the model-generated schema.

[0161] 5. Optimization strategy: Adjust model parameters and training strategies based on performance monitoring and quality feedback information, for example:

[0162] (1) Model fine-tuning: retraining the model using new training data or adjusting training parameters;

[0163] (2) Prompt engineering optimization: Optimize the design of Prompt to improve the accuracy and completeness of the model-generated schema;

[0164] (3) Model selection: Try using different models or model architectures.

[0165] (4) Data augmentation: Use new data augmentation methods to expand training data.

[0166] (5) Resource configuration modification: Update the resources used by the model service to maximize the use of the resources required for reasoning.

[0167] In summary, through the various virtual modules shown in the embodiments of the present application, the quality and efficiency of code understanding and automatic Schema generation can be significantly improved, and the specific effects are as follows:

[0168] Improved training data quality: A standardized data collection and processing process was established to ensure high data quality and consistency. The speed and accuracy of data set generation were improved through automated data processing. The diversity and breadth of training data were expanded through the introduction of Prompt Engineering technology.

[0169] Improved Schema generation quality: A multi-dimensional evaluation system is used to fully ensure the quality of the generated Schema, reducing the need for manual intervention. A complete scoring system is established, covering key dimensions such as Schema accuracy, completeness, and availability. The automated quality control mechanism makes the Schema generation process more stable and reliable, reducing human errors.

[0170] The evaluation system is more complete: a standardized quality evaluation dimension is established to ensure a comprehensive and accurate evaluation of the generated schema; the automated scoring mechanism makes the quality evaluation work more efficient and reduces the errors caused by manual operations; a detailed evaluation report is generated to provide clear data support for subsequent optimization;

[0171] Innovation in model optimization mechanism: supports parallel evaluation and performance comparison of multiple models, and selects the best model and optimization strategy from different perspectives; realizes continuous optimization of the model, and improves the generation effect by continuously adjusting the training data and model parameters based on feedback; has the ability to dynamically adjust and can adjust the optimization direction according to different project requirements.

[0172] Improved automation: Through a comprehensive automated process, the entire chain from training data construction to Schema generation is automated; manual intervention is reduced, labor costs are reduced, and development efficiency is improved; the system's response speed and processing efficiency are improved, and it can quickly adapt to the processing needs of large-scale code bases.

[0173] In general, by comprehensively utilizing the computing resources, large language models, automated data construction, quality assessment, and continuous optimization technologies of the intelligent computing center, the quality and efficiency of schema generation can be significantly improved, manual intervention can be reduced, and the degree of automation in software development and document generation can be increased. This not only enables efficient processing of large-scale code bases, but also provides the development team with accurate and complete interface documentation and schemas, thereby improving the overall efficiency of software development.

[0174] Figure 6 A device for training a Schema generation model based on the computing power of an intelligent computing center according to an embodiment of the present application is shown, such as Figure 6 As shown, the device 60 includes:

[0175] The first acquisition module 601 is used to acquire a training data set, which includes: determining a prompt instruction of a sample task type to be processed; generating a development task corresponding to the sample task type to be processed based on the prompt instruction and the first language model; generating a development requirement corresponding to the sample task type to be processed based on the development task and the first language model; generating a code corresponding to the sample task type to be processed based on the development requirement and the first language model; processing the code to remove duplicate and low-quality code to obtain processed code; generating a Schema corresponding to the sample task type to be processed based on the processed code and the first language model; and using the processed code and the Schema together as a training data set;

[0176] The training module 602 is used to train the second largest language model based on the training data set. The trained second largest language model can output the corresponding target Schema according to the processed code of the input task type to be processed.

[0177] Figure 7 A device for generating a Schema based on the computing power of an intelligent computing center according to an embodiment of the present application is shown. Figure 7 As shown, the device 70 includes:

[0178] The second acquisition module 701 is used to obtain the processed code of the task type to be processed;

[0179] Schema generation module 702, used to input the processed code into the trained second language model to obtain a target Schema corresponding to the task type to be processed;

[0180] Among them, the second largest language model is based on Figure 1 , Figure 2 , Figure 3 Large language models trained using the method shown.

[0181] The device provided in the embodiment of the present application realizes the automatic construction of the training data set based on the powerful computing power and large language model of the intelligent computing center, and greatly improves the training efficiency of the Schema generation model. Specifically, the high-performance computing resources provided by the intelligent computing center make the training process of complex models faster and more efficient, and can process massive data and perform high-frequency model iterations, thereby accelerating the convergence speed of the model; and by automatically constructing the training data set, the necessity of human intervention is reduced, and the probability of human error is significantly reduced, which not only improves the accuracy and consistency of the data, but also ensures that the quality of the generated Schema is higher, which can meet the actual application needs (such as automatic generation of UI interfaces, automatic development of interfaces, automatic generation of test cases, etc.); with the computing power of the intelligent computing center, the model can be tested and optimized multiple times in a short period of time, and quickly adapt to various development needs.

[0182] In summary, the powerful computing power of the intelligent computing center has improved the efficiency and quality of schema generation model training, and also achieved efficient, accurate and automated schema generation during the schema generation model usage phase.

[0183] Please refer to Figure 8An embodiment of the present invention further provides an electronic device 80, including a processor 801, a memory 802, and a computer program stored in the memory 802 and executable on the processor 801. When the computer program is executed by the processor 801, the computer program implements the above-mentioned method for training a Schema generation model based on the computing power of an intelligent computing center and the various processes of the method embodiment for generating a Schema based on the computing power of an intelligent computing center, and can achieve the same technical effect. To avoid repetition, it will not be described here.

[0184] The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for training a Schema generation model based on the computing power of an intelligent computing center and the various processes of the method embodiment for generating a Schema based on the computing power of an intelligent computing center are implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0185] An embodiment of the present invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement the various processes of the above-mentioned method for training a Schema generation model based on the computing power of an intelligent computing center and the method for generating a Schema based on the computing power of an intelligent computing center, and can achieve the same technical effect. To avoid repetition, they are not described here.

[0186] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0187] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present invention.

[0188] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation modes, which are merely illustrative rather than restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are within the protection of the present invention.

Claims

1. A method for training a Schema generation model based on the computing power of an intelligent computing center, characterized in that: The method comprises: Step S1: obtaining a training data set, which includes: step S11: determining a prompt instruction of a sample task type to be processed; step S12: generating a development task corresponding to the sample task type to be processed based on the prompt instruction and the first large language model; step S13: generating a development requirement corresponding to the sample task type to be processed based on the development task and the first large language model; step S14: generating a code corresponding to the sample task type to be processed based on the development requirement and the first large language model; step S15: processing the code to remove duplicate and low-quality code to obtain a processed code; step S16: generating a Schema corresponding to the sample task type to be processed based on the processed code and the first large language model; step S17: using the processed code and the Schema together as a training data set; Step S2: training a second largest language model based on the training data set, wherein the trained second largest language model can output a corresponding target Schema according to the processed code of the input task type to be processed.

2. The method according to claim 1, characterized in that The step S15 comprises: Step S151: Process the code to remove duplicate and low-quality code to obtain processed code, wherein the processed code includes at least one of the following: function, class, variable; the processing includes at least one of the following: cleaning, formatting, syntax checking.

3. The method according to claim 1, characterized in that After step S2, the method further comprises: Step S3: obtaining the target Schema based on the trained second largest language model; Step S4: performing quality assessment on the target Schema to obtain a quality assessment result; Step S5: Based on the quality assessment result, the parameters of the second largest language model are adjusted to obtain an adjusted second largest language model.

4. The method according to claim 3, characterized in that There are multiple target Schemas, and step S4 includes: Step S41: Obtain a first evaluation dimension, wherein the first evaluation dimension includes at least one of the following: necessary field completeness, nested structure completeness, data type definition accuracy, constraint condition rationality, parameter description quality, and example value coverage; Step S42: Based on the first evaluation dimension, detect the quality of the target Schema, and score the quality of the target Schema to obtain a sub-score corresponding to each first evaluation dimension; Step S43: Calculate the average value of the sub-scores corresponding to each first evaluation dimension, and use the average value as the target score of the target Schema; Step S44: According to the target score of each of the target Schemas, the second language model is quality evaluated from a second evaluation dimension to obtain a quality evaluation result, wherein the second evaluation dimension includes: average score, number of Chinese characters, stability, robustness, and failure rate.

5. A method for generating a Schema based on the computing power of an intelligent computing center, characterized in that: The method comprises: Step Sa: Obtain the processed code of the task type to be processed; Step Sb: inputting the processed code into the trained second language model to obtain a target Schema corresponding to the task type to be processed; The second large language model is a large language model trained based on the method according to any one of claims 1 to 4.

6. A device for training a Schema generation model based on the computing power of an intelligent computing center, characterized in that: The device comprises: The first acquisition module is used to acquire a training data set, which includes: determining a prompt instruction of a sample task type to be processed; generating a development task corresponding to the sample task type to be processed based on the prompt instruction and the first large language model; generating a development requirement corresponding to the sample task type to be processed based on the development task and the first large language model; generating a code corresponding to the sample task type to be processed based on the development requirement and the first large language model; processing the code to remove duplicate and low-quality code to obtain a processed code; generating a Schema corresponding to the sample task type to be processed based on the processed code and the first large language model; and using the processed code and the Schema together as a training data set; A training module is used to train a second language model based on the training data set, and the trained second language model can output a corresponding target Schema according to the processed code of the input task type to be processed.

7. A device for generating a Schema based on the computing power of an intelligent computing center, characterized in that: The device comprises: The second acquisition module is used to obtain the processed code of the task type to be processed; A Schema generation module, used for inputting the processed code into the trained second language model to obtain a target Schema corresponding to the task type to be processed; The second large language model is a large language model trained based on the method according to any one of claims 1 to 4.

8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the method for generating a model based on the training of a Schema based on the computing power of an intelligent computing center as described in any one of claims 1 to 4, or the program, when executed by the processor, implements the steps of the method for generating a Schema based on the computing power of an intelligent computing center as described in claim 5.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by the processor, implements the steps of the method for generating a model based on the training of a Schema based on the computing power of an intelligent computing center as described in any one of claims 1 to 4, or, when executed by the processor, implements the steps of the method for generating a Schema based on the computing power of an intelligent computing center as described in claim 5.

10. A computer program product, characterized in that The method comprises computer instructions, which, when executed by the processor, implement the steps of the method for training a Schema generation model based on the computing power of an intelligent computing center as described in any one of claims 1 to 4, or, when executed by the processor, implement the steps of the method for generating a Schema based on the computing power of an intelligent computing center as described in claim 5.