Method, medium and system for generating computer program language from natural language

Through the improved two-way long and short-term memory network and dynamic resource allocation mechanism, the instability of code generation quality caused by the mixing of mathematical description and plain text description is solved, and efficient and stable code generation is achieved.

CN120255904AInactive Publication Date: 2025-07-04WUHAN TECHN COLLEGE OF COMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510405907.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When the prior art deals with the mixing of mathematical descriptions and plain text descriptions, the code generation quality is unstable, feature extraction is insufficient, computational efficiency is low, and resource allocation is static, resulting in limited system performance.

Method used

Feature extraction is performed using an improved two-way long and short-term memory network, combining matrix decomposition and deep learning to process mathematical and plain text description, dynamic resource allocation and quality assurance mechanism, and program code is generated through feature mapping.

Benefits of technology

It improves the accuracy and efficiency of code generation, ensures the stability of code quality, optimizes the utilization of computing resources, and improves system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120255904A_ABST
    Figure CN120255904A_ABST
Patent Text Reader

Abstract

The invention provides a method, medium and system for generating a computer program language through a natural language, and belongs to the technical field of computer models.The method for generating the computer program language through the natural language comprises the steps that firstly, word segmentation processing is conducted on an input text, and mathematical description and plain text description are recognized; matrix decomposition and a deep learning method are respectively adopted to extract features, and feature fusion is carried out through an improved bidirectional long-short-term memory network. And then dynamic resource allocation is carried out based on accuracy judgment and effectiveness evaluation, and the use of the video memory is optimized by adopting an improved zero-redundancy optimizer technology. And finally, a program code syntax tree is generated through the feature mapping model, and the quality of generated codes is ensured through bidirectional alignment verification and modular processing. A closed-loop optimization mechanism is formed in the whole process, the system performance is continuously improved, and the technical problem of unstable code generation quality caused by mixing of mathematical description and plain text description in the process of converting a natural language into a program code is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer models, and more specifically, relates to a method, medium, and system for generating computer programming languages from natural languages. Background Art

[0002] Generating program code from natural language is an important research direction in the field of artificial intelligence and has been widely applied in scenarios such as intelligent programming assistance, automatic code generation, and program understanding. Traditional techniques for generating program code from natural language mainly use rule-based methods and statistical machine learning methods. Rule-based methods convert natural language into program code through predefined templates and grammar rules. Statistical machine learning methods, on the other hand, establish a language model to learn the correspondence between natural language and program code. These methods have achieved certain results in simple scenarios but still face many challenges when dealing with complex mixed descriptions.

[0003] Current techniques generally suffer from problems such as insufficient feature extraction, low computational efficiency, and unstable generated code quality. In terms of feature extraction, existing methods often treat mathematical descriptions and pure text descriptions uniformly, ignoring the differences in semantic structure and expression between the two, resulting in features that cannot accurately reflect the semantic information of the input text. In terms of computational resource allocation, traditional methods usually adopt a fixed resource allocation strategy and fail to dynamically adjust according to the characteristics of the task, causing resource waste or insufficiency. In terms of code generation, existing techniques are difficult to ensure the quality and consistency of the generated code, especially when dealing with inputs containing complex mathematical descriptions, it is easy to produce syntax errors or semantic deviations.

[0004] The root cause of these problems lies in the lack of an effective mechanism in existing techniques to handle the mixed situation of mathematical descriptions and pure text descriptions. Mathematical descriptions have a strict syntax structure and precise semantics, while pure text descriptions are more flexible and ambiguous. The characteristics of the two are significantly different, and different processing strategies are required. In addition, the static allocation method of computational resources cannot adapt to the processing requirements of different types of inputs, affecting the overall performance of the system. These technical problems severely restrict the practical application effect of the technique for generating program code from natural language. In summary, there are technical problems in the prior art regarding the unstable quality of code generation during the conversion from natural language to program code due to the mixture of mathematical descriptions and pure text descriptions. Summary of the Invention

[0005] In view of this, the present invention provides a method, medium, and system for generating computer programming languages from natural languages, which can solve the technical problem of unstable code generation quality in the prior art during the conversion from natural language to program code due to the mixture of mathematical descriptions and pure text descriptions.

[0006] The present invention is implemented as follows: A method for generating a computer programming language from natural language according to a first aspect of the present invention includes the following steps: obtaining a natural language input text, establishing a language model training data set, and performing word segmentation processing on the natural language input text; using an improved bidirectional long short-term memory network for feature extraction, where an attention mechanism module is provided between the input layer and the hidden layer of the improved bidirectional long short-term memory network to capture long-distance dependencies in the input sequence, and a residual connection module is provided between the hidden layer and the output layer to alleviate the difficulty of deep network training; establishing an accuracy determination function, where the accuracy determination function uses a linear combination calculation method of cosine similarity and edit distance, establishing an accuracy tendency determination function for evaluating the matching degree between the generated code and the expected function, and establishing a utility evaluation function for comprehensive evaluation and calculation based on code running time, memory occupancy, and function completion degree; constructing a feature mapping model to generate a program code syntax tree, performing video memory call optimization, optimizing the structure of the program code syntax tree, and outputting the target program code.

[0007] Among them, the step of obtaining the natural language input text is specifically to construct a text preprocessing module to receive the natural language text input by the user, clean the text, including removing special characters, unifying the encoding format, and correcting obvious grammar errors; establish a language model training data collection system, and collect paired data of natural language descriptions and corresponding program codes from open-source code platforms, programming documents, and technical forums; use a text similarity algorithm to deduplicate and screen the samples in the data set to ensure the quality and diversity of the data set.

[0008] Among them, the step of word segmentation processing is specifically to use a word segmentation algorithm based on conditional random fields to segment the input text; use a part-of-speech tagging module to perform part-of-speech tagging on the word segmentation results, and use a hidden Markov model for part-of-speech analysis; establish a mathematical expression recognition module to recognize the mathematical description part in the text based on regular expressions and rule matching methods, including mathematical formulas, numerical calculations, and logical operation contents.

[0009] Among them, the improved bidirectional long short-term memory network is specifically an improvement based on the standard long short-term memory network. A multi-head attention mechanism is added between the input layer and the hidden layer, and the number of attention heads is set to 8, and the dimension of each attention head is 64; a residual connection is added between the hidden layer and the output layer, and the number of residual blocks is set to 3, and each residual block contains two fully connected layers; the batch size is set to 32, the initial value of the learning rate is 0.001, and the cosine annealing strategy is used to dynamically adjust the learning rate.

[0010] Among them, the accuracy determination function specifically establishes a feature vector normalization module to perform normalization processing on the fused feature matrix; constructs a multi-dimensional similarity calculation module, which includes two sub-modules: vector cosine similarity and edit distance. Cosine similarity is used to calculate the angular similarity between feature vectors, and edit distance is used to calculate the structural similarity of text sequences; the weight of cosine similarity is set to 0.6, and the weight of edit distance is set to 0.4.

[0011] Among them, the steps of constructing the feature mapping model specifically adopt an encoder-decoder architecture. The encoder uses a multi-layer transformer network to process the fused feature matrix, and the decoder uses an attention-based recurrent neural network to generate syntax tree nodes; in the encoding process, position encoding technology is used to maintain the position information of the feature sequence. The number of encoder layers is set to 6 layers, and each layer includes a multi-head self-attention mechanism and a feed-forward neural network; the attention temperature parameter is set to 0.7.

[0012] Among them, the utility evaluation function specifically uses a code running time evaluation module, a memory occupancy evaluation module, and a function completion degree evaluation module to evaluate the time complexity, space complexity, and function coverage rate of the code respectively; the weights of each evaluation module are set to 0.3, 0.3, and 0.4 respectively, and the weighted average is calculated to obtain the final utility score.

[0013] Among them, the video memory call optimization specifically dynamically allocates video memory resources according to the task priority and utility score; adopts an improved zero-redundancy optimizer technology for video memory optimization, including a training parameter sharding storage module that disperses the model parameters in different video memory areas, a gradient calculation distributed processing module that uses a pipelined parallel method to process gradient calculation, and a video memory dynamic recycling module that releases useless intermediate variables after each training step.

[0014] The second aspect of the present invention provides a computer-readable storage medium, in which program instructions are stored. When the program instructions run on a computer, they are used to execute the above method for generating a computer program language for natural language.

[0015] The third aspect of the present invention provides a system for generating a computer program language for natural language, including the above computer-readable storage medium. The system can be any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is set inside the system, and a microprocessor for executing the program instructions stored in the computer-readable storage medium is set inside the system.

[0016] Compared with the prior art, a method, medium, and system for generating a computer programming language from natural language provided by the present invention effectively solve the above technical problems by establishing a feature separation and extraction, dynamic resource allocation, and quality assurance mechanism. The method first performs word segmentation on the input text to identify and separate mathematical descriptions and pure text descriptions, and then adopts different feature extraction strategies. For mathematical descriptions, a method based on matrix factorization is used to extract structured features, and for pure text descriptions, a deep learning model is used to extract semantic features. Finally, the two types of features are organically combined by fusing the feature matrices.

[0017] In terms of resource allocation, the present invention introduces a dynamic scheduling strategy based on utility. The feature matching degree is evaluated through an accuracy determination function, computing resources are dynamically allocated according to the task priority and utility score, and an improved zero-redundancy optimizer technology is used to optimize the video memory usage. In this way, the resource allocation can be flexibly adjusted according to the processing requirements of different inputs, significantly improving the computing efficiency. In terms of the quality of code generation, the present invention ensures the correctness and usability of the generated code by establishing a multi-level quality assurance mechanism, including two-way alignment verification, modular processing, and continuous optimization strategies.

[0018] The innovation of this method lies in the organic combination of feature extraction, resource allocation, and quality assurance, forming a complete technical solution. The accuracy of semantic understanding is improved through precise feature extraction, the computing efficiency of the system is enhanced through dynamic resource allocation, and the quality stability of the generated code is ensured through a multi-level quality assurance mechanism. This systematic solution effectively overcomes the key problems in the prior art and solves the technical problem of unstable code generation quality in the prior art during the conversion from natural language to program code due to the mixture of mathematical descriptions and pure text descriptions. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0021] As Figure 1 shown, it is a flowchart of a method for generating a computer programming language from natural language provided in the first aspect of the present invention. This method includes the following steps:

[0022] S01. Obtain a natural language input text and establish a language model training data set, where the language model training data set includes natural language descriptions and corresponding program code samples;

[0023] S02. Tokenize the natural language input text to generate a first word sequence, and identify the mathematical description part and the plain text description part from the first word sequence;

[0024] S03. Construct a first feature matrix for the mathematical description part, and extract a first mathematical feature matrix using the singular value decomposition method; construct a second feature matrix for the plain text description part;

[0025] S04. Use an improved bidirectional long short-term memory network to extract features from the first mathematical feature matrix and the second feature matrix. The improved bidirectional long short-term memory network sets an attention mechanism module between the input layer and the hidden layer, and sets a residual connection module between the hidden layer and the output layer to obtain a fused feature matrix;

[0026] S05. Establish an accuracy determination function, which is used to calculate a first similarity score between the fused feature matrix and a standard feature library;

[0027] S06. Dynamically allocate computing resources according to the first similarity score, split the computing task into computing subtasks, and use the improved zero redundancy optimizer technology for video memory optimization;

[0028] S07. Construct a feature mapping model, which converts the fused feature matrix into a program code syntax tree;

[0029] S08. Establish an accuracy tendency determination function, which is used to evaluate a second similarity score between the program code syntax tree and a training target;

[0030] S09. Construct a utility degree evaluation function according to the second similarity score, which is used to calculate the utility degree score of the program code;

[0031] S10. Optimize the video memory call based on the utility degree score, allocate video memory according to the task priority and the threshold of the utility degree score, and optimize the structure of the program code syntax tree;

[0032] S11. Perform bidirectional alignment verification between the optimized program code syntax tree and the natural language input text to ensure semantic consistency and functional integrity;

[0033] S12. Modularize the verified program code syntax tree to generate a target program code;

[0034] S13. Store the target program code and its corresponding fused feature matrix in a program code library, and update the language model training data set;

[0035] S14, updating the weight coefficients of the training samples in the language model training data set according to the utility scores to optimize the subsequent training process;

[0036] S15, outputting the target program code as a computer program language result generated from a natural language;

[0037] Among them, the improved zero-redundancy optimizer technology includes a training parameter sharding storage module, a gradient calculation distributed processing module and a video memory dynamic recovery module; the attention mechanism module is used to capture long-distance dependencies in the input sequence; the residual connection module is used to alleviate the difficulty of deep network training; the accuracy judgment function adopts a linear combination calculation method of cosine similarity and edit distance; the accuracy tendency judgment function is used to evaluate the degree of match between the generated code and the expected function; the utility evaluation function performs a comprehensive evaluation calculation based on the code running time, memory usage and function completion, which is used to measure the practical value of the generated code.

[0038] The specific implementation methods of the above steps are described in detail below. The specific implementation method of step S01 is to obtain natural language input text and establish a language model training data set. First, it is necessary to build a text preprocessing module, which receives the natural language text input by the user and cleans the text, including removing special characters, unifying the encoding format, and correcting obvious grammatical errors. Then a language model training data collection system is established to collect paired data of natural language descriptions and corresponding program codes from multiple channels such as open source code platforms, program design documents, and technical forums. The collected data is quality evaluated, samples with poor quality are eliminated, and an initial training data set is established. This step uses a text similarity algorithm to remove and screen samples in the data set to ensure the quality and diversity of the data set. Finally, the processed training data is saved in a structured format for subsequent training. The main purpose of this step is to establish a high-quality training data foundation and provide reliable data support for subsequent model training.

[0039] The specific implementation method of step S02 is to process the input text with word segmentation and identify different types of description parts. First, the word segmentation algorithm based on conditional random fields is used to segment the input text, and this algorithm can effectively handle the ambiguity problem in Chinese word segmentation. Then the word segmentation result is tagged with part of speech using a part-of-speech tagging module, and a hidden Markov model is used to perform part-of-speech analysis. Then a mathematical expression recognition module is established, which identifies the mathematical description part in the text based on regular expressions and rule matching methods, including contents such as mathematical formulas, numerical calculations, and logical operations. Finally, the identified mathematical description part and the plain text description part are stored separately, and their relative position relationship in the original text is maintained. The effect of this step is to decompose the complex input text into different semantic units, in preparation for subsequent feature extraction.

[0040] The specific implementation of step S03 is to construct a feature matrix and perform matrix decomposition. For the mathematical description part, first establish a mathematical symbol mapping dictionary to convert mathematical symbols into a standardized representation form. Then construct the first feature matrix, where the rows of the matrix represent different mathematical expressions and the columns represent feature dimensions. Use the singular value decomposition method to decompose the first feature matrix, retaining the largest first k singular values, and the k value is usually set to 80% of the matrix rank. For the pure text description part, use the bag-of-words model and the term frequency-inverse document frequency method to construct the second feature matrix, where each row of the matrix represents a text segment and each column represents a word feature. The purpose of this step is to convert text information into a numerical representation form that can be processed by a computer.

[0041] The specific implementation of step S04 is to use an improved bidirectional long short-term memory network for feature extraction. This network is improved on the basis of the standard long short-term memory network. A multi-head attention mechanism is added between the input layer and the hidden layer, with the number of attention heads set to 8 and the dimension of each attention head set to 64. A residual connection is added between the hidden layer and the output layer, with the number of residual blocks set to 3, and each residual block contains two fully connected layers. Set the batch size to 32, the initial learning rate to 0.001, and use the cosine annealing strategy to dynamically adjust the learning rate. Use a batch normalization layer to process intermediate features to prevent the problems of gradient disappearance and gradient explosion. The purpose of this step is to extract the deep semantic features of the text and fuse the information of the mathematical description and the pure text description.

[0042] The specific implementation of step S05 is to establish an accuracy determination function to calculate the similarity score. First, establish a feature vector normalization module to normalize the fused feature matrix. Then construct a multi-dimensional similarity calculation module, which includes two sub-modules: vector cosine similarity and edit distance. Cosine similarity is used to calculate the angular similarity between feature vectors, and edit distance is used to calculate the structural similarity of text sequences. Set the cosine similarity weight to 0.6 and the edit distance weight to 0.4, and the weighted sum of the two is used as the final similarity score. When the similarity score is greater than 0.85, it is considered that the feature matching degree is relatively high. The purpose of this step is to evaluate the similarity between the generated features and the standard features.

[0043] The specific implementation of step S06 is to perform dynamic allocation of computing resources and video memory optimization. First, a task scheduling strategy is designed based on the first similarity score, and the higher the score, the higher the priority of the task. Then, the computing tasks are divided into multiple subtasks according to the computing dependency relationship, and the granularity of the subtasks is dynamically adjusted according to the video memory size. The improved zero-redundancy optimizer technology is used for video memory optimization, and this technology includes three core modules: the training parameter sharding storage module dispersedly stores the model parameters in different video memory areas, the gradient calculation distributed processing module processes the gradient calculation in a pipelined parallel manner, and the video memory dynamic recycling module timely releases the useless intermediate variables after each training step. The upper limit of the video memory utilization rate is set to 90%, and when the utilization rate exceeds this threshold, the video memory recycling mechanism is triggered. The purpose of this step is to improve the utilization efficiency of computing resources and reduce the video memory occupancy.

[0044] The specific implementation of step S07 is to construct a feature mapping model to convert the fused feature matrix into a program code syntax tree. This model adopts an encoder-decoder architecture. The encoder uses a multi-layer transformer network to process the fused feature matrix, and the decoder uses an attention-based recurrent neural network to generate the syntax tree nodes. During the encoding process, the position encoding technology is used to maintain the position information of the feature sequence. The number of encoder layers is set to 6 layers, and each layer contains a multi-head self-attention mechanism and a feed-forward neural network. The decoder generates the syntax tree in a top-down manner, and the generation of each node depends on the state of the parent node and the output of the encoder. The attention temperature parameter is set to 0.7, which is used to control the randomness of the generation process. The purpose of this step is to convert the abstract feature representation into a program code syntax tree with a clear structure.

[0045] The specific implementation of step S08 is to establish an accuracy tendency determination function to evaluate the code matching degree. First, a code semantic understanding module is constructed, and this module uses the abstract syntax tree analysis method to extract the semantic features of the program code. Then, a function similarity calculation module is established, and a graph neural network is used to process the syntax tree structure to calculate the semantic matching degree between the generated code and the target function description. In the graph neural network, the features of each node are updated through the message passing mechanism, and the number of iterations is set to 5 times. The function similarity threshold is set to 0.8, and when the similarity exceeds this threshold, it is considered that the code has achieved the expected function. The purpose of this step is to ensure that the generated code meets the user's functional requirements.

[0046] The specific implementation of step S09 is to construct a utility degree evaluation function to calculate the practical value of the program code. This function comprehensively considers three dimensions: the code running time evaluation module estimates the time complexity of the code using static analysis methods, the memory occupancy evaluation module analyzes the space complexity and resource usage of the code, and the function completion degree evaluation module checks the function coverage implemented by the code. The weights of each dimension are set to 0.3, 0.3, and 0.4 respectively, and the weighted average is calculated to obtain the final utility degree score. A utility degree benchmark threshold of 0.75 is set, and the code with a score lower than this threshold needs to be optimized. The purpose of this step is to evaluate the quality of the generated code from a practical perspective.

[0047] The specific implementation of step S10 is to perform video memory call optimization and code structure optimization. The video memory call optimization module dynamically allocates video memory resources according to the task priority and utility degree score. When the utility degree score is greater than 0.8, more video memory resources are allocated for code optimization. The code structure optimization module includes multiple optimization strategies: redundant code elimination, loop optimization, variable scope optimization, etc. The program dependence graph is used to analyze the code structure and identify code fragments that can be optimized. The upper limit of the optimization iteration count is set to 10 times. After each optimization, the utility degree score is recalculated until the score improvement amplitude is less than 0.05 or the iteration limit is reached. The purpose of this step is to improve the execution efficiency and resource utilization rate of the code.

[0048] The specific implementation of step S11 is to perform two-way alignment verification. First, use the semantic understanding module to analyze the syntax tree of the optimized program code and extract its functional semantic representation. Then, perform two-way mapping verification between this semantic representation and the original natural language input text. The verification process includes forward verification and reverse verification. Forward verification checks whether the code completely implements all the functions described in the input text, and reverse verification checks whether there are redundant functions in the code that are irrelevant to the input text. A semantic consistency threshold of 0.9 is set. When the consistency scores of both two-way verifications exceed this threshold, the verification is considered passed. The purpose of this step is to ensure the consistency between the generated code and the user requirements.

[0049] The specific implementation of step S12 is to modularize the program code. First, use a code dependence analysis tool to construct a function call graph and identify the functional modules in the code. Then, reorganize the code according to the functional relevance, and organize the code fragments with high relevance into independent functional modules. Each module is encapsulated, and clear interfaces and data flows are defined to ensure the decoupling between modules. The maximum number of code lines for each module is set to 200 lines, and modules exceeding this limit need to be further split. The purpose of this step is to improve the maintainability and reusability of the code.

[0050] The specific implementation of step S13 is to store code and update the training dataset. First, a program code library management system is constructed. This system adopts a distributed storage architecture and includes three sub-modules: code storage, index establishment, and version control. The code storage module uses the file system to store source code files and simultaneously saves the metadata information of the code. The index establishment module uses the inverted index technology to create multi-dimensional retrieval indexes for the code, including dimensions such as functional features, performance metrics, and application scenarios. The version control module adopts a version management method based on a directed acyclic graph to record the evolution history of the code. When updating the training dataset, an incremental learning strategy is used, and it is decided whether to add the newly generated code to the training set according to the utility score of the new code. The training set update threshold is set to 0.8, and when the utility score of the new code exceeds this threshold, it is added to the training set. The purpose of this step is to accumulate high-quality code samples and continuously improve the generation ability of the system.

[0051] The specific implementation of step S14 is to update the weight coefficients of the training samples. First, a sample evaluation module is established. This module evaluates the importance of each training sample based on historical generation results and actual application effects. A dynamic weight adjustment method based on time decay is adopted, and the most recently used samples have higher weights. The weight update formula takes into account the usage frequency, generation effect, and time factor of the samples, and combines the influences of these factors in a non-linear manner. The weight upper limit is set to 1.5, and the lower limit is set to 0.5. Weights outside the range will be truncated to the boundary values. The weight adjustment is performed regularly, and the adjustment period is set to 100 training iterations. The purpose of this step is to optimize the quality distribution of the training data and improve the model training effect.

[0052] The specific implementation of step S15 is to output the program code result. Format the code according to a preset output template and add necessary comments and documentation. The code formatting applies a unified code style standard, including specifications such as indentation, line breaks, and spaces. Finally, the formatted code is output to the user. The purpose of this step is to ensure the standardization and usability of the output code.

[0053] During the entire implementation process, the specific meanings of various technical terms are as follows: Long Short-Term Memory network is a special recurrent neural network structure that can learn long-term dependencies; Conditional Random Field is a probabilistic graphical model suitable for sequence labeling tasks; Attention mechanism is a dynamic weight allocation method used to focus on important feature information; Residual connection is a network structure design method used to solve the problem of difficult training of deep networks; Singular Value Decomposition is a matrix decomposition method used for dimensionality reduction and feature extraction; Term Frequency-Inverse Document Frequency is a text feature representation method that reflects the importance of words; Graph Neural Network is a deep learning model for processing graph-structured data; Zero Redundancy Optimizer is a video memory optimization technology used to reduce storage overhead during training; Abstract Syntax Tree is a tree-shaped data structure representing the program structure; Program Dependence Graph is used to analyze the dependencies between various parts of a program.

[0054] The following is a detailed description of the functions or calculation processes involved in the present invention.

[0055] 1. For the accuracy determination function, it is specifically expressed as follows:

[0056] S total = α·S cos + β·S edit + γ;

[0057] In the formula, S total is the total similarity score; S cos is the cosine similarity component; S edit is the edit distance component; α is the cosine similarity weight coefficient, with a default value of 0.6; β is the edit distance weight coefficient, with a default value of 0.4; γ is the correction term, ranging from 0 to 0.1.

[0058] Among them, the cosine similarity calculation formula is:

[0059]

[0060] In the formula, x i is the i-th component of feature vector 1; y i is the i-th component of feature vector 2; n is the vector dimension.

[0061] The edit distance calculation formula is:

[0062]

[0063] edit In the formula, L

[0064] is the minimum number of edit operations; len1 is the length of sequence 1; len2 is the length of sequence 2.

[0065] U total = w t ·U t + w m ·U m + w f ·U f + δ;

[0066] In the formula, U total is the total utility score; U t is the time complexity score; U m is the space complexity score; U f is the function completion score; w t , w m , w f are the corresponding weights, with default values of 0.3, 0.3, and 0.4 respectively; δ is the correction factor, with a range of -0.1 to 0.1.

[0067] Time complexity score calculation formula:

[0068]

[0069] In the formula, T actual is the actual running time; T baseline is the baseline running time; λ is the time decay coefficient, with a default value of 0.1; t is the running duration.

[0070] Space complexity score calculation formula:

[0071]

[0072] In the formula, M peak is the peak memory occupancy; M total is the total available memory; k is the adjustment coefficient, with a default value of 2.

[0073] Function completion score calculation formula:

[0074]

[0075] In the formula, f i is the completion status of the i-th function point, taking values of 0 or 1; w i is the weight of the i-th function point; m is the total number of function points.

[0076] 3. For the calculation of the attention mechanism, it is specifically expressed as follows:

[0077]

[0078] In the formula, Q is the query matrix; K is the key matrix; V is the value matrix; d k is the dimension of the key vector, with a default value of 64.

[0079] Multi-head attention calculation formula:

[0080] MultiHead(Q, K, V) = Concat(head1, …, head h )W O ;

[0081]

[0082] wherein, is the parameter matrix of the i-th attention head; W O is the output mapping matrix; h is the number of attention heads, and the default value is 8.

[0083] 4. For the residual connection calculation, it is specifically expressed as follows:

[0084] H l = F(H l-1 ) + H l-1 + ∈;

[0085] wherein, H l is the output of the l-th layer; F is the non-linear transformation function of the residual block; H l-1 is the output of the (l-1)-th layer; ∈ is the noise term, which follows a normal distribution with a mean of 0 and a variance of 0.01.

[0086] 5. For the feature mapping process, it is specifically expressed as follows:

[0087]

[0088] wherein, X is the input feature matrix; L is the number of encoder layers, and the default value is 6; PE is the position encoding matrix.

[0089] Position encoding calculation formula:

[0090]

[0091] wherein, pos is the position index; i is the dimension index; d model is the model dimension.

[0092] The construction principles and meanings of these formulas are as follows:

[0093] 1. The accuracy determination function adopts a linear combination of cosine similarity and edit distance because cosine similarity can capture the similarity of vector directions, while edit distance can reflect the similarity of sequence structures. The combination of the two can more comprehensively measure the similarity degree of features;

[0094] 2. The utility degree evaluation function adopts the form of weighted summation. By introducing exponential and power terms, it can perform non-linear mapping on different types of performance indicators, making the evaluation results more in line with actual needs.

[0095] 3. The calculation of the attention mechanism adopts the form of dot product plus scaling. The introduction of the scaling factor can prevent the problem of gradient disappearance.

[0096] 4. The design of the residual connection retains the original feature information by direct addition, which is helpful for the training of deep networks.

[0097] 5. The positional encoding uses trigonometric functions, which can generate unique encodings for different positions and has good additivity and translational invariance.

[0098] The derivation process of some functions is described in detail below.

[0099] 1. Derivation process of the accuracy determination function:

[0100] First, construct the feature vector matrix:

[0101]

[0102] where x ij represents the j-th feature component of the i-th sample; m is the number of samples; n is the feature dimension; the value range of each feature component is [-1, 1].

[0103] Feature vector normalization process:

[0104]

[0105] where x' ij is the normalized feature value; μ j is the mean of the j-th feature; σ j is the standard deviation of the j-th feature.

[0106] The optimization of cosine similarity considers the feature importance weight:

[0107]

[0108] where w i is the importance weight of the i-th feature, which is obtained by optimizing through backpropagation.

[0109] 2. Derivation of the utility degree evaluation function:

[0110] The time complexity score introduces an exponential decay term:

[0111]

[0112] where t i is the execution time of the i-th code block; c i is the number of executions; k is the number of code blocks.

[0113] The space complexity score takes into account the dynamic changes in memory usage:

[0114] M peak = max{M1, M2,..., M T};

[0115] where M t represents the memory occupancy at time t; T is the total execution duration.

[0116] The function completion score matrix is expressed as:

[0117]

[0118] where f ij represents the verification result of the i-th test case for the j-th function point.

[0119] 3. Complete derivation of the attention mechanism:

[0120] Calculation of the query matrix Q, key matrix K, and value matrix V:

[0121] Q = XW Q ;

[0122] K = XW K ;

[0123] V = XW V ;

[0124] where X is the input feature matrix; W Q , W K , W V are learnable parameter matrices.

[0125] Calculation of the attention weight matrix:

[0126]

[0127] where A is the attention weight matrix, and each element a ij represents the association strength between the i-th query vector and the j-th key vector.

[0128] 4. Detailed derivation of the residual connection:

[0129] Specific form of the non-linear transformation function F:

[0130] F(H l-1 ) = W2σ(W1H l-1+b1)+b2;

[0131] Wherein, W1 and W2 are weight matrices; b1 and b2 are bias vectors; σ is an activation function.

[0132] Normalization of the residual block output:

[0133]

[0134] Wherein, γ and β are learnable scaling and translation parameters; μ and σ are the mean and standard deviation; ∈ is a small constant to prevent division by zero.

[0135] 5. Calculation of the Transformer layer of the feature mapping:

[0136] TransformerLayer(X) = FFN(MultiHead(X)) + X;

[0137] FFN(X) = max(0, XW1 + b1)W2 + b2;

[0138] Wherein, FFN is a feedforward neural network; W1, W2, b1, and b2 are network parameters.

[0139] The second aspect of the present invention provides a computer-readable storage medium, in which program instructions are stored, and when the program instructions run on a computer, they are used to execute the above method for generating a computer program language for natural language.

[0140] The third aspect of the present invention provides a system for generating a computer program language for natural language, including the above computer-readable storage medium, and the system is any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is arranged inside the system, and a microprocessor for executing the program instructions stored in the computer-readable storage medium is arranged inside the system.

[0141] Specifically, the principle of the present invention is: The core principle of this solution is a processing mechanism based on feature separation and fusion. There are essential differences in the semantic structure between mathematical descriptions and pure text descriptions. Separation processing can make full use of their respective characteristics. Mathematical descriptions have clear syntax rules and calculation logics, and are suitable for using mathematical methods such as matrix decomposition to extract features; pure text descriptions contain rich semantic information and are suitable for processing using deep learning models. Through the feature fusion mechanism, the two types of features are organically combined, retaining both the accuracy of mathematical descriptions and the semantic richness of text descriptions.

[0142] The principle of dynamic resource allocation is based on adaptive adjustment according to task characteristics. Different types of inputs require different computing resources. Through accuracy determination and utility evaluation, the system can accurately assess the resource requirements of tasks. The improved zero-redundancy optimizer technology realizes the efficient utilization of video memory resources through parameter sharding storage and dynamic recycling mechanisms. This dynamic resource allocation mechanism based on task characteristics can effectively avoid problems of resource waste or insufficiency.

[0143] The principle of the quality assurance mechanism is based on multi-level verification and optimization. Through two-way alignment verification, the consistency between the generated code and the input requirements is ensured. Through modular processing, the maintainability of the code is improved. Through continuous optimization strategies, the generation quality is continuously improved. This multi-level quality assurance mechanism forms a complete closed loop, ensuring the quality stability of the generated code.

[0144] A specific Embodiment 1 of the present invention is provided below. The specific implementation of each step in this Embodiment 1 is described in detail as follows.

[0145] The specific implementation of step S01 is to obtain the natural language input text and establish a language model training dataset. First, a text preprocessing module is constructed. This module receives the natural language text input by the user and includes three links: text cleaning, normalization, and error correction. In the text cleaning link, the regular expression matching method is used to identify and remove special characters, including emojis, repeated punctuation, illegal characters, etc., while retaining the basic semantic information of the text. In the normalization link, the Unicode normalization algorithm is used to uniformly convert texts in different encoding formats to UTF8 encoding to ensure the consistency of subsequent processing. In the error correction link, the n-gram model is used to detect and correct spelling and grammar errors in the text. The order of the n-gram model is set to 3, and the error detection threshold is set to 0.8. Then, a language model training data collection system is established. This system collects paired data of natural language descriptions and corresponding program codes from multiple channels such as open-source code platforms, programming documents, and technical forums through web crawler technology. During the collection process, the depth-first search strategy is used, and the maximum collection depth is set to 5 layers, and the collection interval for a single website is not less than 1 second. The quality of the collected data is evaluated, and the evaluation indicators include three dimensions: text integrity, code executability, and pairing accuracy. Text integrity is evaluated by calculating the text length and keyword density. It is required that the text length is between 50 and 1000 characters, and the keyword density is not less than 0.05. Code executability is checked for syntax errors and runtime dependencies through static code analysis tools. It is required that the code can be successfully compiled and there are no obvious logical errors. Pairing accuracy uses the text similarity algorithm to calculate the similarity between the natural language description and the code comment, and the similarity threshold is set to 0.7. An initial training dataset is established through these screening criteria, and then the text similarity algorithm is used to deduplicate and screen the dataset. The calculation of text similarity uses the TF-IDF vector space model to convert the text into a high-dimensional feature vector and calculate the cosine similarity between the vectors. The similarity threshold is set to 0.85, and samples with a similarity higher than this threshold are regarded as duplicate samples. Finally, the processed training data is saved in a structured format, a distributed storage system is used for data management, and an asynchronous writing strategy is used to improve the data processing efficiency. The core purpose of this step is to establish a high-quality training data foundation and ensure the effect of subsequent model training through strict data cleaning and quality control. Throughout the process, algorithms such as regular expressions, n-gram models, depth-first search, and TF-IDF are used. The selection of these algorithms is based on their excellent performance and wide application in the fields of text processing and data collection. The threshold parameters for data processing are determined through a large number of experimental verifications. The selection of these parameters balances data quality and processing efficiency. At the same time, the use of a distributed storage system and an asynchronous writing strategy effectively improves the efficiency of large-scale data processing.

[0146] The specific implementation of step S02 is to perform word segmentation on the natural language input text and identify different types of descriptive parts. First, a word segmentation algorithm based on conditional random fields is used to segment the input text. The feature functions of the conditional random field model include three types: character features, position features, and transition features. The feature weights are learned through the maximum likelihood estimation method. The character features consider the combination forms of single characters, double characters, and triple characters. The position features consider the relative position information of characters in the sentence. The transition features consider the transition probabilities between adjacent labels. The Viterbi algorithm is used in the word segmentation process to solve the optimal annotation sequence, and the time complexity of the algorithm is O(n·m 2 ), where n is the sentence length and m is the number of labels. Then, a part-of-speech tagging module is used to perform part-of-speech tagging on the word segmentation result, and a hidden Markov model is used for part-of-speech analysis. The parameters of the hidden Markov model include the initial probability distribution π, the transition probability matrix A, and the emission probability matrix B, and these parameters are estimated from the annotated corpus through a supervised learning method. Next, a mathematical expression recognition module is established, and this module identifies the mathematical description parts in the text based on regular expressions and rule matching methods. The mathematical expression recognition adopts a hierarchical recognition strategy. First, basic mathematical symbols and operators are recognized, and then composite expressions and formula structures are recognized. The recognition rules are implemented using a finite state automaton, and the state transition diagram of the automaton includes multiple sub-modules such as numerical recognition, operator recognition, parenthesis matching, and variable recognition. For the recognized mathematical description parts and pure text description parts, their relative position relationships in the original text are maintained, and an index mapping table is established to record the start position and length information of the segments.

[0147] The specific implementation of step S03 is to construct a feature matrix and perform matrix factorization. For the mathematical description parts, first, a mathematical symbol mapping dictionary is established, which contains the standardized representations of common mathematical symbols, operators, function names, etc. The symbol mapping is stored in a hash table structure, supporting fast lookup with a time complexity of O(1). Then, a first feature matrix M1 is constructed according to the symbol occurrence frequency and position information. Each row of the matrix represents a mathematical expression, and each column represents a feature dimension. The feature dimensions include information such as symbol frequency, symbol position, and symbol combination. The singular value decomposition method is used to decompose the first feature matrix:

[0148] M1 = U∑V T ;

[0149] In the formula, U is the left singular vector matrix; ∑ is the singular value diagonal matrix; V is the right singular vector matrix. The largest first k singular values are retained, and the k value is usually set to 80% of the matrix rank to obtain the reduced-dimensional first mathematical feature matrix. For the pure text description parts, a second feature matrix M2 is constructed using the bag-of-words model and the term frequency-inverse document frequency method. The term frequency calculation formula is:

[0150]

[0151] In the formula, n ij is the number of occurrences of word i in document j. The inverse document frequency calculation formula is:

[0152]

[0153] In the formula, N is the total number of documents; df i is the number of documents containing word i. The final feature weight is:

[0154] w ij = tf ij ·idf i .

[0155] The specific implementation of step S04 is to use an improved bidirectional long short-term memory network for feature extraction. This network is improved on the basis of the standard long short-term memory network, adding a multi-head attention mechanism between the input layer and the hidden layer. The number of attention heads is set to 8, and the dimension of each attention head is 64. The calculation of the multi-head attention mechanism follows the aforementioned formula:

[0156] MultiHead(Q, K, V) = Concat(head1,..., head h )W O ;

[0157]

[0158] A residual connection is added between the hidden layer and the output layer. The number of residual blocks is set to 3, and each residual block contains two fully connected layers. The calculation of the residual connection follows the aforementioned formula:

[0159] H l = F(H l-1 ) + H l-1 + ∈.

[0160] Set the batch size to 32, the initial learning rate to 0.001, and adopt a cosine annealing strategy to dynamically adjust the learning rate. The learning rate adjustment formula is:

[0161]

[0162] In the formula, η t is the learning rate at the t-th step; η min is the minimum learning rate, set to 0.00001; η max is the maximum learning rate, set to 0.001; T is the total number of steps. Use a batch normalization layer to process the intermediate features. The normalization formula is:

[0163]

[0164] where, μ B is the batch mean; is the batch variance; γ and β are learnable parameters; ∈ is a small constant to prevent division by zero, with a value of 0.00001.

[0165] The specific implementation of step S05 is to construct an accuracy determination function to calculate the similarity score. First, a feature vector normalization module is established to perform normalization processing on the fused feature matrix. The normalization uses the aforementioned formula:

[0166]

[0167] Then, a multi-dimensional similarity calculation module is constructed, which includes two sub-modules: vector cosine similarity and edit distance. The cosine similarity calculation uses the aforementioned formula:

[0168]

[0169] The edit distance is calculated using the dynamic programming algorithm, and the state transition equation is defined as:

[0170] dp[i][j] = min{dp[i - 1][j] + 1, dp[i][j - 1] + 1, dp[i - 1][j - 1] + cost};

[0171] where, cost is the cost of the substitution operation, which is 0 when the characters are the same and 1 when the characters are different. Finally, the total similarity score is calculated:

[0172] S total = α·S cos + β·S edit + γ;

[0173] For the parameter settings in the formula, refer to the aforementioned description. This step realizes the precise evaluation of feature matching by combining different similarity measurement methods.

[0174] The specific implementation of step S06 is to perform dynamic allocation of computing resources and video memory optimization. First, a task scheduling strategy is designed according to the first similarity score, and the higher the score, the higher the task priority. The task priority calculation formula is:

[0175] P = α·S total + β·T remain + γ·R available ;

[0176] where, P is the task priority score; S total is the similarity score; T remain is the remaining time; R availableis the proportion of available resources; α, β, and γ are weight coefficients, taking values of 0.5, 0.3, and 0.2 respectively. Then, the computing tasks are divided into multiple subtasks according to the computing dependency relationship, and the granularity of the subtasks is dynamically adjusted according to the video memory size. The task division uses a directed acyclic graph to represent the dependency relationship, and the critical path algorithm is used to determine the task execution order. The improved zero-redundancy optimizer technology is adopted for video memory optimization, and this technology includes three core modules. The training parameter sharding storage module stores the model parameters dispersedly in different video memory areas, and the sharding size calculation formula is:

[0177]

[0178] In the formula, S chunk is the sharding size; M total is the total video memory size; N param is the total number of parameters; α is the sharding coefficient, taking a value of 0.8. The gradient calculation distributed processing module processes the gradient calculation in a pipelined parallel manner, with a pipeline depth of 4, and the computing load balance factor of each stage does not exceed 1.2. The video memory dynamic recycling module releases useless intermediate variables after each training step, and the recycling trigger condition is:

[0179]

[0180] In the formula, M used is the used video memory; M total is the total video memory.

[0181] The specific implementation manner of step S07 is to construct a feature mapping model to convert the fused feature matrix into a program code syntax tree. This model adopts an encoder-decoder architecture. The encoder uses a multi-layer transformer network to process the fused feature matrix, and the decoder uses an attention-based recurrent neural network to generate syntax tree nodes. The calculation formula for the encoding process is:

[0182]

[0183] The calculation formula for position encoding is:

[0184]

[0185] The decoder generates the syntax tree in a top-down manner, and the generation probability of each node is:

[0186] P(n t |n <t ) = softmax(W·h t +b);

[0187] In the formula, n t is the current node; n <t is the historical node sequence; h tis in the hidden state; W and b are model parameters. Node types include declaration nodes, expression nodes, control flow nodes, etc. The dependencies between nodes are modeled through the attention mechanism.

[0188] The specific implementation of step S08 is to establish an accuracy tendency determination function to evaluate the code matching degree. First, a code semantic understanding module is constructed, which uses the abstract syntax tree analysis method to extract the semantic features of the program code. The semantic feature extraction adopts a tree-shaped recurrent neural network, and the state update formula of the network is:

[0189] h t = tanh(W x ·x t + W h ·h t-1 + b);

[0190] In the formula, h t is the node state; x t is the node feature; W x , W h , and b are network parameters. Then, a function similarity calculation module is established, and a graph neural network is used to process the syntax tree structure. The message passing mechanism of the graph neural network is defined as:

[0191]

[0192] In the formula, m i is the message received by node i; h i is the node state; N(i) is the neighbor set of node i; MLP is a multi-layer perceptron; GRU is a gated recurrent unit.

[0193] The specific implementation of step S09 is to construct a utility degree evaluation function to calculate the practical value of the program code. This function comprehensively considers three dimensions: code running time, memory occupancy, and function completion degree. The calculation formulas for each dimension are as follows:

[0194]

[0195] The formula for the final utility degree score is:

[0196] U total = w t ·U t + w m ·U m + w f ·U f + δ;

[0197] The meanings and values of the parameters in the formula are as described above. The utility degree evaluation result is used to guide the code optimization process. When the evaluation score is lower than the threshold of 0.75, the optimization mechanism is triggered.

[0198] The specific implementation of step S10 is to optimize the video memory call and the code structure. The video memory call optimization module dynamically allocates video memory resources according to the task priority and the utility score. The video memory allocation strategy is implemented using a priority queue, and the queue priority calculation formula is:

[0199]

[0200] In the formula, Priority i is the priority of task i; U total is the utility score; M i is the video memory requirement of task i; M total is the total video memory capacity; T i is the estimated execution time of task i; w1, w2, w3 are weight coefficients, with values of 0.5, 0.3, and 0.2 respectively. The code structure optimization module includes multiple optimization strategies. Redundant code elimination uses data flow analysis to define the set of active variables:

[0201] Live in (B) = Use B ∪ (Live out (B) - Def B );

[0202] Live out (B) = ∪ S∈succ(B) Live in (S);

[0203] In the formula, B represents a basic block; Use B is the use set; Def B is the definition set; succ(B) is the set of successor basic blocks. Loop optimization uses techniques such as loop invariant code motion and loop unrolling. The loop unrolling factor calculation formula is:

[0204]

[0205] In the formula, C cache is the cache size; S iter is the storage requirement per iteration; F max is the maximum unrolling factor, set to 8.

[0206] The specific implementation of step S11 is to perform bidirectional alignment verification. First, use the semantic understanding module to analyze the syntax tree of the optimized program code and extract its functional semantic representation. The semantic representation uses a vector space model:

[0207]

[0208] In the formula, V code is the code semantic vector; vi is the vector of basic semantic units; w i is the weight coefficient; n is the number of semantic units. The two-way mapping verification is performed between this semantic representation and the original natural language input text, and the verification process includes forward verification and backward verification. The forward verification calculates the code implementation degree:

[0209]

[0210] In the formula, f i is the requirement feature; c i is the code feature; match() is the feature matching function; w i is the feature weight. The backward verification calculates the code purity:

[0211]

[0212] In the formula, r i is the redundant code segment; size() is the code scale measurement function; k is the number of redundant segments.

[0213] The specific implementation of step S12 is to modularize the program code. First, use the code dependency analysis tool to construct the function call graph, and the calculation formula for the dependency relationship strength is:

[0214]

[0215] In the formula, D ij is the dependency strength between modules i and j; C ij is the number of common calls; C i , C j are the respective call times. Then, reorganize the code according to the functional relevance, and the module division standard is:

[0216]

[0217] In the formula, Q is the module division quality; A ij is the actual dependency strength; P ij is the expected dependency strength; m is the number of modules.

[0218] The specific implementation of step S13 is to store the code and update the training dataset. The code storage module uses a distributed file system, and the data chunking strategy is:

[0219]

[0220] In the formula, S block is the block size; S max , S min is the size limit; S total is the total data volume; N nodesis the number of nodes. The index building module uses the inverted index technology, and the index item weight calculation formula is:

[0221] W term = tf·idf·norm l ·norm t ;

[0222] In the formula, tf is the term frequency; idf is the inverse document frequency; norm l , norm t are the length normalization factor and the time decay factor respectively.

[0223] The specific implementation of step S14 is to update the weight coefficient of the training sample. The sample importance evaluation formula is:

[0224] I s = α·F s + β·T s + γ·R s ;

[0225] In the formula, I s is the sample importance; F s is the usage frequency; T s is the time factor; R s is the effect score; α, β, γ are the weight coefficients. The weight update formula is:

[0226]

[0227] In the formula, W s is the sample weight; λ is the learning rate; is the average importance.

[0228] The specific implementation of step S15 is to output the program code result. Optionally, the code quality evaluation indicators include:

[0229]

[0230] In the formula, Q total is the overall quality score; q i is the single index score; w i is the index weight; n is the number of indicators. The code formatting adopts a unified style specification, including indentation rules, naming rules, comment rules, etc. Finally, a code description document is generated. Optionally, the document integrity scoring formula is:

[0231]

[0232] In the formula, C doc is the document integrity score; d i is the document item score; w iis the project weight; k is the number of document items.

[0233] To better understand and implement the present invention, Example 2 of a specific application scenario of the present invention is provided below: A research team developed an automatic code generation system using the method of the present invention and verified its application in actual projects.

[0234] First, the research team collected 1 million pairs of natural language descriptions and corresponding program codes from multiple sources such as GitHub, Stack Overflow, and programming document libraries. After data cleaning and quality assessment, 500,000 high-quality training samples were finally retained. These samples cover mainstream programming languages such as Python, Java, and C++, and involve multiple application scenarios such as algorithm implementation, data processing, and interface development. After preprocessing these data, an initial training dataset was established, and the basic statistical information of the dataset is shown in Table 1 below:

[0235] Table 1 Basic Statistical Information Table

[0236] Programming Language Number of Samples Average Code Length Average Text Length Pairing Accuracy Python 200000 156 lines 86 words 0.92 Java 150000 203 lines 92 words 0.89 C++ 150000 189 lines 88 words 0.87

[0237] In the word segmentation processing stage, the research team used a conditional random field model for Chinese word segmentation. The training corpus size of the model was 5 million, and the word segmentation accuracy reached 97.8%. For English texts, the NLTK toolkit was used for word segmentation and part-of-speech tagging, with an accuracy of 98.5%. The mathematical expression recognition module used a rule-based method, with an identification accuracy of 95.6%.

[0238] In the feature matrix construction stage, the dimension of the first feature matrix constructed for the mathematical description part was 50000×256. After singular value decomposition, 205 main feature dimensions were retained, and the feature retention rate was 80.2%. The dimension of the second feature matrix constructed for the pure text description part was 50000×1024, and the term frequency-inverse document frequency method was used to calculate the feature weights. The evaluation results of the feature extraction effect are shown in Table 2 below:

[0239] Table 2 Feature Extraction Evaluation Table

[0240] Feature Type Original Dimension Dimension after Dimensionality Reduction Information Retention Rate Computation Time Consumed Mathematical Feature 256 205 0.802 1.2 seconds Text Feature 1024 819 0.785 2.5 seconds

[0241] The specific configuration of the improved bidirectional long short-term memory network includes: input layer dimension 1024, hidden layer dimension 512, output layer dimension 256, number of attention heads 8, dimension of each attention head 64, and number of residual blocks 3. The network training used a batch size of 32, an initial learning rate of 0.001, and 100 training epochs. The main performance indicators are shown in Table 3 below:

[0242] Table 3 Performance Indicator Table

[0243] Metric Name Training Set Performance Validation Set Performance Test Set Performance Accuracy 0.886 0.862 0.853 Recall 0.892 0.875 0.867 F1 Score 0.889 0.868 0.860

[0244] In the actual application of the accuracy determination function, the cosine similarity weight is set to 0.6, the edit distance weight is set to 0.4, and the correction term value is 0.05. The calculation process of the function uses the aforementioned formula: S total = 0.6·S cos + 0.4·S edit + 0.05.

[0245] In terms of video memory optimization configuration, the upper limit of video memory usage is set to 90%. When the usage rate exceeds this threshold, the video memory recycling mechanism is triggered. The training parameter sharding storage uses 32 shards, and the size of each shard is about 256MB. The pipeline depth of gradient calculation is 4, and the load balancing factor is controlled within 1.15. These optimization measures reduce the peak video memory occupancy during the model training process from 32GB to 18GB, improving the utilization efficiency of computing resources.

[0246] In the actual application test, the research team selected 100 actual project requirements for testing, and these requirements cover different difficulty levels and function types. The test results show that for simple requirements (the code volume is less than 100 lines), the code generation accuracy rate of the system reaches 92%; for medium-complexity requirements (the code volume is 100 - 500 lines), the accuracy rate is 85%; for complex requirements (the code volume is more than 500 lines), the accuracy rate is 76%. The specific test data is shown in Table 4 below:

[0247] Table 4 Test Data Table

[0248] Requirement Complexity Number of Samples Generation Correct Rate Average Generation Time Memory Occupancy Simple Requirement 40 0.92 2.5 seconds 1.2 GB Medium Requirement 35 0.85 5.8 seconds 2.8 GB Complex Requirement 25 0.76 12.3 seconds 4.5 GB

[0249] Compared with the traditional natural language generation program code method, the present invention has the following significant advantages:

[0250] 1. The traditional method usually uses a single sequence-to-sequence model and cannot effectively handle the differences between mathematical descriptions and pure text descriptions. However, the present invention improves the accuracy of feature extraction by separately constructing feature matrices and fusing them, and the code generation accuracy rate is increased by 15% on average.

[0251] 2. The traditional method has information loss problems when processing long sequence inputs. The present invention significantly improves the ability to capture long-distance dependency relationships through an improved bidirectional long short-term memory network and an attention mechanism. For code generation tasks with more than 500 lines of code, the accuracy rate is increased by 23%.

[0252] 3. The traditional method often ignores the video memory optimization problem, resulting in too high video memory occupancy during training. However, the present invention realizes the dynamic scheduling and optimization of video memory by improving the zero-redundancy optimizer technology, reducing the peak video memory occupancy by 43.75%.

[0253] 4. The traditional methods lack a comprehensive evaluation mechanism for the quality of the generated code. By establishing a multi-dimensional utility degree evaluation function, the present invention achieves an accurate evaluation of the code quality, and the evaluation accuracy rate is increased by 28%.

[0254] 5. The training data update mechanism of the traditional methods is relatively simple. By means of a dynamic adjustment strategy for sample weights, the present invention improves the continuous learning ability of the model, and the improvement speed of the model performance with the increase of training data is increased by 35%.

[0255] This embodiment fully verifies the practical value of the present invention in the field of natural language generation of computer programming languages. It not only significantly improves the accuracy and efficiency of code generation, but also greatly reduces the consumption of computing resources, providing a new solution for the technological development of related fields.

[0256] It should be noted that the detailed explanations of the variables involved in the present invention are shown in Tables 5 and 6 below.

[0257] Table 5 Variable Explanation Table (First Part)

[0258]

[0259]

[0260] Table 5 Variable Explanation Table (First Part)

[0261]

[0262] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention.

Claims

1. A method for generating a computer programming language from natural language, characterized in that, It includes the following steps: Obtain a natural language input text, establish a language model training dataset, and perform word segmentation processing on the natural language input text; Adopt an improved bidirectional long short-term memory network for feature extraction. The improved bidirectional long short-term memory network sets an attention mechanism module between the input layer and the hidden layer to capture long-distance dependencies in the input sequence, and sets a residual connection module between the hidden layer and the output layer to alleviate the difficulty of deep network training. Establish an accuracy determination function. The accuracy determination function adopts a linear combination calculation method of cosine similarity and edit distance. Establish an accuracy tendency determination function to evaluate the matching degree between the generated code and the expected function. Establish a utility evaluation function for comprehensive evaluation and calculation based on code running time, memory occupancy, and function completion degree. Construct a feature mapping model to generate a program code syntax tree, perform video memory call optimization, optimize the structure of the program code syntax tree, and output the target program code.

2. The method for generating a computer programming language from a natural language according to claim 1, wherein, The step of obtaining the natural language input text is specifically to construct a text preprocessing module to receive the natural language text input by the user, and clean the text, including removing special characters, unifying the encoding format, and correcting obvious grammar errors. Establish a language model training data collection system to collect paired data of natural language descriptions and corresponding program codes from open-source code platforms, program design documents, and technical forums; Use a text similarity algorithm to deduplicate and screen the samples in the dataset to ensure the quality and diversity of the dataset.

3. The method for generating a computer programming language from a natural language according to claim 1, wherein The step of word segmentation processing is specifically to use a word segmentation algorithm based on conditional random fields to segment the input text; use a part-of-speech tagging module to perform part-of-speech tagging on the word segmentation results, and use a hidden Markov model for part-of-speech analysis. Establish a mathematical expression recognition module to recognize the mathematical description part in the text based on regular expressions and rule matching methods, including mathematical formulas, numerical calculations, and logical operation contents.

4. The method for generating a computer programming language from a natural language according to claim 1, wherein The improved bidirectional long short-term memory network is specifically an improvement based on the standard long short-term memory network. Add a multi-head attention mechanism between the input layer and the hidden layer, with the number of attention heads set to 8 and the dimension of each attention head set to 64. Add a residual connection between the hidden layer and the output layer, with the number of residual blocks set to 3, and each residual block contains two fully connected layers. Set the batch size to 32, the initial value of the learning rate to 0.001, and adopt a cosine annealing strategy to dynamically adjust the learning rate.

5. The method for generating a computer programming language from a natural language according to claim 1, wherein The accuracy determination function is specifically to establish a feature vector normalization module to normalize the fused feature matrix; construct a multi-dimensional similarity calculation module, including two sub-modules of vector cosine similarity and edit distance. Cosine similarity is used to calculate the angular similarity between feature vectors, and edit distance is used to calculate the structural similarity of text sequences. Set the weight of cosine similarity to 0.6 and the weight of edit distance to 0.

4.

6. The method for generating a computer programming language from a natural language according to claim 1, wherein The steps of constructing the feature mapping model specifically adopt an encoder-decoder architecture. The encoder uses a multi-layer Transformer network to process the fused feature matrix, and the decoder uses an attention-based recurrent neural network to generate syntax tree nodes. In the encoding process, positional encoding technology is used to maintain the position information of the feature sequence. The number of encoder layers is set to 6, and each layer contains a multi-head self-attention mechanism and a feed-forward neural network. The attention temperature parameter is set to 0.

7.

7. The method for generating a computer programming language from a natural language according to claim 1, wherein The utility degree evaluation function specifically uses a code running time evaluation module, a memory occupancy evaluation module, and a function completion degree evaluation module to evaluate the time complexity, space complexity, and function coverage rate of the code respectively. The weights of each evaluation module are set to 0.3, 0.3, and 0.4 respectively, and the weighted average is calculated to obtain the final utility degree score.

8. The method for generating a computer programming language from natural language according to claim 1, wherein The video memory call optimization specifically dynamically allocates video memory resources according to the task priority and the utility degree score. The improved zero-redundancy optimizer technology is adopted for video memory optimization, including that the training parameter sharding storage module stores the model parameters dispersedly in different video memory areas, the gradient calculation distributed processing module processes the gradient calculation in a pipeline parallel manner, and the video memory dynamic recycling module releases useless intermediate variables after each training step.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions, and when the program instructions run on a computer, they are used to execute the method for a natural language generation computer programming language according to any one of claims 1-8.

10. A system for generating a computer programming language from natural language, characterized in that, It includes the computer-readable storage medium according to claim 9. The system is any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is set inside the system, and a microprocessor for executing the program instructions stored in the computer-readable storage medium is set inside the system.