A method for establishing a mathematical description language large model, a medium and a system

By constructing a multi-level mathematical logic reasoning network model and a contribution value evaluation memory optimization strategy, the problem of unreasonable memory allocation in the training of large models of mathematical description languages ​​is solved, realizing efficient utilization of memory resources and improving model training efficiency.

CN120181223BActive Publication Date: 2025-11-18WUHAN TECHN COLLEGE OF COMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510228132.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-11-18
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

In existing technologies, unreasonable allocation of video memory during the training of large models using mathematical description languages ​​leads to low utilization of computing resources, and existing video memory management methods have failed to effectively solve the problem of low utilization of video memory resources.

Method used

By constructing a mathematical logic reasoning network model that includes a semantic understanding layer, a mathematical symbol mapping layer, a logical relation transformation layer, a mathematical unit organization layer, and a unit relation construction layer, and combining a memory optimization allocation strategy based on contribution values, the gradient backpropagation algorithm is used to quantify and evaluate the contribution weight of each layer, and the memory allocation strategy, computing thread allocation, and training batch are dynamically adjusted to achieve efficient utilization of memory resources.

Benefits of technology

It achieves efficient utilization of computing resources, avoids memory fragmentation and resource waste, improves model training efficiency, and ensures the prediction accuracy of mathematical description language conversion tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181223B_ABST
    Figure CN120181223B_ABST
Patent Text Reader

Abstract

The application provides a mathematical description language large model establishment method, medium and system, belonging to the technical field of computer model, which first acquires corpus information through constructing a training data set, then carries out word segmentation and feature extraction on natural language description text, constructs an inexact description matrix, and carries out structured processing on standard mathematical description language text to form an accurate description matrix. On this basis, a display memory pool is established and training resources are pre-allocated, and a multi-level mathematical logic reasoning network model is constructed. Through contribution value evaluation and accuracy calculation, resource optimization allocation based on importance is realized, and a display memory pool management mechanism is adopted to ensure efficient use of resources. Through the quantitative evaluation mechanism and dynamic optimization strategy, the efficient training of the mathematical description language conversion model is realized, and the technical problem of low utilization of computing resources caused by unreasonable allocation of display memory in the training process of the mathematical description language large model in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer modeling technology, and more specifically, relates to a method, medium, and system for establishing large models using mathematical description languages. Background Technology

[0002] In the fields of artificial intelligence and natural language processing, converting natural language descriptions into standard mathematical description languages ​​is a crucial task, significant for intelligent education, intelligent tutoring, and scientific research. Traditional methods for converting mathematical description languages ​​mainly include rule-based and statistical methods. Rule-based methods map natural language to mathematical expressions by manually designing grammatical rules and conversion templates. This method requires a large number of manually written rules and struggles to handle complex mathematical concepts and logical relationships. Statistical methods employ machine learning algorithms to learn conversion models from large-scale corpora, including Conditional Random Fields (CRF) and Support Vector Machines (SVM). While these methods possess some adaptability, they require high-quality and high-quantity training data and struggle to capture deep semantic information.

[0003] With the development of deep learning technology, neural network-based mathematical description language conversion models have gradually become a research hotspot. These methods achieve end-to-end conversion from natural language to mathematical descriptions by constructing deep neural networks. However, the training process of deep learning models requires substantial computational resources, especially GPU memory. Existing training methods typically employ fixed GPU memory allocation strategies, failing to consider the varying importance of different model layers, leading to inefficient use of GPU memory. Furthermore, because mathematical description languages ​​have strict grammatical rules and logical constraints, models need to simultaneously handle multiple tasks such as semantic understanding, symbol mapping, and logical relation construction, and these tasks have significantly different computational resource requirements.

[0004] Current memory management methods mainly fall into two categories: static allocation and dynamic allocation. Static allocation pre-allocates a fixed amount of memory to each layer of the model. While simple to implement, it struggles to adapt to changing resource demands during training, easily leading to memory waste or insufficiency. Dynamic allocation adjusts memory allocation based on demand during training, but frequent memory allocation and deallocation operations incur additional computational overhead, impacting training efficiency. Furthermore, existing methods lack a quantitative evaluation mechanism for the contribution of each model layer, failing to achieve importance-based resource optimization and thus struggling to address the problem of low memory utilization. In other words, current technologies suffer from the technical problem of low computational resource utilization during the training of large mathematical description language models due to unreasonable memory allocation. Summary of the Invention

[0005] In view of this, the present invention provides a method, medium and system for building large models of mathematical description languages, which can solve the technical problem of low utilization of computing resources due to unreasonable allocation of video memory during the training of large models of mathematical description languages ​​in the prior art.

[0006] The present invention is implemented as follows: The first aspect of the present invention provides a method for establishing a large mathematical description language model, comprising the following steps: acquiring natural language description texts and their corresponding standard mathematical description language texts from a training corpus to establish a training dataset; establishing a memory pool and pre-allocating memory blocks required for training, dividing the memory blocks into basic memory blocks and dynamic memory blocks, wherein the basic memory blocks are used to store fixed parameters and the dynamic memory blocks are used to store intermediate calculation results; constructing a mathematical logic reasoning network model including a semantic understanding layer, a mathematical symbol mapping layer, a logical relation transformation layer, a mathematical unit organization layer, and a unit relation construction layer; establishing a unit contribution value evaluation model based on the accuracy score of the verification prediction results, and calculating the contribution weight of each layer using a gradient backpropagation algorithm; determining the memory allocation amount, the number of computation threads, and the number of training batches according to the contribution weights calculated by the unit contribution value evaluation model; optimizing the memory allocation of the memory pool according to the memory allocation amount until the accuracy score of the verification prediction results reaches a preset accuracy threshold.

[0007] Specifically, the steps for establishing the training dataset involve constructing a training corpus by building a mathematical knowledge graph, extracting mathematical concepts, theorems, formulas, and problem-solving methods from mathematical textbooks, papers, and professional literature, constructing a set of mathematical knowledge nodes, establishing inclusion, inheritance, combination, and transformation relationships between mathematical knowledge nodes, and storing the mathematical knowledge nodes and their relationships in a structured manner.

[0008] The preset accuracy thresholds include a preset semantic accuracy threshold, a preset symbol accuracy threshold, a preset structural accuracy threshold, and a preset logical accuracy threshold, wherein the preset semantic accuracy threshold is 0.85, the preset symbol accuracy threshold is 0.9, the preset structural accuracy threshold is 0.8, and the preset logical accuracy threshold is 0.85.

[0009] The semantic understanding layer uses a bidirectional long short-term memory network to extract the contextual semantic information of the text; the mathematical symbol mapping layer uses an attention mechanism to realize the mapping and conversion of semantic features to mathematical symbols; the logical relationship conversion layer uses a graph neural network to establish the logical relationship between mathematical symbols; the mathematical unit organization layer uses a recurrent neural network to organize mathematical logical units; and the unit relationship construction layer uses a tree-type recurrent neural network to establish the hierarchical relationship between mathematical logical units.

[0010] The total size of the video memory block is determined based on the size of the non-precise description matrix, the size of the precise description matrix, and the preset intermediate calculation result storage requirements. The basic video memory block accounts for 60% of the total video memory, and the dynamic video memory block accounts for 40% of the total video memory.

[0011] The method for determining the contribution weights is as follows: the semantic understanding layer mainly affects semantic accuracy, the mathematical symbol mapping layer mainly affects symbol accuracy, the logical relation transformation layer mainly affects logical accuracy, and the mathematical unit organization layer and the unit relation construction layer mainly affect structural accuracy.

[0012] The method for determining the amount of video memory allocated is as follows: the greater the contribution weight, the more video memory is allocated, and the specific ratio is the square root of the contribution weight; the method for determining the number of computing threads is as follows: the greater the contribution weight, the more computing threads are allocated, and the specific number is the base number of threads multiplied by the contribution weight; the method for determining the number of training batches is as follows: the greater the contribution weight, the more training batches are generated, and the specific number of batches is the base number of batches multiplied by the contribution weight.

[0013] The video memory allocation optimization includes merging free video memory in the video memory block, reclaiming idle video memory in the video memory block, reallocating the priority of the video memory block according to the contribution weight, and updating the usage status of the video memory block, wherein the usage status includes usage time, access frequency and idle ratio.

[0014] A second aspect of the present invention provides a computer-readable storage medium storing program instructions, which, when executed in a computer, are used to perform the above-described method for establishing a large model of a mathematical description language.

[0015] A third aspect of the present invention provides a system for establishing a large model of a mathematical description language, comprising the aforementioned computer-readable storage medium. The system is any one of a computer, a server, or a microcontroller. The computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor that executes the program instructions stored in the computer-readable storage medium.

[0016] Compared with existing technologies, the present invention provides a method, medium and system for establishing a large model of a mathematical description language. The method for establishing a large model of a mathematical description language proposed in this invention constructs a mathematical logic reasoning network model that includes a semantic understanding layer, a mathematical symbol mapping layer, a logical relation transformation layer, a mathematical unit organization layer and a unit relation construction layer. Combined with a memory optimization allocation strategy based on contribution value, it achieves efficient utilization of computing resources.

[0017] This method first establishes a unit contribution value evaluation model and uses the gradient backpropagation algorithm to quantify the impact of each network layer on model performance. Then, based on the contribution weights, it dynamically adjusts the memory allocation strategy, computation thread allocation, and training batch settings. Through a memory pool management mechanism, it achieves dynamic scheduling and optimized allocation of memory resources, effectively avoiding memory fragmentation and resource waste. Simultaneously, this method employs a multi-layered model structure design, constructing dedicated processing modules tailored to the characteristics of mathematical description languages. Through reasonable resource allocation, it ensures the computational efficiency of each module.

[0018] This invention establishes a quantitative decision-making mechanism for resource allocation by combining memory allocation with model performance contribution, thus solving the problem of low resource utilization efficiency caused by traditional fixed allocation or simple dynamic allocation methods. Simultaneously, through overall management and optimized scheduling of the memory pool, efficient utilization of memory resources is achieved, improving model training efficiency. Based on a strategy of accuracy evaluation and incremental training, the predictive accuracy of the model is guaranteed, providing reliable technical support for mathematical description language conversion tasks. Attached Figure Description

[0019] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0021] like Figure 1 The diagram shown is a flowchart of a method for establishing a large model of a mathematical description language according to the first aspect of this invention. This method includes the following steps:

[0022] S01. Obtain the natural language description text and its corresponding standard mathematical description language text from the training corpus, and establish a training dataset. The training dataset includes the natural language description text, the standard mathematical description language text, and a mathematical logic unit structure. The mathematical logic unit structure includes a set of mathematical symbols, a set of mathematical functions, a set of mathematical operations, a set of mathematical relations, and a set of logical reasoning rules.

[0023] S02. The natural language description text is segmented into words, semantic features, numerical features and logical relationship features are extracted, and an imprecise description matrix is ​​constructed. The imprecise description matrix is ​​composed of a semantic vector matrix, a numerical feature matrix and a relational feature matrix. The semantic vector matrix is ​​used to represent mathematical concepts in natural language, the numerical feature matrix is ​​used to represent numerical information in natural language, and the relational feature matrix is ​​used to represent logical relationships in natural language.

[0024] S03. The standard mathematical description language text is subjected to structured processing to extract mathematical symbols, operators and mathematical logic units, and a precise description matrix is ​​constructed. The precise description matrix includes a mathematical symbol matrix, an operator matrix and a mathematical logic unit matrix. The mathematical logic unit matrix is ​​used to store the hierarchical relationship and calling relationship between mathematical logic units.

[0025] S04. Establish a memory pool and pre-allocate memory blocks required for training. Determine the total size of the memory blocks based on the size of the inaccurate description matrix, the size of the accurate description matrix, and the preset storage requirements for intermediate calculation results. Divide the memory blocks into basic memory blocks and dynamic memory blocks. The basic memory blocks are used to store the inaccurate description matrix, the accurate description matrix, and fixed parameters. The dynamic memory blocks are used to store intermediate calculation results.

[0026] S05. Construct a mathematical logic reasoning network model, which includes a semantic understanding layer, a mathematical symbol mapping layer, a logical relation transformation layer, a mathematical unit organization layer, and a unit relation construction layer.

[0027] S06. Input the inaccurate description matrix into the mathematical logic reasoning network model. Convert the semantic vector matrix into a mathematical concept representation through the semantic understanding layer. Map the mathematical concept representation into a sequence of mathematical symbols through the mathematical symbol mapping layer. Construct the logical relationship between the mathematical symbol sequences through the logical relationship conversion layer. Organize the mathematical symbol sequences and the logical relationship into mathematical logic units through the mathematical unit organization layer. Establish the hierarchical relationship and calling relationship between the mathematical logic units through the unit relationship construction layer. Generate a mathematical logic unit tree structure and predicted mathematical description text.

[0028] S07. Calculate the accuracy score between the predicted mathematical description text and the standard mathematical description language text. The accuracy score includes semantic accuracy score, symbolic accuracy score, structural accuracy score and logical accuracy score. Optimize and train the mathematical logic reasoning network model based on the accuracy score.

[0029] S08. Obtain the natural language description text from the validation dataset, and repeat steps S02 and S06 to obtain the validation prediction results.

[0030] S09. Calculate the accuracy score of the verification prediction result. When the accuracy score of the verification prediction result is lower than the preset accuracy threshold, add the verification dataset to the training dataset for incremental training. The preset accuracy threshold includes a preset semantic accuracy threshold, a preset symbol accuracy threshold, a preset structural accuracy threshold, and a preset logical accuracy threshold.

[0031] S10. Based on the accuracy score of the verification prediction results, establish a unit contribution value evaluation model, and use the gradient backpropagation algorithm to calculate the contribution weights of the semantic understanding layer, the mathematical symbol mapping layer, the logical relation transformation layer, the mathematical unit organization layer and the unit relation construction layer.

[0032] S11. Based on the contribution weights calculated by the unit contribution value evaluation model, determine the memory allocation, number of computing threads, and number of training batches for the semantic understanding layer, the mathematical symbol mapping layer, the logical relation conversion layer, the mathematical unit organization layer, and the unit relation construction layer.

[0033] S12. Optimize the memory allocation of the memory pool according to the memory allocation amount. The memory allocation optimization includes merging the free memory in the memory block, reclaiming the idle memory in the memory block, reallocating the priority of the memory block according to the contribution weight, and updating the usage status of the memory block until the accuracy score of the verification prediction result reaches the preset accuracy threshold, thereby obtaining the trained mathematical description language model.

[0034] The mathematical logic unit tree structure includes unit type identifier, input parameter information, output result information, calling method information, and constraint information. The unit type identifier is used to identify the functional category of the mathematical logic unit. The input parameter information is used to receive data input. The output result information is used to return the calculation result. The calling method information is used to define the interaction interface between the mathematical logic units. The constraint information is used to limit the value range of the input parameter information.

[0035] The specific implementation methods of the above steps are described in detail below. Step S01 involves constructing a training corpus by establishing a mathematical knowledge graph. First, mathematical concepts, theorems, formulas, and problem-solving methods are extracted from mathematics textbooks, papers, and professional literature to construct a set of mathematical knowledge nodes. Next, the relationships between these mathematical knowledge nodes are established, including inclusion, inheritance, combination, and transformation relationships between concepts. Then, the mathematical knowledge nodes and their relationships are stored in a structured manner to form a mathematical knowledge graph. Based on this graph, mathematical problems described in natural language and their standard mathematical expressions are collected from online resources, teaching materials, and exam question banks to establish an initial corpus. The initial corpus is cleaned to remove duplicate, incomplete, and erroneous samples. Finally, the cleaned corpus is divided into a training set and a validation set in an 8:2 ratio. The training set is used for model training, and the validation set is used for model evaluation. The purpose of this step is to construct a high-quality training dataset, providing a reliable data foundation for subsequent model training.

[0036] The specific implementation of step S02 involves first using the Jieba word segmentation tool to segment the natural language description text, obtaining a word sequence; then using a pre-trained word vector model, word2vec, to extract semantic features from the word sequence, mapping each word to a 300-dimensional word vector to construct a semantic vector matrix; next, using regular expressions to extract numerical information from the text, including integers, decimals, fractions, percentages, etc., and normalizing the numerical information to construct a numerical feature matrix; finally, using dependency parsing techniques to extract logical relationships from the word sequence, including causal relationships, conditional relationships, and parallel relationships, encoding the logical relationships into a relational feature matrix. The purpose of this step is to convert the natural language description text into a computer-processable feature representation, providing input for subsequent mathematical description conversion.

[0037] The specific implementation of step S03 involves first performing lexical analysis on the standard mathematical description language text to identify mathematical symbols, operators, keywords, and identifiers; then performing syntactic analysis to construct an abstract syntax tree representing the hierarchical structure of mathematical expressions; next, performing semantic analysis to extract the dependencies and calling relationships between mathematical logic units; and finally, organizing the analysis results into a precise description matrix, where the mathematical symbol matrix stores symbol type and attribute information, the operator matrix stores operation priority and associativity information, and the mathematical logic unit matrix stores the hierarchical relationships and calling relationships between units. The purpose of this step is to convert the standard mathematical description language text into a structured matrix representation, providing a target output for model training.

[0038] The specific implementation of step S04 involves first estimating the GPU memory requirements during training based on the size of the training dataset and the model structure, and calculating the total size of the required GPU memory blocks. Then, the GPU memory blocks are divided into basic GPU memory blocks and dynamic GPU memory blocks. The basic GPU memory blocks account for 60% of the total GPU memory and are used to store model parameters and fixed data, while the dynamic GPU memory blocks account for 40% of the total GPU memory and are used to store intermediate calculation results. Next, a GPU memory pre-allocation strategy is adopted to allocate the GPU memory space required for training in advance, avoiding frequent GPU memory allocation and release during training. Finally, a GPU memory management mechanism is established to achieve dynamic scheduling and reclamation of GPU memory blocks. The purpose of this step is to optimize GPU memory usage efficiency and improve model training speed.

[0039] The specific implementation of step S05 involves first constructing a semantic understanding layer, using a bidirectional long short-term memory network to extract contextual semantic information from the text; then constructing a mathematical symbol mapping layer, using an attention mechanism to map semantic features to mathematical symbols; next, constructing a logical relationship transformation layer, using a graph neural network to establish logical relationships between mathematical symbols; then, constructing a mathematical unit organization layer, using a recurrent neural network to organize mathematical symbols and their logical relationships into mathematical logical units; and finally, constructing a unit relationship construction layer, using a tree-type recurrent neural network to establish hierarchical relationships between mathematical logical units. The purpose of this step is to construct a neural network model capable of understanding natural language and generating mathematical descriptions.

[0040] The specific implementation of step S06 involves first inputting the imprecise description matrix into the mathematical logic reasoning network model, and then extracting deep semantic features of the text through the semantic understanding layer. Next, the semantic features are converted into a sequence of mathematical symbols through the mathematical symbol mapping layer, and the mapping accuracy is optimized using a cross-entropy loss function. Then, the logical relationship transformation layer constructs the logical relationships between mathematical symbols, and the relationship classification loss function is used to optimize the relationship recognition accuracy. Next, the mathematical unit organization layer organizes the mathematical symbols and their logical relationships into mathematical logical units, and the structural similarity loss function is used to optimize the organization's rationality. Finally, the unit relationship construction layer generates a mathematical logical unit tree structure, and the tree distance loss function is used to optimize the tree structure's correctness. The purpose of this step is to realize the conversion process from natural language to mathematical description.

[0041] The specific implementation of step S07 involves first calculating the semantic accuracy score, using cosine similarity to calculate the semantic similarity between the predicted text and the standard text, with a semantic accuracy threshold set to 0.85; then calculating the symbol accuracy score, using edit distance to calculate the matching degree between the predicted symbol sequence and the standard symbol sequence, with a symbol accuracy threshold set to 0.9; next, calculating the structural accuracy score, using tree structure similarity to calculate the consistency between the predicted tree structure and the standard tree structure, with a structural accuracy threshold set to 0.8; and finally, calculating the logical accuracy score, using a graph matching algorithm to calculate the correspondence between the predicted logical relationship and the standard logical relationship, with a logical accuracy threshold set to 0.85. The purpose of this step is to evaluate the accuracy of the model's prediction results and provide guidance for model optimization.

[0042] The specific implementation of step S08 involves first obtaining the natural language description text from the validation dataset; then repeating step S02 to segment the validation text, extract semantic features, numerical features, and logical relationship features, and construct an imprecise description matrix; next, repeating step S06 to input the imprecise description matrix into the trained mathematical logic reasoning network model to generate a mathematical logic unit tree structure and predicted mathematical description text. The purpose of this step is to verify the model's generalization ability and test its performance on unseen data.

[0043] The specific implementation of step S09 involves first calculating the accuracy scores of the verification prediction results, including semantic accuracy, symbolic accuracy, structural accuracy, and logical accuracy; then comparing the accuracy scores with preset thresholds, where the semantic accuracy threshold is 0.85, the symbolic accuracy threshold is 0.9, the structural accuracy threshold is 0.8, and the logical accuracy threshold is 0.85; if any accuracy score is lower than the corresponding threshold, the verification sample is added to the training dataset; next, the expanded training dataset is retrained, and the model parameters are optimized using the gradient descent algorithm. The purpose of this step is to improve the prediction accuracy of the model through incremental training.

[0044] The specific implementation of step S10 involves first constructing a unit contribution value evaluation model using a multilayer perceptron structure. The input is the intermediate features of each layer, and the output is the accuracy score. Then, the gradient backpropagation algorithm is used to calculate the contribution of each layer to the accuracy score. Specifically, the semantic understanding layer primarily affects semantic accuracy, the mathematical symbol mapping layer primarily affects symbolic accuracy, the logical relation transformation layer primarily affects logical accuracy, and the mathematical unit organization layer and unit relation construction layer primarily affect structural accuracy. Next, the contribution weight of each layer is determined based on the magnitude of the gradient value. The purpose of this step is to quantify the importance of each network layer, providing a basis for resource allocation.

[0045] The specific implementation of step S11 involves first determining the memory allocation for each layer based on its contribution weight. A higher contribution weight results in a larger memory allocation, specifically calculated as the square root of the contribution weight. Next, the number of computation threads for each layer is determined based on the contribution weight. A higher contribution weight results in a larger number of computation threads, specifically the base number of threads multiplied by the contribution weight. Finally, the number of training batches for each layer is determined based on the contribution weight. A higher contribution weight results in a larger number of training batches, specifically the base number of batches multiplied by the contribution weight. The purpose of this step is to rationally allocate computational resources according to the importance of each layer, thereby improving training efficiency.

[0046] The specific implementation of step S12 involves first defragmenting the memory blocks in the memory pool, merging adjacent free memory blocks; then reclaiming memory blocks that have not been used for a long time and marking them as available; next, prioritizing the memory blocks according to their contribution weights, with layers having higher contribution weights receiving memory resources first; and finally, periodically updating the usage status of the memory blocks, including usage time, access frequency, and idle ratio. The purpose of this step is to optimize the efficiency of memory resource utilization and ensure the smooth progress of model training. This step continues until all accuracy scores of the verification prediction results reach preset thresholds, indicating that model training is complete.

[0047] The mathematical logic unit tree structure in this method adopts an object-oriented design approach, with each mathematical logic unit containing a complete functional definition and interface specification. Unit type identifiers distinguish different types of mathematical operations or functions, facilitating targeted processing by the model; input parameter information defines the data types and format requirements that the unit can accept, ensuring the standardization of data input; output result information specifies the return format of the calculation result, facilitating data retrieval by subsequent units; calling method information defines the interaction rules between units, ensuring the correctness of data flow; and constraint information limits the range of input parameter values, guaranteeing the validity of the calculation results. This structured design achieves modular organization and flexible invocation of mathematical logic units.

[0048] The inaccurate description matrix includes a semantic vector matrix, a numerical feature matrix, and a relational feature matrix, wherein the semantic vector matrix is ​​specifically represented as follows:

[0049]

[0050] In the formula, s ij Let be the vector value of the i-th word in the j-th semantic dimension; m is the word sequence length; n is the semantic vector dimension, with a value of 300; w k v is the weight coefficient of the k-th context word; k Let be the word vector of the k-th context word; K is the size of the context window; and α is the semantic adjustment factor.

[0051] The derivation of the semantic vector matrix equation in the inaccurate description matrix is ​​as follows:

[0052] 1. First, basic word vectors are obtained based on the Word2Vec model. Initial vector representations of each word are obtained by training on a large-scale corpus.

[0053] 2. Considering the influence of contextual semantic information, a window size K is introduced, with K ranging from 3 to 7. The optimal value is determined through grid search.

[0054] 3. To reflect the importance of contextual words in different positions, a weighting coefficient w is introduced. k The calculation formula is as follows:

[0055]

[0056] In the formula, k is the position of the context word, c is the position of the target word, and ∈ is the smoothing factor, which ranges from 0.01 to 0.1.

[0057] 4. Finally, a semantic adjustment factor α is introduced to optimize the vector representation, which is then optimized using the backpropagation algorithm.

[0058] The numerical feature matrix is ​​specifically represented as follows:

[0059]

[0060] In the formula, n ij p is the value of the i-th value in the j-th feature dimension; p is the number of values; q is the feature dimension; x t Let t be the t-th numerical sample; μ be the numerical mean; T be the total number of samples; and β be the numerical adjustment factor.

[0061] The principle behind establishing the numerical characteristic matrix equation is as follows:

[0062] 1. Standardize the original values ​​to eliminate the influence of dimensions;

[0063] 2. Calculate the statistical characteristics of the numerical sample, including mean, variance, kurtosis, skewness, etc.

[0064] 3. Variance is used to characterize the distribution of numerical values. Variance is chosen over standard deviation because variance is more sensitive to outliers.

[0065] 4. A numerical adjustment factor β is introduced to adjust the feature distribution, and its value is determined by minimizing the reconstruction error.

[0066] The relation feature matrix is ​​specifically represented as follows:

[0067]

[0068] In the formula, r ij Let u be the feature value representing the relationship between the i-th word and the j-th word; u and v are the word sequence lengths; c l d represents the weight coefficient of the l-th type of relation; l is the feature vector of the l-th type of relation; is the total number of relation types; γ is the relation adjustment factor.

[0069] The principle behind establishing the relational characteristic matrix equation is as follows:

[0070] 1. Identify the types of relationships between words based on dependency parsing;

[0071] 2. Encode the relation type as a vector representation using one-hot encoding;

[0072] 3. Introduce the relation weight coefficient c l This indicates the importance of different types of relationships, and its calculation formula is:

[0073]

[0074] In the formula, f l Let l be the frequency of the l-th type relation in the training set;

[0075] 4. Finally, the feature representation is optimized using the relationship adjustment factor γ.

[0076] The precise description matrix includes a mathematical symbol matrix, an operator matrix, and a mathematical logic unit matrix, wherein the mathematical symbol matrix is ​​specifically represented as follows:

[0077]

[0078] In the formula, m ij Let f be the value of the i-th symbol in the j-th feature dimension; a is the length of the symbol sequence; b is the feature dimension; f h g represents the weighting coefficient for the h-th class of symbols; h denoted as the feature vector of the h-th symbol; H represents the total number of symbol types; and δ is the symbol adjustment factor.

[0079] The mathematical symbols in the matrix are precisely described, and the derivation of the matrix equation is as follows:

[0080] 1. Construct a dictionary of mathematical symbols, containing commonly used mathematical symbols and their attribute information;

[0081] 2. Encode symbolic attribute information into a vector representation;

[0082] 3. Introduce the symbol weighting coefficient f h This reflects the frequency of use of different types of symbols;

[0083] 4. Optimize the symbol representation by adjusting the symbol adjustment factor δ.

[0084] The specific representation of the operator matrix is ​​as follows:

[0085]

[0086] In the formula, o ij Let be the value of the i-th operator on the j-th feature dimension; e is the length of the operator sequence; f is the feature dimension; q p z represents the weight coefficient for the p-th type of operation; p Let be the feature vector of the p-th type of operation; P be the total number of operation types; and θ be the operation adjustment factor.

[0087] The principle behind establishing the equation for the operator matrix is ​​as follows:

[0088] 1. Define the operator precedence system, including arithmetic operations, exponentiation, function operations, etc.;

[0089] 2. Encode operator attributes as vector representations;

[0090] 3. Introduce the calculation weight coefficient q p This reflects the complexity of different types of operations;

[0091] 4. Optimize the computation representation by adjusting the factor θ.

[0092] The mathematical logic unit matrix is ​​specifically represented as follows:

[0093]

[0094] In the formula, u ij The relationship value between the i-th logic unit and the j-th logic unit; g and h are the number of logic units; k y Let l be the weight coefficient of the y-th type of relation; y Let be the feature vector of the y-th type of relation; Y be the total number of relation types; and λ be the unit adjustment factor.

[0095] The principle behind establishing the mathematical logic unit matrix equation is as follows:

[0096] 1. Construct a relationship diagram between logical units;

[0097] 2. Encode relation types as vector representations;

[0098] 3. Introduce the relation weight coefficient k y This reflects the importance of different relationship types;

[0099] 4. The relationship is optimized by using the unit adjustment factor λ.

[0100] The formula for calculating the accuracy score is expressed as follows:

[0101] Semantic precision score:

[0102] In the formula, A s Score for semantic precision; x i and y i denoted as the i-th semantic vector of the predicted text and the standard text, respectively; n is the dimension of the semantic vector; and η is the semantic accuracy adjustment factor.

[0103] Symbol precision score:

[0104] In the formula, A m Score for symbol precision; D i L is the edit operand at position i; p and L s ξ represents the lengths of the predicted sequence and the standard sequence, respectively; ξ is the sign precision adjustment factor.

[0105] Structural accuracy score:

[0106] In the formula, A t Score for structural accuracy; N cN represents the set of common nodes for the predicted tree structure and the standard tree structure. p and N s ω represents the node sets of the predicted tree structure and the standard tree structure, respectively; ω is the structure accuracy adjustment factor.

[0107] Logical precision score:

[0108] In the formula, A l Score for logical precision; E c E is the set of common edges between predicted logical relations and standard logical relations. p and E s φ represents the edge set of the predicted logical relation and the standard logical relation, respectively; φ is the logical precision adjustment factor.

[0109] In the accuracy score calculation, the semantic accuracy score is calculated using cosine similarity. Cosine similarity is chosen because it is not sensitive to vector length and can better reflect the degree of semantic similarity. The symbolic accuracy score is calculated based on edit distance, and normalization is applied to make the score range between 0 and 1. The structural accuracy score uses the Dice coefficient to calculate tree structure similarity. This coefficient is symmetric and insensitive to sample imbalance. The logical accuracy score uses the Jaccard coefficient to calculate graph structure similarity. This coefficient can effectively measure the degree of overlap between two sets.

[0110] The specific formula for calculating the contribution weight is as follows:

[0111]

[0112] In the formula, W i The contribution weights of the i-th layer are: L is the total loss function; O i ψ represents the output of the i-th layer; ψ is the contribution weight adjustment factor.

[0113] The contribution weight calculation adopts a method based on the gradient of the loss function, and the specific steps are as follows:

[0114] 1. Calculate the partial derivatives of the loss function with respect to the output of each layer;

[0115] 2. Calculate the relative contribution of each layer's output to the loss function;

[0116] 3. Introduce a contribution weight adjustment factor for optimization;

[0117] This equation can quantify the impact of each network layer on the model's performance.

[0118] The specific formula for calculating the video memory allocation is as follows:

[0119]

[0120] In the formula, M i M represents the memory allocation for the i-th layer; base Basic memory size; W i σ represents the contribution weight of the i-th layer; σ is the memory allocation adjustment factor.

[0121] The formula for calculating the number of threads is expressed as follows:

[0122] T i =T base ×W i +τ;

[0123] In the formula, T i T represents the number of computation threads in the i-th layer; base Base number of threads; W i τ represents the contribution weight of the i-th layer; τ is the thread number adjustment factor.

[0124] The formula for calculating the number of training batches is expressed as follows:

[0125] B i =B base ×W i +ρ;

[0126] In the formula, B i B is the number of training batches for the i-th layer; base Basic batch number; W i ρ represents the contribution weight of the i-th layer; ρ is the batch number adjustment factor.

[0127] The equations for calculating memory allocation, the number of computation threads, and the number of training batches are all designed based on contribution weights. The memory allocation uses a square root relationship to reduce uneven resource distribution caused by weight differences; the number of computation threads and the number of training batches use a linear relationship to highlight the priority of important layers. The purpose of these equations is to achieve a reasonable allocation of computational resources and improve training efficiency.

[0128] The values ​​of each adjustment factor range from 0.1 to 0.5, and the optimal value is determined through cross-validation; the base memory, base number of threads, and base number of batches are determined according to the hardware configuration.

[0129] The values ​​of all adjustment factors were determined through cross-validation, and the specific steps are as follows:

[0130] 1. Set the candidate value range of the adjustment factor to 0.1 to 0.5, with a step size of 0.1;

[0131] 2. Perform a grid search on the validation set;

[0132] 3. Select the combination of adjustment factor values ​​that optimizes model performance.

[0133] The methods for determining each basic parameter are as follows:

[0134] 1. The base video memory size is determined based on 20% of the GPU video memory size;

[0135] 2. The base number of threads is determined based on 50% of the number of CPU cores;

[0136] 3. The base batch size is determined based on 10% of the training dataset size.

[0137] The principle behind constructing the above equations is as follows:

[0138] 1. The non-precise description matrix adopts a weighted summation form, taking into account the influence of contextual information, and introduces an adjustment factor for fine-tuning;

[0139] 2. The precise description matrix also adopts a weighted summation form, taking into account the importance of different types of features, and is optimized through adjustment factors;

[0140] 3. Different similarity measurement methods are used to calculate the accuracy score: cosine similarity is used for semantic accuracy, edit distance is used for symbolic accuracy, tree structure similarity is used for structural accuracy, and graph matching similarity is used for logical accuracy.

[0141] 4. The contribution weights are calculated using the gradient method, taking into account the sensitivity of the loss function to the output of each layer;

[0142] 5. Resource allocation calculation takes into account the impact of contribution weights. The memory allocation adopts a square root relationship to smooth out differences, and the number of threads and batches adopts a linear relationship to highlight their importance.

[0143] A second aspect of the present invention provides a computer-readable storage medium storing program instructions, which, when executed in a computer, are used to perform the above-described method for establishing a large model of a mathematical description language.

[0144] A third aspect of the present invention provides a system for establishing a large model of a mathematical description language, comprising the aforementioned computer-readable storage medium. The system is any one of a computer, a server, or a microcontroller. The computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor that executes the program instructions stored in the computer-readable storage medium.

[0145] Specifically, the principle of this invention is as follows: Based on deep learning and resource optimization theory, the technical solution achieves efficient training for mathematical description language conversion by constructing a hierarchical neural network model and a dynamic resource management mechanism. In terms of model structure design, a multi-layered processing architecture is adopted, with each layer optimized for a specific task. The semantic understanding layer uses a bidirectional long short-term memory network to effectively capture the contextual semantic information of the text; the mathematical symbol mapping layer achieves accurate mapping from semantic features to mathematical symbols through an attention mechanism; the logical relationship conversion layer uses a graph neural network to establish logical relationships between mathematical symbols; and the mathematical unit organization layer and unit relationship construction layer are responsible for constructing a complete mathematical logic system.

[0146] In terms of resource management, this invention innovatively proposes a memory allocation optimization strategy based on contribution values. The contribution of each layer to model performance is calculated using the gradient backpropagation algorithm, establishing a quantitative basis for memory allocation. A square root relationship is used to determine the amount of memory allocated, considering both differences in importance and avoiding excessive resource skewness. A linear relationship is adopted in the allocation of computation threads and training batches, highlighting the priority of important layers and ensuring the execution efficiency of critical computational tasks. Simultaneously, through unified management of the memory pool, dynamic scheduling and optimized allocation of memory resources are achieved, effectively solving the problems of memory fragmentation and resource waste.

[0147] The following provides a specific embodiment 1 of the present invention, and the specific implementation of each step in this embodiment 1 is described in detail below.

[0148] The specific implementation of step S01 involves constructing a training corpus by establishing a mathematical knowledge graph. First, mathematical concepts, theorems, formulas, and problem-solving methods are extracted from mathematics textbooks, papers, and professional literature to construct a set of mathematical knowledge nodes. Next, the relationships between these mathematical knowledge nodes are established. Then, the mathematical knowledge nodes and their relationships are stored in a structured manner to form a mathematical knowledge graph. Based on this graph, mathematical problems described in natural language and their standard mathematical expressions are collected from online resources, teaching materials, and exam question banks to establish an initial corpus. The initial corpus is then cleaned to remove duplicate, incomplete, and erroneous samples. Finally, the cleaned corpus is divided into a training set and a validation set in an 8:2 ratio. The training set is used for model training, and the validation set is used for model evaluation. The reason for using a knowledge graph is that it can effectively express the hierarchical structure and relationships of mathematical knowledge, facilitating subsequent model learning and understanding of mathematical concepts. The relationships between knowledge nodes are represented using a directed graph, specifically defined as follows:

[0149] G = (V, E, R);

[0150] In the formula, V is the set of nodes, representing mathematical concepts; E is the set of edges, representing the connections between nodes; and R is the set of relation types, including hierarchical relations and attribute relations. The purpose of this step is to construct a high-quality training dataset, providing a reliable data foundation for subsequent model training.

[0151] The specific implementation of step S02 is as follows: First, a word segmentation tool is used to segment the natural language descriptive text to obtain a word sequence; then, a pre-trained word vector model is used to extract the semantic features of the word sequence, and the semantic feature calculation formula is used to map each word into a 300-dimensional word vector to construct a semantic vector matrix.

[0152]

[0153] In the formula, s ij Let be the vector value of the i-th word in the j-th semantic dimension; m is the word sequence length; n is the semantic vector dimension, with a value of 300; w k v is the weight coefficient of the k-th context word; k Let be the word vector of the k-th context word; K be the size of the context window; and α be the semantic adjustment factor. The weight coefficients are calculated as follows:

[0154]

[0155] In the formula, k is the position of the context word, c is the position of the target word, and ∈ is the smoothing factor, ranging from 0.01 to 0.1. Next, regular expressions are used to extract numerical information from the text, including integers, decimals, fractions, percentages, etc. After normalizing the numerical information, a numerical feature matrix is ​​constructed.

[0156]

[0157] In the formula, n ij p is the value of the i-th value in the j-th feature dimension; p is the number of values; q is the feature dimension; x t Let be the t-th numerical sample; μ be the numerical mean; T be the total number of samples; and β be the numerical adjustment factor. Finally, dependency parsing techniques are used to extract logical relations from the word sequence, including causal, conditional, and parallel relations, and these logical relations are encoded into a relational feature matrix.

[0158]

[0159] In the formula, r ij Let u be the feature value representing the relationship between the i-th word and the j-th word; u and v are the word sequence lengths; c l d represents the weight coefficient of the l-th type of relation; lLet be the feature vector of the l-th type of relation; L be the total number of relation types; and γ be the relation adjustment factor. This step converts the natural language description text into a computer-processable feature representation, providing input for subsequent mathematical description transformations.

[0160] The specific implementation of step S03 is to first perform lexical analysis on the standard mathematical description language text to identify mathematical symbols, operators, keywords, and identifiers, and construct a mathematical symbol matrix:

[0161]

[0162] In the formula, m ij Let f be the value of the i-th symbol in the j-th feature dimension; a is the length of the symbol sequence; b is the feature dimension; f h g represents the weighting coefficient for the h-th class of symbols; h Let be the feature vector of the h-th symbol; H be the total number of symbol types; and δ be the symbol adjustment factor. Then, syntax analysis is performed to construct the operator matrix:

[0163]

[0164] In the formula, o ij Let be the value of the i-th operator on the j-th feature dimension; e is the length of the operator sequence; f is the feature dimension; q p z represents the weight coefficient for the p-th type of operation; p Let be the feature vector of the p-th type of operation; P be the total number of operation types; and θ be the operation adjustment factor. Next, semantic analysis is performed to extract the dependencies and calling relationships between mathematical logic units, constructing a mathematical logic unit matrix:

[0165]

[0166] In the formula, u ij The relationship value between the i-th logic unit and the j-th logic unit; g and h are the number of logic units; k y Let l be the weight coefficient of the y-th type of relation; y Let be the feature vector of the y-th relation; Y be the total number of relation types; and λ be the unit adjustment factor. The purpose of this step is to convert standard mathematical description language text into a structured matrix representation, providing the target output for model training.

[0167] The specific implementation of step S04 is as follows: First, estimate the GPU memory requirements during training based on the size of the training dataset and the model structure, and calculate the total size of the required GPU memory blocks; then, divide the GPU memory blocks into basic GPU memory blocks and dynamic GPU memory blocks, where the basic GPU memory blocks account for 60% of the total GPU memory and are used to store model parameters and fixed data, and the dynamic GPU memory blocks account for 40% of the total GPU memory and are used to store intermediate calculation results; the formula for calculating the size of the basic GPU memory blocks is:

[0168] M base =M total ×0.6×η;

[0169] In the formula, M total η is the total video memory size; η is the video memory utilization factor, ranging from 0.8 to 0.9. The formula for calculating the size of the dynamic video memory block is:

[0170] M dynamic =M total ×0.4×η;

[0171] Next, a pre-allocation strategy for GPU memory is adopted to allocate the GPU memory space required for training in advance, avoiding frequent GPU memory allocation and release during training. Finally, a GPU memory management mechanism is established to realize the dynamic scheduling and reclamation of GPU memory blocks. The purpose of this step is to optimize GPU memory usage efficiency and improve model training speed.

[0172] The specific implementation of step S05 is as follows: First, a semantic understanding layer is constructed, and a bidirectional long short-term memory network is used to extract the contextual semantic information of the text. The state update formula is as follows:

[0173] h t =f(W h ×[h t-1 x t ]+b h );

[0174] In the formula, h t Let W be the hidden state at time t; h Here is the weight matrix; x t b is the input vector; h is the bias term; f is the activation function. Then, a mathematical symbol mapping layer is constructed, employing an attention mechanism to achieve the mapping transformation from semantic features to mathematical symbols. The attention weight calculation formula is:

[0175]

[0176] In the formula, a ij For attention weights; e ij is the energy function value; n is the sequence length. Next, a logical relationship transformation layer is constructed, using a graph neural network to establish the logical relationships between mathematical symbols. The node update formula is:

[0177]

[0178] In the formula, W represents the state of node i at time t. v Here is the weight matrix; m ij b is the edge weight; N(i) is the set of neighbors of node i; vis the bias term; g is the activation function. Then, a mathematical unit organization layer is constructed, using a recurrent neural network to organize mathematical symbols and their logical relationships into mathematical logical units. Finally, a unit relationship construction layer is built, using a tree-type recurrent neural network to establish the hierarchical relationships between mathematical logical units. The purpose of this step is to construct a neural network model capable of understanding natural language and generating mathematical descriptions.

[0179] The specific implementation of step S06 is to first input the imprecise description matrix into the mathematical logic reasoning network model, and the processing formula of the input layer is:

[0180] X = W i ×I+b i ;

[0181] In the formula, X represents the input layer output; W i I is the input layer weight matrix; b is the input matrix; i This is the bias term. Then, a semantic understanding layer extracts deep semantic features from the text; next, a mathematical symbol mapping layer converts the semantic features into a sequence of mathematical symbols, and a cross-entropy loss function is used to optimize the mapping accuracy.

[0182]

[0183] In the formula, L ce For cross-entropy loss; y i This is a real label; The probability is predicted; n is the number of samples. Then, a logical relationship transformation layer is used to construct the logical relationships between mathematical symbols; finally, a unit relationship construction layer is used to generate a mathematical logical unit tree structure. The purpose of this step is to realize the conversion process from natural language to mathematical description.

[0184] The formula for converting the semantic vector matrix into a mathematical concept representation through the semantic understanding layer is as follows:

[0185] C = f(W) c ×S+b c );

[0186] In the formula, C represents a mathematical concept; W c S is the transformation weight matrix; S is the semantic vector matrix; b c is the bias term; f is the hyperbolic tangent activation function. The formula for mapping mathematical concept representations to sequences of mathematical symbols through the mathematical symbol mapping layer is:

[0187]

[0188] In the formula, P ij e represents the probability that the i-th concept is mapped to the j-th symbol; ijWhere K is the mapping energy and K is the size of the symbol dictionary. The formula for calculating the logical relationships between sequences of mathematical symbols constructed through the logical relation transformation layer is as follows:

[0189] R ij =g(W r ×[h i h j ]+b r );

[0190] In the formula, R ij W is the relationship vector between symbols i and j. r h is the relational transformation matrix. i and h j For the hidden state of the symbol; b r is the bias term; g is the sigmoid activation function. The calculation formula for organizing mathematical symbol sequences and logical relationships into mathematical logic units through the mathematical unit organization layer is:

[0191] U t =σ(W u ×[x t h t-1 ]+b u )×tanh(W c ×[x t r t ×h t-1 ]+b c );

[0192] In the formula, U t Let x be the mathematical logic unit representation of time t; t For input symbols; h t-1 The hidden state of the previous time step; r t To reset the door; W u and W c b is the weight matrix; u and b c This is a bias term. The formula for calculating the hierarchical and calling relationships between mathematical logic units, established through the unit relationship construction layer, is as follows:

[0193] T ij =softmax(W t ×[U i U j ]+b t );

[0194] In the formula, T ij The probability distribution of the relationship between unit i and unit j; W t b is the weight matrix of the tree structure. t This is a bias term.

[0195] The specific implementation of step S07 is to first calculate the semantic accuracy score, and then use cosine similarity to calculate the semantic similarity between the predicted text and the standard text:

[0196]

[0197] In the formula, A s Score for semantic precision; x i and y i Let be the i-th semantic vector of the predicted text and the standard text, respectively; n be the dimension of the semantic vector; and η be the semantic precision adjustment factor. Then, the symbol precision score is calculated:

[0198]

[0199] In the formula, A m Score for symbol precision; D i L is the edit operand at position i; p and L s ξ represents the lengths of the predicted sequence and the standard sequence, respectively; ξ is the sign precision adjustment factor.

[0200] First, calculate the overall loss function:

[0201] L total =α1L s +α2L m +α3L t +α4L l +λΩ(W);

[0202] In the formula, L s L m L t L l The loss terms are semantic loss, symbolic loss, structural loss, and logical loss, respectively; α1, α2, α3, and α4 are the loss weights; λ is the regularization coefficient; and Ω(W) is the L2 regularization term. Then, an adaptive moment estimation optimization algorithm is used to update the parameters.

[0203] m t =β1m t-1 +(1-β1)g t ;

[0204]

[0205] In the formula, m t and v t β1 and β2 are the first and second order momentum; β1 and β2 are the momentum decay rates; g t η is the gradient; η is the learning rate; ∈ is the numerical stability factor.

[0206] The purpose of this step is to evaluate the accuracy of the model's predictions and provide guidance for model optimization.

[0207] The specific implementation of step S08 involves first acquiring the natural language description text from the verification dataset and processing it using the established data processing workflow; repeating step S02, segmenting the verification text, extracting semantic features, and constructing a semantic vector matrix S; extracting numerical features and constructing a numerical feature matrix N; extracting logical relation features and constructing a relational feature matrix R; then repeating step S06, inputting the constructed imprecise description matrix into the trained mathematical logic reasoning network model, and generating a mathematical logic unit tree structure and predicted mathematical description text through sequential processing of the semantic understanding layer, mathematical symbol mapping layer, logical relation transformation layer, mathematical unit organization layer, and unit relation construction layer. This step uses the same feature extraction and model reasoning workflow as the training phase, ensuring the consistency of the verification process.

[0208] The specific implementation of step S09 is to first calculate the accuracy score of the verification prediction results, including semantic accuracy A. s Symbol precision A m Structural accuracy A t And logical precision A l Then, the accuracy scores for each item are compared with preset thresholds, where the preset thresholds for semantic accuracy are 0.85, symbolic accuracy are 0.9, structural accuracy are 0.8, and logical accuracy are 0.85. If any accuracy score is lower than the corresponding threshold, the validation sample is added to the training dataset. The expanded training dataset is then retrained, and the model parameters are optimized using the gradient descent algorithm.

[0209]

[0210] In the formula, W new The updated weights; W old The weights are as follows: α is the learning rate; L is the loss function. This represents the gradient of the loss function with respect to the weights. The purpose of this step is to improve the model's prediction accuracy through incremental training.

[0211] The specific implementation of step S10 is as follows: First, a unit contribution value evaluation model is constructed, using a multilayer perceptron structure. The intermediate features of each layer are input, and the accuracy score is output. Then, the gradient backpropagation algorithm is used to calculate the contribution weight of each layer.

[0212]

[0213] In the formula, W i The contribution weights of the i-th layer are: L is the total loss function; O iLet be the output of the i-th layer; ψ is the contribution weight adjustment factor. The formula for calculating the total loss function is:

[0214] K=λ1L s +λ2L m +λ3L t +λ4L l ;

[0215] In the formula, L s L m L t L l These are semantic loss, symbolic loss, structural loss, and logical loss, respectively; λ1, λ2, λ3, and λ4 are the weight coefficients for each loss term. The purpose of this step is to quantitatively evaluate the importance of each network layer, providing a basis for resource allocation.

[0216] The specific implementation of step S11 is to first determine the memory allocation of each layer based on the contribution weight:

[0217]

[0218] In the formula, M i M represents the memory allocation for the i-th layer; base Basic memory size; W i Let σ be the contribution weight for the i-th layer; σ is the memory allocation adjustment factor. Then, the number of computation threads for each layer is determined based on the contribution weights.

[0219] T i =T base ×W i +τ;

[0220] In the formula, T i T represents the number of computation threads in the i-th layer; base τ is the base number of threads; τ is the thread number adjustment factor. Finally, the number of training batches for each layer is determined based on the contribution weights.

[0221] B i =B base ×W i +ρ;

[0222] In the formula, B i B is the number of training batches for the i-th layer; base ρ is the base batch number; ρ is the batch number adjustment factor. The purpose of this step is to rationally allocate computational resources according to the importance of each layer, thereby improving training efficiency.

[0223] The specific implementation of step S12 is to first calculate the utilization efficiency index of the video memory block:

[0224]

[0225] In the formula, E i U represents the utilization efficiency of the i-th memory block. i For effective usage time; T i For total allocated time; F i For access frequency; F max Maximum access frequency; S i To occupy space; S total This is the total video memory space. Then, memory defragmentation is performed:

[0226]

[0227] In the formula, M compact For the consolidated contiguous video memory space; M i γ is the size of the i-th memory block; i This refers to the fragmentation rate. Next, idle video memory is reclaimed:

[0228]

[0229] In the formula, M recycle The size of the reclaimable video memory; t i Idle time; T threshold I represents the recovery threshold; I is the indicator function. Then, priority redistribution is performed based on contribution weights:

[0230]

[0231] In the formula, P i For memory block priority; W i Contribute weights to the corresponding layer; E i For efficiency; For average efficiency; D i For data dependency; D max Maximum dependency. Last updated video memory usage status:

[0232]

[0233] In the formula, and These represent the old and new states, respectively; ΔS i Let be the state change variable; α be the smoothing factor. Through these optimization operations, efficient utilization and dynamic adjustment of GPU memory resources are achieved until the model performance meets the expected requirements.

[0234] To better understand and implement this invention, a specific application scenario, Example 2, is provided below: A research team at a university is conducting research on an intelligent education project. The project goal is to build an intelligent system capable of converting natural language mathematical problem descriptions into standard mathematical expressions. The research team uses the method of this invention, and the specific implementation process is as follows:

[0235] In the first step, the research team collected 50,000 math problems from high school textbooks, test papers, and exercise sets, including word problems, geometry problems, and algebra problems. Each data point included a natural language description and the corresponding standard mathematical expression. After data cleaning and filtering, 45,000 valid data points were obtained, including 36,000 in the training set and 9,000 in the validation set. Examples of the training data are shown in Table 1.

[0236] Table 1: Examples of Training Data

[0237]

[0238] The second step involves word segmentation and feature extraction of the natural language description text. A word segmentation tool is used, with a word vector dimension of 300 and a context window size of 5. Numerical information is extracted and normalized, with a feature dimension of 50. Dependency parsing is used to extract logical relationships, with 20 relationship types. The final imprecise description matrix contains a semantic vector matrix of 200×300 dimensions, a numerical feature matrix of 50×50 dimensions, and a relational feature matrix of 200×200 dimensions.

[0239] The third step involves structuring standard mathematical expressions. A mathematical symbol dictionary is constructed, containing 100 basic operators and 200 mathematical function symbols. The symbol feature dimension is set to 100, and the operation feature dimension to 50. A mathematical logic unit tree is built through syntax analysis, with an average of 8 logic units per expression. In the resulting precise description matrix, the mathematical symbol matrix has a dimension of 300×100, the operator matrix has a dimension of 100×50, and the mathematical logic unit matrix has a dimension of 300×300.

[0240] In the fourth step, the research team used an NVIDIA A100 GPU with 32GB of video memory for model training. A base memory block allocation of 19.2GB was used to store model parameters and fixed data; a dynamic memory block allocation of 12.8GB was used to store intermediate computation results. The memory utilization factor was set to 0.85, and the smoothing factor was set to 0.05.

[0241] The fifth step is to construct a 5-layer mathematical logic reasoning network model. The semantic understanding layer uses a bidirectional long short-term memory network with a hidden layer dimension of 512; the mathematical symbol mapping layer uses a multi-head attention mechanism with 8 attention heads; the logical relation transformation layer uses a graph convolutional network with a convolution kernel size of 3; the mathematical unit organization layer uses a gated recurrent unit network with a hidden layer dimension of 256; and the unit relation construction layer uses a tree-type recurrent neural network with an output dimension of 128.

[0242] Step 6: Model training employs an iterative approach with a batch size of 64. The semantic accuracy threshold is set to 0.85, the symbolic accuracy threshold to 0.9, the structural accuracy threshold to 0.8, and the logical accuracy threshold to 0.85. During training, resource allocation is optimized by dynamically evaluating the contribution weights of each layer. The initial contribution weights and resource allocation for each layer are shown in Table 2.

[0243] Table 2: Contribution Weights and Resource Allocation at Each Level

[0244] Network layer Contribution weight Video memory allocation (GB) Calculate the number of threads Number of training batches Semantic understanding layer 0.3 3.84 12 1000 Symbol mapping layer 0.25 3.20 10 800 Relationship Transformation Layer 0.2 2.56 8 600 Unit organization level 0.15 1.92 6 400 Relationship Building Layer 0.1 1.28 4 200

[0245] Step 7: Monitor GPU memory usage in real time during training, updating the memory status every 10 batches. The trigger threshold for memory defragmentation is set to 30%, and the time threshold for memory block reclamation is set to 100 training batches. By dynamically adjusting the memory allocation strategy, the memory utilization rate during model training reaches 92%.

[0246] After 50 training epochs, the model achieved the following performance metrics on the validation set: semantic accuracy of 0.87, symbolic accuracy of 0.92, structural accuracy of 0.83, and logical accuracy of 0.88. For a new input natural language mathematical problem description, the model can generate the corresponding standard mathematical expression within 0.5 seconds, achieving an accuracy of 85%.

[0247] Traditional training methods for mathematical description language models employ a fixed memory allocation strategy, typically distributing memory evenly across network layers or using a simple proportional allocation based on the number of parameters in each layer. This approach fails to consider the actual importance and changing resource requirements of each layer during training, leading to insufficient memory in some critical layers, impacting training performance, while other layers suffer from memory waste. Furthermore, due to the lack of an effective memory management mechanism, memory fragmentation is severe, with actual memory utilization often only reaching 60% to 70%.

[0248] This invention employs a dynamic memory optimization allocation strategy based on contribution values. By evaluating the contribution of each layer to model performance in real time, it achieves precise allocation of memory resources. Combined with a memory pool management mechanism, it effectively solves the memory fragmentation problem, increasing memory utilization to over 90%. In this embodiment, by rationally allocating computing resources, the model training speed is increased by 40% compared to traditional methods, and higher prediction accuracy is achieved within the same training cycle. This demonstrates that this invention successfully solves the problem of low computing resource utilization caused by unreasonable memory allocation, providing an effective solution for efficient training of mathematical description language models.

[0249] The following is a specific application scenario of the present invention, Example 3: A research team conducts research based on the method of the present invention, and the specific implementation process is as follows:

[0250] In the first step, the research team collected 100,000 English papers in mathematics, physics, and computer science from databases such as arXiv and Web of Science, and extracted 500,000 pairs of natural language descriptions and mathematical formulas. Through manual screening and quality assessment, incomplete, erroneous, and duplicate data were removed, ultimately yielding 400,000 valid data pairs. The dataset was then divided into a training set of 280,000 pairs, a validation set of 80,000 pairs, and a test set of 40,000 pairs in a 7:2:1 ratio. An example of the training data is shown in Table 3.

[0251] Table 3: Examples of Training Data

[0252]

[0253] The second step involves feature extraction and matrix construction of the natural language description text. First, a BERT pre-trained model is used for text encoding to obtain a 768-dimensional text representation vector. Then, a bidirectional Long Short-Term Memory (LSTM) network is used to extract contextual features, with the hidden layer dimension set to 512. Specialized marking is applied to technical terms and mathematical symbols to construct a feature dictionary containing 5000 technical terms and 1000 mathematical symbols. The final imprecise description matrix consists of the following three parts:

[0254] 1. Semantic vector matrix: 1000×768 dimensions, with feature weighting using an attention mechanism: The context window size K is set to 7, and the weight coefficient w k The calculation uses relative position encoding, and the smoothing factor ∈ is set to 0.05.

[0255] 2. Numerical feature matrix: 200×100 dimensions, extracted using statistical features: The numerical adjustment factor β was determined to be 0.15 through cross-validation.

[0256] 3. Relation Feature Matrix: 1000×1000 dimensions, using dependency parsing: The total number of relation types L is set to 50, and the relation adjustment factor γ is set to 0.2.

[0257] The third step involves structuring the standard mathematical expressions. A parser identifies mathematical symbols and operators, constructing a mathematical logic unit tree. The resulting precise description matrix includes:

[0258] 1. Mathematical symbol matrix: 500×200 dimensions: The total number of symbol types H is 300, and the symbol adjustment factor δ is set to 0.1.

[0259] 2. Operator matrix: 200×100 dimensions: The total number of operation types P is 100, and the operation adjustment factor θ is set to 0.15.

[0260] 3. Mathematical logic unit matrix: 500×500 dimensions: The total number of relation types Y is 80, and the unit adjustment factor λ is set to 0.25.

[0261] Fourth, the research team used an 8-GPU NVIDIA A100 cluster for distributed training, with each GPU configured with 80GB of video memory. The video memory allocation scheme is as follows:

[0262] 1. Basic Video Memory Configuration: Each graphics card is allocated 48GB of basic video memory for storing model parameters and fixed data; 32GB of dynamic video memory is allocated for storing intermediate calculation results. The video memory utilization factor η is set to 0.9. The basic video memory amount is calculated as follows: M base =M total ×0.6×η=80×0.6×0.9=43.2GB.

[0263] 2. Dynamic video memory configuration: M dynamic =M total ×0.4×η=80×0.4×0.9=28.8GB.

[0264] Fifth, construct a 5-layer mathematical logic reasoning network model, with the following configuration parameters for each layer:

[0265] 1. Semantic understanding layer: It adopts a 12-layer Transformer encoder with 8 attention heads per layer, a hidden layer dimension of 1024, a feedforward network dimension of 4096, and a dropout rate of 0.1.

[0266] 2. Mathematical symbol mapping layer: adopts a cross-attention mechanism with 16 attention heads, 64 key-value dimensions, and 512 mapping hidden layer dimensions.

[0267] 3. Logical Relationship Transformation Layer: A graph attention network is used with 4 layers, 8 attention heads per layer, 256 feature dimensions, residual connections, and layer normalization.

[0268] 4. Mathematical Unit Organization Layer: A recurrent neural network with LSTM units, 512 hidden layer dimensions, 3 layers, and bidirectional processing is adopted.

[0269] 5. Unit Relationship Construction Layer: A tree-structured LSTM network is used, with an input dimension of 256, a hidden layer dimension of 512, and a maximum tree depth of 10.

[0270] Step 6: The following optimization strategies are used during the training process:

[0271] 1. Loss function settings: Total loss function: L total =α1L s +α2L m +α3L t +α4L l +λΩ(W); where α1=0.3, α2=0.25, α3=0.25, α4=0.2, and the regularization coefficient λ=0.001.

[0272] 2. Optimizer Configuration: The Adam optimizer is used, with an initial learning rate of 0.0001, β1 = 0.9, β2 = 0.999, and weight decay of 0.01. Cosine annealing learning rate scheduling is used, with a minimum learning rate of 1e-6.

[0273] 3. Training batch settings: The initial batch size is 128, a gradient accumulation strategy is adopted, and the cumulative steps are 4. The total number of training rounds is 100, and the patience index of the early stopping strategy is set to 10.

[0274] Step 7: Dynamic resource allocation strategy based on contribution weights:

[0275] 1. Contribution weight calculation: The contribution weight adjustment factor ψ is set to 0.1.

[0276] 2. Calculation of video memory allocation: The memory allocation adjustment factor σ is set to 0.2.

[0277] 3. Calculate thread allocation: T i =T base ×W i +τ; The base number of threads is set to 16, and the thread number adjustment factor τ is set to 0.15.

[0278] 4. Training batch allocation: B i =B base ×W i +ρ; The base batch number is set to 1000, and the batch number adjustment factor ρ is set to 0.1.

[0279] After 100 rounds of training, the model achieved the following performance metrics on the validation set: 1. Semantic accuracy: 0.893; 2. Symbolic accuracy: 0.921; 3. Structural accuracy: 0.856; 4. Logical accuracy: 0.878.

[0280] In practical application testing, for newly input natural language mathematical descriptions, the model can generate the corresponding standard mathematical expression within 1 second, with an accuracy of 88.5%. The model's performance in handling different types of mathematical expressions is shown in Table 4.

[0281] Table 4: Processing performance of different types of mathematical expressions

[0282] expression type Sample size accuracy Average processing time (ms) Algebraic expressions 10000 91.2% 180 Calculus formula 8000 87.5% 320 Probability and Statistics 7000 89.6% 250 Linear Algebra 9000 86.8% 420 Geometric formulas 6000 90.3% 290

[0283] Traditional methods suffer from the following main problems: 1. They employ a fixed-ratio memory allocation strategy, failing to dynamically adjust resource allocation based on the actual needs of each layer during model training; 2. They lack a quantitative evaluation mechanism for the contribution of each layer, resulting in inaccurate resource allocation; 3. Their simple memory management mechanism easily leads to memory fragmentation, affecting training efficiency. These problems result in traditional methods generally achieving memory utilization rates below 70% and slower model training speeds.

[0284] In comparison, the present invention has the following advantages: 1. By dynamically adjusting the GPU memory allocation through contribution value evaluation, the GPU memory utilization rate is increased to 92%; 2. The incremental training strategy based on accuracy evaluation improves the model's prediction accuracy by 5 percentage points; 3. The GPU memory pool management mechanism effectively reduces GPU memory fragmentation, increasing training speed by 40%; 4. The multi-level model structure design improves the processing capability of complex mathematical expressions. Practice has proven that the present invention successfully solves the technical problem of low GPU memory resource utilization, providing a feasible solution for the efficient training of mathematical description language models.

[0285] It should be noted that the variables involved in this invention are explained in detail in Tables 5 and 6 below.

[0286] Table 5. Variable Explanation Table (Part 1)

[0287]

[0288]

[0289] Table 6. Variable Explanation Table (Part Two)

[0290]

[0291] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for establishing a large model of a mathematical description language, characterized in that, Includes the following steps: Acquire natural language description texts and their corresponding standard mathematical description language texts from the training corpus to establish a training dataset; perform word segmentation on the natural language description texts, extract semantic features, numerical features, and logical relationship features, and construct an imprecise description matrix; perform structured processing on the standard mathematical description language texts, extract mathematical symbols, operators, and mathematical logic units, and construct an accurate description matrix; establish a memory pool and pre-allocate memory blocks required for training, dividing the memory blocks into basic memory blocks and dynamic memory blocks. The basic memory blocks are used to store the imprecise description matrix, the accurate description matrix, and fixed parameters. The dynamic memory blocks are used to store intermediate calculations. The results are calculated; a mathematical logic reasoning network model is constructed, including a semantic understanding layer, a mathematical symbol mapping layer, a logical relation transformation layer, a mathematical unit organization layer, and a unit relation construction layer; the inaccurate description matrix is ​​input into the mathematical logic reasoning network model to generate a mathematical logic unit tree structure and predicted mathematical description text; the accuracy score between the predicted mathematical description text and the standard mathematical description language text is calculated; a unit contribution value evaluation model is established based on the accuracy score of the verified prediction results; the gradient backpropagation algorithm is used to calculate the contribution weight of each layer; the amount of video memory allocation, the number of computing threads, and the number of training batches are determined according to the contribution weights calculated by the unit contribution value evaluation model. The memory allocation of the memory pool is optimized based on the memory allocation amount until the accuracy score of the verification prediction result reaches a preset accuracy threshold.

2. The method for establishing a large model of a mathematical description language according to claim 1, characterized in that, The steps for establishing the training dataset are as follows: a training corpus is constructed by building a mathematical knowledge graph. Mathematical concepts, theorems, formulas and problem-solving methods are extracted from mathematical textbooks, papers and professional literature. A set of mathematical knowledge nodes is constructed, and the inclusion, inheritance, combination and transformation relationships between mathematical knowledge nodes are established. The mathematical knowledge nodes and their relationships are then stored in a structured manner.

3. The method for establishing a large model of a mathematical description language according to claim 1, characterized in that, The preset accuracy thresholds include a preset semantic accuracy threshold, a preset symbol accuracy threshold, a preset structural accuracy threshold, and a preset logical accuracy threshold, wherein the preset semantic accuracy threshold is 0.85, the preset symbol accuracy threshold is 0.9, the preset structural accuracy threshold is 0.8, and the preset logical accuracy threshold is 0.

85.

4. The method for establishing a large model of a mathematical description language according to claim 1, characterized in that, The semantic understanding layer uses a bidirectional long short-term memory network to extract contextual semantic information of the text; the mathematical symbol mapping layer uses an attention mechanism to realize the mapping and conversion of semantic features to mathematical symbols; the logical relationship conversion layer uses a graph neural network to establish logical relationships between mathematical symbols; the mathematical unit organization layer uses a recurrent neural network to organize mathematical logical units; and the unit relationship construction layer uses a tree-type recurrent neural network to establish hierarchical relationships between mathematical logical units.

5. The method for establishing a large model of a mathematical description language according to claim 1, characterized in that, The total size of the video memory block is determined based on the size of the non-precise description matrix, the size of the precise description matrix, and the preset intermediate calculation result storage requirements. The basic video memory block accounts for 60% of the total video memory, and the dynamic video memory block accounts for 40% of the total video memory.

6. The method for establishing a large model of a mathematical description language according to claim 1, characterized in that, The method for determining the contribution weights is as follows: the semantic understanding layer mainly affects semantic accuracy, the mathematical symbol mapping layer mainly affects symbol accuracy, the logical relation transformation layer mainly affects logical accuracy, and the mathematical unit organization layer and the unit relation construction layer mainly affect structural accuracy.

7. The method for establishing a large model of a mathematical description language according to claim 1, characterized in that, The method for determining the amount of video memory allocated is as follows: the greater the contribution weight, the more video memory is allocated, and the specific ratio is the square root of the contribution weight. The method for determining the number of computing threads is as follows: the greater the contribution weight, the more computing threads are allocated, and the specific number is the base number of threads multiplied by the contribution weight. The method for determining the number of training batches is as follows: the greater the contribution weight, the more training batches are allocated, and the specific number of batches is the base number of batches multiplied by the contribution weight.

8. The method for establishing a large model of a mathematical description language according to claim 1, characterized in that, The memory allocation optimization includes merging free memory in the memory block, reclaiming idle memory in the memory block, reallocating the priority of the memory block according to the contribution weight, and updating the usage status of the memory block, which includes usage time, access frequency, and idle ratio.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions, which, when executed in a computer, are used to perform the method for establishing a large model of a mathematical description language as described in any one of claims 1-8.

10. A system for establishing large-scale models of a mathematical description language, characterized in that, The system includes the computer-readable storage medium of claim 9, wherein the system is any one of a computer, a server, or a microcontroller, the computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor that executes the program instructions stored in the computer-readable storage medium.

Citation Information

Patent Citations

  • Method for optimizing speech recognition process, and device thereof and storage medium

    CN113205818A

  • Method and device for improving throughput of large language model

    CN117349032A