An overload transportation large model implementation method and system fusing multi-stage decision and constraint rule feedback

The large-scale heavy-haul railway transportation model, which incorporates multi-stage decision-making and constraint rule feedback, solves the problem of relying on human experience in heavy-haul railway transportation and enables the intelligent and efficient generation of line technical standards and transportation plans.

CN122433894APending Publication Date: 2026-07-21SHUOHUANG RAILWAY DEV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHUOHUANG RAILWAY DEV
Filing Date
2026-04-10
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

The determination of line technical standards and transportation organization schemes in existing heavy-haul railway transportation relies on manual experience, resulting in low processing efficiency, insufficient intelligence, and difficulty in adapting to complex application needs.

Method used

A large-scale heavy-haul railway transportation model with multi-stage decision-making and constraint rule feedback is adopted. The model is constructed through a three-stage training system. Combining the knowledge retrieval, parameter derivation and scheme generation stages, a constraint rule feedback mechanism is introduced to verify compliance and safety, and automatically correct and recalculate.

Benefits of technology

It improves the intelligence level and processing efficiency of heavy-haul railway transportation organization, reduces reliance on personnel experience, and ensures the safety and accuracy of output results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122433894A_ABST
    Figure CN122433894A_ABST
Patent Text Reader

Abstract

The application relates to a heavy-load railway transportation large model implementation method and system fusing multi-stage decision and constraint rule feedback, and relates to the technical field of heavy-load railway transportation. The application can improve the processing efficiency and intelligent level of heavy-load railway transportation organization. The method comprises the following steps: determining a base model to be trained, training the base model to obtain a heavy-load railway transportation large model; constructing a multi-stage intelligent decision generation architecture for heavy-load railway transportation, dividing the reasoning process into a knowledge retrieval stage, a parameter derivation stage and a standard or scheme generation stage; introducing a reasoning checking mechanism based on constraint rule feedback, checking the compliance and safety of candidate standards or candidate schemes output by the standard or scheme generation stage; in the case of passing the checking, outputting the candidate standard as a standard output result or the candidate scheme as a scheme output result; in the case of failing to pass the checking, triggering secondary reasoning and parameter recalculation to obtain a standard correction result or a scheme correction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of heavy-haul railway transportation technology, and in particular to a method, system, computer equipment, computer-readable storage medium, and computer program product for implementing a large-scale heavy-haul railway transportation model that integrates multi-stage decision-making and constraint rule feedback. Background Technology

[0002] In the field of heavy-haul railway transportation, the determination of line technical standards and the generation of transportation organization plans usually involve a large number of professional factors such as line conditions, traction capacity and vehicle formation, and must comply with relevant technical standards and regulations.

[0003] At present, the traditional transportation organization methods still mainly rely on human experience and offline computing tools to determine the technical standards of the line or to formulate transportation organization plans. This not only involves a large workload and slow response speed, but also relies heavily on human experience, resulting in problems such as low processing efficiency and low level of intelligence. It is difficult to adapt to the complex application needs of modern heavy-haul railway transportation with multiple scenarios and multiple constraints. Summary of the Invention

[0004] Based on this, it is necessary to provide a method, system, computer equipment, computer-readable storage medium, and computer program product for implementing a large-scale heavy-haul railway transportation model that integrates multi-stage decision-making and constraint rule feedback, in order to address the above-mentioned technical problems.

[0005] Firstly, this application provides a method for implementing a large-scale heavy-haul railway transportation model that integrates multi-stage decision-making and constraint rule feedback, including:

[0006] The base model to be trained is determined, and the base model is trained through a preset three-stage training system to obtain a large heavy-haul railway transportation model with specialized reasoning capabilities for heavy-haul transportation scenarios. The large heavy-haul railway transportation model is then adapted to the computing power platform.

[0007] A multi-stage intelligent decision generation architecture for heavy-haul railway transportation is constructed, and the reasoning process of the heavy-haul railway transportation model is divided into a knowledge retrieval stage, a parameter derivation stage, and a standard or scheme generation stage, which are executed sequentially.

[0008] A reasoning verification mechanism based on constraint rule feedback is introduced to verify the compliance and security of the candidate standards or candidate solutions output during the standard or solution generation stage.

[0009] If the verification passes, the candidate standard is used as the standard output result or the candidate solution is used as the solution output result; if the verification fails, a correction prompt is automatically generated and secondary inference and parameter recalculation are triggered to obtain the standard correction result or the solution correction result.

[0010] In one embodiment, training the base model through a preset three-stage training system to obtain a large-scale heavy-haul railway transportation model with specialized reasoning capabilities for heavy-haul transportation scenarios includes:

[0011] Multi-source heterogeneous data on heavy-haul railway equipment and transportation operations are acquired, and the data undergoes format conversion and data cleaning to form a standardized domain corpus. Based on the standardized domain corpus, a continuous language modeling task is used to enable the base model to learn key knowledge in the transportation domain, resulting in an initial heavy-haul railway transportation model. The initial heavy-haul railway transportation model is then supervised and fine-tuned using a low-rank adaptive fine-tuning method to obtain the final heavy-haul railway transportation model.

[0012] In one embodiment, the step of employing a low-rank adaptive fine-tuning method to supervise the fine-tuning of the initial heavy-haul railway transportation model to obtain the heavy-haul railway transportation model includes:

[0013] In the base model, a target weight matrix is ​​selected, and two low-rank matrices are introduced for each target weight matrix. An increment matrix is ​​constructed based on the low-rank matrices. In the forward computation of the base model, the original output of the base model is added to the output of the increment matrix. During the supervised fine-tuning process, the parameters of the low-rank matrices are updated and the target weight matrices are kept frozen. The updated low-rank matrices are merged with the target weight matrices to obtain the large model of heavy-haul railway transportation.

[0014] In one embodiment, the method further includes:

[0015] A dedicated knowledge base for heavy-haul railway transportation is constructed, and various types of data are managed in a structured and vectorized manner through the dedicated knowledge base. In the knowledge retrieval stage, the heavy-haul railway transportation big model searches the dedicated knowledge base using a combination of semantic retrieval and keyword retrieval based on the user-input computational task requirements and intermediate requests generated by the system, and dynamically recalls relevant domain knowledge as the knowledge retrieval results.

[0016] In one embodiment, the method further includes:

[0017] In the parameter derivation stage, the heavy-haul railway transportation big model, based on the knowledge retrieval results and combined with the specific transportation task objectives, uses mathematical models and optimization algorithms from the tool library to quantify and calculate the key technical parameters of the objectives, thereby obtaining the parameter derivation results.

[0018] In one embodiment, the method further includes:

[0019] During the standard or scheme generation stage, the heavy-haul railway transportation big model is used as the core. Semantic information obtained from the knowledge retrieval results and quantitative results obtained from the parameter derivation results are integrated to generate line technical standards or transportation organization schemes that meet engineering constraints and business objectives, which serve as candidate standards or candidate schemes.

[0020] Secondly, this application also provides a large-scale model implementation system for heavy-haul railway transportation that integrates multi-stage decision-making and constraint rule feedback, including:

[0021] The model training module is used to determine the base model to be trained, train the base model through a preset three-stage training system to obtain a large heavy-haul railway transportation model with specialized reasoning capabilities for heavy-haul transportation scenarios, and adapt the large heavy-haul railway transportation model to the computing power platform.

[0022] The architecture generation module is used to construct a multi-stage intelligent decision generation architecture for heavy-haul railway transportation, which divides the reasoning process of the heavy-haul railway transportation model into a knowledge retrieval stage, a parameter derivation stage, and a standard or scheme generation stage that are executed sequentially.

[0023] The inference verification module is used to introduce an inference verification mechanism based on constraint rule feedback to verify the compliance and security of the candidate standards or candidate solutions output during the standard or solution generation stage.

[0024] The inference recalculation module is used to output the candidate standard as the standard output result or the candidate scheme as the scheme output result when the verification passes; when the verification fails, it automatically generates a correction prompt and triggers secondary inference and parameter recalculation to obtain the standard correction result or the scheme correction result.

[0025] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0026] A base model to be trained is determined, and the base model is trained through a preset three-stage training system to obtain a large-scale heavy-haul railway transportation model with specialized reasoning capabilities for heavy-haul transportation scenarios. The large-scale heavy-haul railway transportation model is then adapted to a computing platform. A multi-stage intelligent decision generation architecture for heavy-haul railway transportation is constructed, dividing the reasoning process of the large-scale heavy-haul railway transportation model into a knowledge retrieval stage, a parameter derivation stage, and a standard or scheme generation stage, which are executed sequentially. A reasoning verification mechanism based on constraint rule feedback is introduced to verify the compliance and security of the candidate standards or candidate schemes output in the standard or scheme generation stage. If the verification passes, the candidate standard is output as the standard output result or the candidate scheme is output as the scheme output result. If the verification fails, a correction prompt is automatically generated and secondary reasoning and parameter recalculation are triggered to obtain the standard correction result or the scheme correction result.

[0027] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:

[0028] A base model to be trained is determined, and the base model is trained through a preset three-stage training system to obtain a large-scale heavy-haul railway transportation model with specialized reasoning capabilities for heavy-haul transportation scenarios. The large-scale heavy-haul railway transportation model is then adapted to a computing platform. A multi-stage intelligent decision generation architecture for heavy-haul railway transportation is constructed, dividing the reasoning process of the large-scale heavy-haul railway transportation model into a knowledge retrieval stage, a parameter derivation stage, and a standard or scheme generation stage, which are executed sequentially. A reasoning verification mechanism based on constraint rule feedback is introduced to verify the compliance and security of the candidate standards or candidate schemes output in the standard or scheme generation stage. If the verification passes, the candidate standard is output as the standard output result or the candidate scheme is output as the scheme output result. If the verification fails, a correction prompt is automatically generated and secondary reasoning and parameter recalculation are triggered to obtain the standard correction result or the scheme correction result.

[0029] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:

[0030] A base model to be trained is determined, and the base model is trained through a preset three-stage training system to obtain a large-scale heavy-haul railway transportation model with specialized reasoning capabilities for heavy-haul transportation scenarios. The large-scale heavy-haul railway transportation model is then adapted to a computing platform. A multi-stage intelligent decision generation architecture for heavy-haul railway transportation is constructed, dividing the reasoning process of the large-scale heavy-haul railway transportation model into a knowledge retrieval stage, a parameter derivation stage, and a standard or scheme generation stage, which are executed sequentially. A reasoning verification mechanism based on constraint rule feedback is introduced to verify the compliance and security of the candidate standards or candidate schemes output in the standard or scheme generation stage. If the verification passes, the candidate standard is output as the standard output result or the candidate scheme is output as the scheme output result. If the verification fails, a correction prompt is automatically generated and secondary reasoning and parameter recalculation are triggered to obtain the standard correction result or the scheme correction result.

[0031] The aforementioned method, system, computer equipment, computer-readable storage medium, and computer program product for implementing a large-scale heavy-haul railway transportation model integrating multi-stage decision-making and constraint rule feedback, on the one hand, utilize a multi-stage intelligent decision-making generation architecture for heavy-haul railways to process knowledge retrieval, parameter derivation, and standard or scheme generation in a layered manner. This enables the large-scale heavy-haul railway transportation model to fully integrate facility and equipment parameters and regulations, completing technical standard derivation and transportation scheme generation. On the other hand, by introducing a verification mechanism based on constraint rule feedback, the candidate standards or candidate schemes output by the large-scale heavy-haul railway transportation model are automatically verified, and secondary reasoning is triggered when verification fails, thereby effectively improving the reliability of the reasoning results in terms of safety, accuracy, and engineering applicability. This application, based on a computing platform and with a large-scale heavy-haul railway model as its core, achieves complex business processes such as line technical standard derivation and transportation scheme generation based on transportation expert knowledge question answering through collaborative design of a multi-stage intelligent decision-making generation architecture and constraint rule feedback mechanism. This reduces reliance on human experience while improving the processing efficiency, intelligence level, and autonomous controllability of heavy-haul railway transportation organization. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0033] Figure 1 This is an application environment diagram of a method for implementing a large-scale heavy-haul railway transportation model that integrates multi-stage decision-making and constraint rule feedback in one embodiment.

[0034] Figure 2 This is a flowchart illustrating a method for implementing a large-scale heavy-haul railway transportation model that integrates multi-stage decision-making and constraint rule feedback in one embodiment.

[0035] Figure 3 This is a flowchart illustrating the model training steps in one embodiment;

[0036] Figure 4 This is a flowchart illustrating a specific embodiment of a method for implementing a large-scale heavy-haul railway transportation model that integrates multi-stage decision-making and constraint rule feedback.

[0037] Figure 5 This is a technical roadmap for a large-scale heavy-haul railway transportation model implementation system that integrates multi-stage decision-making and constraint rule feedback in an application embodiment.

[0038] Figure 6 This is a structural block diagram of a large-scale heavy-haul railway transportation model implementation system that integrates multi-stage decision-making and constraint rule feedback in one embodiment.

[0039] Figure 7 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0040] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0041] The implementation method for a large-scale heavy-haul railway transportation model that integrates multi-stage decision-making and constraint rule feedback provided in this application can be applied to, for example... Figure 1 The application environment shown illustrates this. In this environment, the terminal can communicate with the server via a network. The data storage system can store the data that the server needs to process. The data storage system can be integrated onto the server or located on the cloud or other network servers. In situations such as... Figure 1 In the application environment shown, the terminal can be, but is not limited to, various personal computers, laptops, smartphones, and tablets. The server can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0042] In one embodiment, such as Figure 2 As shown, a large-scale model method for heavy-haul railway transportation that integrates multi-stage decision-making and constraint rule feedback is presented. This method can be applied to... Figure 1 In the terminal, the method may include the following steps:

[0043] Step S201: Determine the base model to be trained, train the base model through a preset three-stage training system to obtain a large heavy-haul railway transportation model with specialized reasoning capabilities for heavy-haul transportation scenarios, and adapt the large heavy-haul railway transportation model to the computing platform.

[0044] It is important to note that the selection of the base model is the core starting point for the entire system construction. When migrating or customizing to a specific domain, an excellent base model can effectively reduce the dependence on the scale of domain data and annotation costs. Knowledge transfer and capability alignment can be achieved through minimal fine-tuning or hint learning, thereby significantly improving the overall performance, stability, and scalability of the system.

[0045] Specifically, firstly, at the computing power platform and model level, the terminal constructs a domain expert model for heavy-haul railway transportation scenarios through a three-stage training system that includes "domain data generation and standardization - incremental pre-training - supervised fine-tuning". Secondly, it deeply adapts the training and inference process of the heavy-haul transportation model to the architectural features of domestic CPUs, GPUs and AI accelerators in terms of operator implementation and parallel strategies. Finally, it completes the entire process of localized training and inference on the domestic computing power platform, enabling the model to systematically learn core knowledge such as heavy-haul railway equipment parameters, transportation organization schemes, industry regulations and technical standards, and possess a deep understanding of heavy-haul transportation scenarios and professional reasoning capabilities, thus realizing intelligent question answering by transportation experts.

[0046] As an example, the heavy-haul railway transportation large model can use the Kunpeng 920 CPU cluster and Ascend 910B AI cluster as a powerful computing base model, openEuler and CANN as a stable and reliable system and driving foundation, and MindSpore, MindSpeed-LLM and MindIE as the core computing and acceleration middleware, together forming a large model support platform that is software and hardware collaborative, autonomous and controllable, and covers the entire life cycle from training to inference.

[0047] (1) Hardware infrastructure architecture

[0048] The hardware system adopts a layered deployment strategy, consisting of three parts: an AI analysis server that undertakes core computing tasks, a front-end server that provides application services, and an edge computing unit responsible for real-time on-site processing, in order to meet the different performance and latency requirements of training, inference, and edge deployment. The specific configuration is shown in Table 1.

[0049] Table 1 Hardware Configuration Table

[0050]

[0051] The application server serves as the system front-end, responsible for hosting applications such as web services, API interfaces, and user management.

[0052] The AI ​​analytics server, acting as a central computing node, undertakes distributed training, batch inference, and deep data analysis tasks for large models. Extensive memory is used for loading large models; high-speed NVMe storage ensures the retrieval of massive training data; and a high-speed interconnect network is key to achieving efficient parallel computing across multiple GPUs and reducing communication overhead.

[0053] Edge computing units serve as inference nodes at the station end. Deployed at the data source or business site, they are responsible for the high-availability deployment of lightweight models, performing real-time, low-latency micro-data inference and analysis, and meeting the business's requirements for privacy, real-time performance, and bandwidth conservation.

[0054] (2) System software stack and core components

[0055] The software environment is built on a domestic open-source operating system and uses Huawei Ascend computing platform as the underlying hardware layer. It is equipped with a full-scenario AI framework and a large model acceleration library, forming a complete large model support stack. The versions and functions of each key component are shown in Table 2.

[0056] Table 2 Software Configuration Table

[0057]

[0058] The operating system can be Huawei openEuler 22.03 LTS, a domestic Linux distribution designed for server scenarios. It offers long-term support and features high performance, high security, and native optimization and compatibility with Kunpeng and Ascend hardware.

[0059] CANN is the Ascend computing architecture and operator library, responsible for managing the computing, storage, and communication resources of the Ascend AI processor. It provides a highly optimized deep learning operator library and is the cornerstone for upper-layer AI frameworks to leverage hardware performance.

[0060] The MindSpore full-scenario AI framework supports deployment across edge, cloud, and device environments. It is responsible for the distributed training, debugging, and export of large models and serves as a key layer for decoupling algorithms from hardware.

[0061] MindIE is the Ascend inference engine, which provides high-performance, low-latency local inference services for large models. It supports mainstream model formats and has been deeply optimized for Ascend chips. It is the core component for model deployment in the production environment of this system.

[0062] MindSpeed-LLM is a large model acceleration framework built on MindSpore and CANN. It provides high-performance Transformer model components, dynamic batching, and memory optimization techniques to improve the efficiency of large model task processing.

[0063] Step S202: Construct a multi-stage intelligent decision generation architecture for heavy-haul railway transportation, dividing the reasoning process of the heavy-haul railway transportation model into a knowledge retrieval stage, a parameter derivation stage, and a standard or scheme generation stage, which are executed sequentially.

[0064] Specifically, at the application architecture level, the terminal constructs a multi-stage intelligent decision-making generation architecture for heavy-haul railways. This architecture includes three stages: knowledge retrieval, parameter derivation, and standard or scheme generation. In the knowledge retrieval stage, knowledge in areas such as railway infrastructure data and technical regulations is systematically organized, semantically understood, and accurately retrieved. In the parameter derivation stage, optimization algorithms are used to calculate and derive key technical parameters such as locomotive traction quality, train frequency, and train formation structure. Finally, in the standard or scheme generation stage, the heavy-haul railway transportation big data model integrates the knowledge retrieval and parameter derivation results to achieve intelligent reasoning and generation of railway technical standards and transportation organization schemes.

[0065] Step S203 introduces a reasoning verification mechanism based on constraint rule feedback to verify the compliance and security of candidate standards or candidate solutions output during the standard or solution generation stage.

[0066] Specifically, during the solution generation process, the system introduces a verification mechanism based on constraint rule feedback to automatically review the solution results output by the large model, focusing on verifying the consistency of technical parameters, regulatory compliance, and operational safety. Simultaneously, this layer incorporates a multi-level constraint rule system, including hard constraints (such as national standards, industry specifications, and equipment limit parameters) and soft constraints (such as empirical rules and operational efficiency preferences), performing compliance and consistency verification in real time during parameter calculation.

[0067] When the derivation results do not meet the constraints, the system will automatically record the cause of the conflict and return a correction prompt to the large model through the constraint rule feedback mechanism. This triggers parameter recalculation or adjustment of input assumptions, thus forming a closed-loop reasoning process of "derivation-verification-correction" until the preset engineering and safety constraints are met. This layer effectively avoids the problem of relying solely on language generation without engineering calculation support, providing a reliable and verifiable technical parameter foundation for intelligent question answering by transportation experts, route technical standard derivation, and transportation scheme generation.

[0068] Step S204: If the verification passes, the candidate standard is used as the standard output result or the candidate solution is used as the solution output result; if the verification fails, a correction prompt is automatically generated and secondary inference and parameter recalculation are triggered to obtain the standard correction result or the solution correction result.

[0069] Specifically, the terminal obtains the verification results of compliance and security checks. If the verification is passed, the candidate standard is used as the standard output result or the candidate solution is used as the solution output result. If the verification fails, a correction prompt is automatically generated and secondary inference and parameter recalculation are triggered to obtain the standard correction result or the solution correction result.

[0070] In this embodiment, on the one hand, a multi-stage intelligent decision-making generation architecture for heavy-haul railways is adopted, which layers knowledge retrieval, parameter derivation, and standard or scheme generation. This enables the large-scale heavy-haul railway transportation model to fully integrate facility and equipment parameters and regulations, completing the derivation of technical standards and the generation of transportation schemes. On the other hand, by introducing a verification mechanism based on constraint rule feedback, the candidate standards or candidate schemes output by the large-scale heavy-haul railway transportation model are automatically verified, and secondary reasoning is triggered when verification fails, thereby effectively improving the reliability of the reasoning results in terms of safety, accuracy, and engineering applicability. This application is based on a domestic computing power platform and takes a large-scale model in the field of heavy-haul railways as its core. Through the collaborative design of a multi-stage intelligent decision-making generation architecture and a constraint rule feedback mechanism, and based on transportation expert knowledge question answering, it realizes complex business operations such as line technical standard derivation and transportation scheme generation, reducing reliance on human experience while improving the processing efficiency, intelligence level, and autonomous controllability of heavy-haul railway transportation organization.

[0071] In one embodiment, such as Figure 3 As shown, in step S201 above, the base model is trained through a preset three-stage training system to obtain a large heavy-haul railway transportation model with specialized reasoning capabilities for heavy-haul transportation scenarios. This may include the following steps:

[0072] Step S301: Obtain multi-source heterogeneous data on heavy-haul railway equipment and transportation operations, perform format conversion and data cleaning on the multi-source heterogeneous data, and form a standardized domain corpus.

[0073] Step S302: Based on the standardized domain corpus, the base model learns key knowledge in the transportation domain through a continuous language modeling task to obtain an initial heavy-haul railway transportation model.

[0074] Step S303: The initial heavy-haul railway transportation model is supervised and fine-tuned using a low-rank adaptive fine-tuning method to obtain the heavy-haul railway transportation model.

[0075] Specifically, the terminal trains a domain expert model through a full-process training process, enabling it to learn relevant knowledge in the heavy-haul transportation field, understand heavy-haul railway transportation scenarios, and provide professional decision support. To ensure the model possesses the professional understanding and task capabilities required for heavy-haul transportation scenarios, this embodiment constructs a three-stage training system consisting of "domain data generation—incremental pre-training—supervised fine-tuning." This system effectively transfers general-purpose models to the heavy-haul transportation domain and is a key process for achieving intelligent question answering by transportation experts. The specific steps are as follows:

[0076] (1) Data acquisition and standardization

[0077] The data sources for heavy-haul railway equipment and transportation operations are complex and diverse, encompassing various media such as engineering design documents, transportation plans, industry standards and specifications, and textbooks and research literature. To achieve systematic data integration, standardized preprocessing of all heterogeneous data is required. First, textual data is converted to structured Markdown format, and table fields and formulas are parsed. Scanned document images are processed using OCR (Optical Character Recognition) technology to extract text, which is then manually proofread. Paper documents are digitized and then subjected to OCR processing. Subsequently, a rigorous data cleaning process is implemented to remove redundant, erroneous, inconsistent, and missing information from the original data. Deduplication uses a text similarity-based algorithm to identify duplicates, retaining data with high version completeness. Error correction combines spell detection, grammar analysis, and standardized vocabulary for automatic replacement, with final manual review.

[0078] (2) Incremental pre-training

[0079] The incremental pre-training phase, based on standardized domain corpora, uses continuous language modeling tasks to enable the model to fully learn key knowledge such as line equipment parameters, transportation organization schemes, industry regulations, and relevant standards. This process not only allows the model to master the terminology system, business concepts, and semantic structures unique to the heavy-haul railway field, but also enables it to establish relational representations between domain knowledge, achieving an effective evolution from a "general language model" to a "domain knowledge model." Compared to training a large model from scratch, incremental pre-training can build a deep cognitive ability of the model for heavy-haul transportation scenarios while significantly shortening training time and reducing computational resource consumption.

[0080] (3) Supervision and fine-tuning

[0081] Supervised fine-tuning employs the LoRA (Low-Rank Adaptation) fine-tuning method, which introduces only incremental parameters that are much smaller than the original parameter size while freezing the pre-trained weights. This reduces the fine-tuning cost and maintains the stability of the model structure.

[0082] In one embodiment, step S302 above, which involves using a low-rank adaptive fine-tuning method to supervise and fine-tune the initial heavy-haul railway transportation model to obtain the heavy-haul railway transportation model, may include the following steps:

[0083] In the base model, a target weight matrix is ​​selected, and two low-rank matrices are introduced for each target weight matrix. An increment matrix is ​​constructed based on the low-rank matrices. In the forward computation of the base model, the original output of the base model is added to the output of the increment matrix. During the supervised fine-tuning process, the parameters of the low-rank matrices are updated and the target weight matrices are kept frozen. The updated low-rank matrices are then merged with the target weight matrices to obtain the large model of heavy-haul railway transportation.

[0084] Among them, the low-rank adaptive method is a parameter-efficient fine-tuning method that aims to efficiently adapt a large pre-trained model by introducing low-rank matrices, significantly reducing the number of training parameters while maintaining model performance. Its core principle is to freeze the original weights of the pre-trained model and add low-rank matrices only next to the weight matrices of specific layers. By training these low-rank matrices, task-specific feature variations can be captured.

[0085] Specifically, the fine-tuning process of LoRA is as follows:

[0086] Step 1: Select the target layer. Common choices include the Query / Key / Value projection matrix or the fully connected layer in the Transformer model. Prioritize fine-tuning the projection matrix of the attention layer.

[0087] Step 2: Decompose the original weights: Create two low-rank matrices A and B for each target weight matrix. A: Initialize using a Gaussian distribution, and B: Initialize as a matrix of all zeros to ensure that the original model output is not disturbed during the initial training phase.

[0088] Step 3: Modify forward propagation: In the forward computation of the target layer, add the original output to the output of the low-rank matrix;

[0089] Step 4: Freeze the original parameters: During the supervised fine-tuning process, only the parameters of A and B are updated, and the original weights remain frozen;

[0090] Step 5: Supervised fine-tuning process: The learning rate setting is usually smaller than that for full parameter fine-tuning, and the batch size can be larger due to the low memory usage;

[0091] Step 6, Parameter Merging (Inference Stage): Merge A and B with the original weights to obtain new weights. There is no additional computational overhead during inference, and the speed is consistent with the original model. It can be exported as a single model file for easy deployment.

[0092] The advantages of LoRA are as follows:

[0093] 1. Reduce the number of training parameters: LoRA only trains a small number of new parameters (usually about 0.1%-1% of the original model).

[0094] 2. Faster training speed: Due to fewer parameter updates, training iterations are faster, and it can quickly adapt to a small amount of task data.

[0095] 3. Flexible deployment: LoRA can be merged into the original model during inference, or dynamically loaded / unloaded without having to save the entire model again.

[0096] In one embodiment, the method of this application further includes the following steps:

[0097] A dedicated knowledge base for heavy-haul railway transportation operations is constructed. This knowledge base enables structured management and vectorized indexing of various types of data. During the knowledge retrieval phase, the heavy-haul railway transportation big data model searches the dedicated knowledge base using a combination of semantic and keyword retrieval based on the user-input computational task requirements and intermediate requests generated by the system. It also dynamically retrieves relevant domain knowledge as the knowledge retrieval results.

[0098] It should be noted that knowledge retrieval is the foundation of the multi-stage decision generation architecture of the heavy-haul railway transportation big model. Its core function is to achieve unified management, in-depth understanding and precise retrieval of knowledge in the field of heavy-haul railways, providing authoritative and traceable knowledge basis for subsequent parameter derivation and scheme generation.

[0099] Specifically, this stage first requires building a dedicated knowledge base for heavy-haul railway transportation operations, centrally storing and structurally managing multi-source data such as regulations, technical standards, line parameters, transportation organization plans, and historical cases. The system supports access to various data types, including text, tables, drawings, and question-and-answer pairs, and uses natural language processing technology to clean, denoise, and semantically annotate the raw data.

[0100] The knowledge base construction supports multiple segmentation strategies, including structured segmentation based on title level, intelligent segmentation based on character length, and custom segmentation based on business rules. Users can manually correct the segmentation results after previewing them to ensure the integrity and professional accuracy of the knowledge units. After segmentation, the knowledge content will undergo semantic embedding processing through a vectorization model to form a vector index that can be efficiently retrieved by large models. The knowledge base can also be re-vectorized after model upgrades or algorithm adjustments.

[0101] During the retrieval phase, based on the user-input transportation task requirements or intermediate requests generated by the system, a combination of semantic and keyword retrieval is used to dynamically retrieve the most relevant regulations, technical constraints, and empirical rules from the knowledge base. These are then output in a standardized data structure for the subsequent parameter derivation and solution generation phases. This mechanism ensures that the large model's reasoning process is always based on authentic, authoritative, and controllable domain knowledge, avoiding reasoning results that deviate from professional standards.

[0102] In one embodiment, the method of this application further includes the following steps:

[0103] In the parameter derivation stage, the heavy-haul railway transportation big model, based on the knowledge retrieval results and combined with the specific transportation task objectives, uses mathematical models and optimization algorithms from the tool library to quantify and calculate the key technical parameters of the objectives, thereby obtaining the parameter derivation results.

[0104] The key technical parameters of the target may include locomotive traction mass, number of train pairs, and formation structure.

[0105] It should be noted that parameter derivation is a key link connecting domain knowledge with transportation scheme generation. Its main function is to calculate and derive the key technical parameters involved in heavy-haul railway transportation.

[0106] Specifically, after receiving information such as route conditions, equipment parameters, and regulatory restrictions from the knowledge retrieval output, this stage, in conjunction with specific transportation task objectives, calculates key technical parameters such as locomotive traction quality, train operation pairs, and formation structure using mathematical models and optimization algorithms from the tool library. The parameter calculation process can be dynamically configured according to different routes, locomotive and rolling stock types, and different transportation organization modes, exhibiting good versatility and scalability. Parameter derivation adopts a modular design, with different parameter calculation modules operating independently and being called on demand, outputting results through a unified data interface.

[0107] In one embodiment, the method of this application further includes the following steps:

[0108] During the standard or scheme generation stage, the heavy-haul railway transportation big model is used as the core. Semantic information obtained from knowledge retrieval results and quantitative results obtained from parameter derivation results are integrated to generate line technical standards or transportation organization schemes that meet engineering constraints and business objectives, which can then be used as candidate standards or candidate schemes.

[0109] It should be noted that the scheme generation is a high-level reasoning and decision-making stage of the multi-stage decision generation architecture. Its main function is to complete the intelligent reasoning and generation of transportation expert intelligent question answering, heavy-haul railway line technical standards and transportation schemes based on the full integration of domain knowledge and parameter derivation results.

[0110] Specifically, this stage uses a large model in the field of heavy-duty transportation as the core reasoning engine. It integrates the semantic information such as regulations and route conditions output from the knowledge retrieval stage with the quantitative calculation results output from the parameter derivation stage to form a unified model. Through multi-task decomposition and collaborative reasoning mechanisms, it generates transportation organization plans that meet engineering constraints and business objectives, and presents them in a combination of structured and visual forms, realizing intelligent support for the entire process from demand input to plan output.

[0111] In one embodiment, such as Figure 4 As shown, a specific embodiment of a method for implementing a large-scale heavy-haul railway transportation model that integrates multi-stage decision-making and constraint rule feedback is provided, which specifically includes the following steps:

[0112] Step S401: Determine the base model to be trained, acquire multi-source heterogeneous data on heavy-haul railway equipment and transportation business, perform format conversion and data cleaning on the multi-source heterogeneous data to form a standardized domain corpus; based on the standardized domain corpus, through a continuous language modeling task, enable the base model to learn key knowledge in the transportation domain and obtain an initial large model of heavy-haul railway transportation.

[0113] Step S402: Select the target weight matrix in the base model, introduce two low-rank matrices for each target weight matrix, and construct the increment matrix based on the low-rank matrix; in the forward calculation of the base model, add the original output of the base model to the output of the increment matrix; during the supervised fine-tuning process, update the parameters of the low-rank matrix and keep the target weight matrix frozen, merge the updated low-rank matrix with the target weight matrix to obtain the heavy-haul railway transportation large model, and adapt the heavy-haul railway transportation large model to the computing platform.

[0114] Step S403: Construct a multi-stage intelligent decision generation architecture for heavy-haul railway transportation, dividing the reasoning process of the heavy-haul railway transportation model into a knowledge retrieval stage, a parameter derivation stage, and a standard or scheme generation stage, which are executed sequentially.

[0115] Step S404 introduces a reasoning verification mechanism based on constraint rule feedback to verify the compliance and security of candidate standards or candidate solutions output during the standard or solution generation stage; if the verification passes, the candidate standard is used as the standard output result or the candidate solution is used as the solution output result; if the verification fails, a correction prompt is automatically generated and secondary reasoning and parameter recalculation are triggered to obtain the standard correction result or the solution correction result.

[0116] The beneficial effects of the above embodiments are as follows:

[0117] (1) Improve the work efficiency of staff and the system's decision-making ability. By introducing a large model of heavy-haul railway transportation, we can achieve deep semantic understanding and comprehensive reasoning of complex transportation scenarios, assist staff in making decisions, and improve the automation and intelligence level of transportation organization.

[0118] (2) To achieve unified management and efficient reuse of domain knowledge, systematically organize and semantically model knowledge in areas such as line facility data, technical regulations and industry standards, solve the problems of knowledge dispersion and difficulty in collaboration, and improve knowledge retrieval efficiency and application consistency;

[0119] (3) Support for automatic derivation of key parameters and intelligent generation of schemes: Through a multi-stage decision-making architecture and optimization algorithm, key parameters such as locomotive traction quality, train formation and number of trains are automatically deduced, and transportation schemes are intelligently generated to improve the efficiency and scientific nature of scheme formulation.

[0120] (4) Enhance the compliance and reliability of results by introducing a constraint rule feedback and verification mechanism to verify the compliance, security and consistency of model output, and trigger secondary inference correction when necessary, so as to effectively ensure the reliability and engineering availability of results;

[0121] (5) Highly professional domain expert model. Incremental pre-training and fine-tuning enable the model to possess professional knowledge of heavy-haul railways, which significantly improves the accuracy of transportation scheme generation and equipment parameter deduction compared to open-source models;

[0122] (6) It can achieve deep adaptation and independent control of domestic computing power platforms. Hardware, operating system, AI framework and large model can all be domestic solutions. It can complete the localized deployment of the entire process of training and inference of domain large model. It can run independently in a closed network environment and meet the application requirements of high security and independent control.

[0123] To more clearly illustrate the implementation method of the large-scale heavy-haul railway transportation model integrating multi-stage decision-making and constraint rule feedback provided in the embodiments of this application, the following specific description uses an application embodiment to illustrate this implementation method. In one embodiment, such as Figure 5 As shown, this application also provides a method for implementing a large-scale heavy-haul railway transportation model that integrates multi-stage decision-making and constraint rule feedback, specifically including the following steps:

[0124] (1) Data acquisition:

[0125] The types and quantities of data are shown in Table 3:

[0126] Table 3 Publicly Available Data in the Field

[0127]

[0128] (2) Selection of base model:

[0129] In this embodiment, the domestically developed large model DeepSeek-R1-Distill-Llama-70B can be selected as the base, mainly based on its performance advantages in language understanding, logical reasoning, and adaptation to the domestic ecosystem.

[0130] (3) Model pre-training and fine-tuning training:

[0131] The training data used for pre-training consisted of a specially compiled corpus related to heavy-haul railway transportation. After data format conversion, noise filtering, relevance testing, and repeatability testing, approximately 15,000 data entries were used for training. The initial learning rate was lr = 1e-4, the training precision was bfloat16, data-parallel training was used, DeepSpeed ​​ZeRO-3 was used for optimization, the gradient accumulation step was 16, the validation set split ratio was 0.1, and the loss function was the cross-entropy loss function.

[0132] The fine-tuning training model selected was a large model after incremental pre-training. The fine-tuning dataset was compiled from relevant corpora in the field of heavy-haul railway transportation, totaling 10,000 question-answer pairs, with a validation set split ratio of 0.1. The initial learning rate was lr = 1e-4. During training, a cosine similarity strategy was used to dynamically adjust the learning rate, gradually decreasing it as training progressed to ensure the stability of the model when learning professional knowledge and to avoid excessive fluctuations in model parameters due to an excessively large learning rate, which would affect the training effect. The gradient accumulation steps were 16, and the LoRA strategy was used for fine-tuning. The LoRA configuration parameters are shown in Table 4.

[0133] Table 4 LoRA Fine-tuning Configuration Parameters

[0134]

[0135] (4) Model inference performance:

[0136] Table 5 shows the inference performance metrics of the model after deployment on the server. The results indicate that the model exhibits good stability and throughput under different concurrency conditions, with a 100% success rate across all concurrency levels, demonstrating reliable service capabilities. Regarding token generation, the model's output speed increased from 26.78 tokens / s to 481.24 tokens / s with increasing concurrency, demonstrating good concurrent scalability and effectively utilizing batch inference to improve overall generation efficiency. Furthermore, the system maintained an acceptable response time even with 100 concurrent requests, without significant performance degradation. The initial response time and subsequent response times increased with increasing concurrency, but the overall fluctuation remained stable, with no timeouts or errors, indicating a reasonable queuing and scheduling mechanism design.

[0137] Generally speaking, the model can still maintain a high success rate, high token generation speed and controllable latency under high concurrency conditions, demonstrating good engineering stability and inference service capabilities, and is suitable for handling a large number of continuous information inference requirements in the heavy-haul railway transportation scenario.

[0138] Table 5 Model Inference Performance Test

[0139]

[0140] (5) Model evaluation metrics:

[0141] To quantitatively evaluate the performance of the model trained on the corpus in the field of heavy-haul railway transportation, two widely recognized automated evaluation metrics, BLEU and ROUGE, are adopted. Among them, the BLEU metric focuses on measuring the matching degree of the generated text and the reference text in terms of n-gram precision; the ROUGE metric focuses on evaluating the coverage of the key information of the reference text by the generated text from the perspective of recall rate. The experimental results show that the trained model significantly outperforms the open-source model in the listed evaluation metrics. Specifically, in terms of the Bleu-1 metric, the score of the trained model is improved by 0.118 compared with the open-source model, indicating that the trained model performs better in terms of the matching degree and generation quality at the lexical level. Rouge-1, Rouge-2 and Rouge-L have achieved improvements of 0.297, 0.189 and 0.290 respectively, reflecting that the trained model has significant advantages in terms of information coverage, sentence continuity and overall structural similarity. Considering these metrics comprehensively, it can be seen that the trained model can not only generate content more in line with the reference text, but also perform more prominently in capturing key semantics and maintaining text coherence, indicating that the adopted training method effectively improves the generation ability of the model in the target task. The model evaluation metrics are shown in Table 6.

[0142] Table 6 Model Evaluation Metrics

[0143]

[0144] (6) Intelligent Q&A for Transportation Experts:

[0145] Input question:

[0146] What does "average transit time" mean?

[0147] Model answer:

[0148] Definition:

[0149] Average transit time (denoted as t 中 ) is an indicator to measure the average residence time of transit wagons (i.e., freight wagons that need to be reorganized) in the station, reflecting the efficiency of a single transit operation of freight wagons.

[0150] formula:

[0151]

[0152] In the above formula, This represents the total parking hours of all connecting vehicles during the reporting period. This refers to the total number of transit vehicles during the reporting period; and The total dwell hours for both un-adjusted and adjusted transfer vehicles are respectively. and These represent the number of vehicles transferring between different modes of transport, namely, those without transfers and those with transfers.

[0153] Features:

[0154] The time with adjustment (requiring revision of assignments) is much longer than the time without adjustment (passing directly without revision), which is the main factor affecting the time.

[0155] Measures to shorten transit time: Increase the "no-tuning ratio" (the proportion of no-tuning transit vehicles to the total number of incoming vehicles).

[0156] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0157] Based on the same inventive concept, this application also provides a system for implementing the above-mentioned method for implementing a large-scale heavy-haul railway transportation model that integrates multi-stage decision-making and constraint rule feedback. The solution provided by this system is similar to the implementation scheme described in the above method. Therefore, the specific limitations of one or more embodiments of the system for implementing a large-scale heavy-haul railway transportation model that integrates multi-stage decision-making and constraint rule feedback provided below can be found in the limitations of the method for implementing a large-scale heavy-haul railway transportation model that integrates multi-stage decision-making and constraint rule feedback described above, and will not be repeated here.

[0158] In one exemplary embodiment, such as Figure 6As shown, a large-scale model implementation system for heavy-haul railway transportation that integrates multi-stage decision-making and constraint rule feedback is provided. This system may include:

[0159] The model training module 601 is used to determine the base model to be trained, train the base model through a preset three-stage training system, obtain a large heavy-haul railway transportation model with specialized reasoning capabilities for heavy-haul transportation scenarios, and adapt the large heavy-haul railway transportation model to a domestic computing power platform.

[0160] Architecture generation module 602 is used to build a multi-stage intelligent decision generation architecture for heavy-haul railway transportation, which divides the reasoning process of the large model of heavy-haul railway transportation into a knowledge retrieval stage, a parameter derivation stage and a standard or scheme generation stage executed sequentially.

[0161] The inference verification module 603 is used to introduce an inference verification mechanism based on constraint rule feedback to verify the compliance and security of candidate standards or candidate solutions output during the standard or solution generation stage.

[0162] The inference recalculation module 604 is used to output the candidate standard as the standard output result or the candidate solution as the solution output result when the verification passes; when the verification fails, it automatically generates a correction prompt and triggers secondary inference and parameter recalculation to obtain the standard correction result or the solution correction result.

[0163] In one embodiment, the model training module 601 is further configured to acquire multi-source heterogeneous data on heavy-haul railway equipment and transportation operations, perform format conversion and data cleaning on the multi-source heterogeneous data to form a standardized domain corpus; based on the standardized domain corpus, through a continuous language modeling task, enable the base model to learn key knowledge in the transportation domain to obtain an initial heavy-haul railway transportation model; and use a low-rank adaptive fine-tuning method to perform supervised fine-tuning on the initial heavy-haul railway transportation model to obtain a large heavy-haul railway transportation model.

[0164] In one embodiment, the model training module 601 is further configured to select target weight matrices in the base model, introduce two low-rank matrices for each target weight matrix, construct an increment matrix based on the low-rank matrices, add the original output of the base model to the output of the increment matrix in the forward computation of the base model, update the parameters of the low-rank matrices and keep the target weight matrices frozen during the supervised fine-tuning process, and merge the updated low-rank matrices with the target weight matrices to obtain a large model for heavy-haul railway transportation.

[0165] In one embodiment, the system may further include: a knowledge retrieval module for constructing a dedicated knowledge base for heavy-haul railway transportation business, and for performing structured management and vectorized indexing of various types of data through the dedicated knowledge base; in the knowledge retrieval stage, the heavy-haul railway transportation big model performs retrieval in the dedicated knowledge base by combining semantic retrieval and keyword retrieval based on the computational task requirements input by the user and the intermediate requests generated by the system, and dynamically recalls relevant domain knowledge as the knowledge retrieval results.

[0166] In one embodiment, the system may further include: a parameter derivation module, which, during the parameter derivation stage, uses the heavy-haul railway transportation big model to quantitatively calculate the key technical parameters of the target based on the knowledge retrieval results and in combination with the specific transportation task objectives, by calling mathematical models and optimization algorithms in the tool library, to obtain the parameter derivation results.

[0167] In one embodiment, the system may further include: a result generation module, used in the standard or scheme generation stage, taking the heavy-haul railway transportation big model as the core, integrating semantic information obtained from knowledge retrieval results and quantitative results obtained from parameter derivation results, to generate line technical standards or transportation organization schemes that meet engineering constraints and business objectives, as candidate standards or candidate schemes.

[0168] The various modules in the aforementioned large-scale heavy-haul railway transportation model implementation system, which integrates multi-stage decision-making and constraint rule feedback, can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the computer device's memory as software, so that the processor can call and execute the corresponding operations of each module.

[0169] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 7As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a method for realizing a large-scale heavy-haul railway transportation model that integrates multi-stage decision-making and constraint rule feedback. The display unit of the computer device is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0170] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0171] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0172] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0173] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0174] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0175] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0176] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0177] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for implementing a large-scale heavy-haul railway transportation model that integrates multi-stage decision-making and constraint rule feedback, characterized in that, The method includes: The base model to be trained is determined, and the base model is trained through a preset three-stage training system to obtain a large heavy-haul railway transportation model with specialized reasoning capabilities for heavy-haul transportation scenarios. The large heavy-haul railway transportation model is then adapted to the computing power platform. A multi-stage intelligent decision generation architecture for heavy-haul railway transportation is constructed, and the reasoning process of the heavy-haul railway transportation model is divided into a knowledge retrieval stage, a parameter derivation stage, and a standard or scheme generation stage, which are executed sequentially. A reasoning verification mechanism based on constraint rule feedback is introduced to verify the compliance and security of the candidate standards or candidate solutions output during the standard or solution generation stage. If the verification passes, the candidate standard is used as the standard output result or the candidate solution is used as the solution output result; if the verification fails, a correction prompt is automatically generated and secondary inference and parameter recalculation are triggered to obtain the standard correction result or the solution correction result.

2. The method according to claim 1, characterized in that, The process of training the base model through a pre-defined three-stage training system to obtain a large-scale heavy-haul railway transportation model with specialized reasoning capabilities for heavy-haul transportation scenarios includes: Acquire multi-source heterogeneous data on heavy-haul railway equipment and transportation operations, perform format conversion and data cleaning on the multi-source heterogeneous data, and form a standardized domain corpus; Based on the standardized domain corpus, the base model learns key knowledge in the transportation domain through a continuous language modeling task, thus obtaining an initial large-scale heavy-haul railway transportation model. The initial heavy-haul railway transportation model is subjected to supervised fine-tuning using a low-rank adaptive fine-tuning method to obtain the final heavy-haul railway transportation model.

3. The method according to claim 2, characterized in that, The method of employing a low-rank adaptive fine-tuning method to supervise the fine-tuning of the initial heavy-haul railway transportation model, thereby obtaining the heavy-haul railway transportation model, includes: In the base model, a target weight matrix is ​​selected, and two low-rank matrices are introduced for each target weight matrix. An incremental matrix is ​​constructed based on the low-rank matrices. In the forward computation of the base model, the original output of the base model is added to the output of the increment matrix; During the supervised fine-tuning process, the parameters of the low-rank matrix are updated while the target weight matrix is ​​kept frozen. The updated low-rank matrix is ​​then merged with the target weight matrix to obtain the large model of heavy-haul railway transportation.

4. The method according to claim 1, characterized in that, The method further includes: Construct a dedicated knowledge base for heavy-haul railway transportation operations, and use the dedicated knowledge base to perform structured management and vectorized indexing of various types of data; During the knowledge retrieval stage, the heavy-haul railway transportation big data model searches the dedicated knowledge base using a combination of semantic retrieval and keyword retrieval based on the user-input computational task requirements and the intermediate requests generated by the system, and dynamically recalls relevant domain knowledge as the knowledge retrieval results.

5. The method according to claim 4, characterized in that, The method further includes: In the parameter derivation stage, the heavy-haul railway transportation big model, based on the knowledge retrieval results and combined with the specific transportation task objectives, uses mathematical models and optimization algorithms from the tool library to quantify and calculate the key technical parameters of the objectives, thereby obtaining the parameter derivation results.

6. The method according to claim 5, characterized in that, The method further includes: During the standard or scheme generation stage, the heavy-haul railway transportation big model is used as the core. Semantic information obtained from the knowledge retrieval results and quantitative results obtained from the parameter derivation results are integrated to generate line technical standards or transportation organization schemes that meet engineering constraints and business objectives, which serve as candidate standards or candidate schemes.

7. A system for implementing a large-scale heavy-haul railway transportation model that integrates multi-stage decision-making and constraint rule feedback, characterized in that, The system includes: The model training module is used to determine the base model to be trained, train the base model through a preset three-stage training system to obtain a large heavy-haul railway transportation model with specialized reasoning capabilities for heavy-haul transportation scenarios, and adapt the large heavy-haul railway transportation model to the computing power platform. The architecture generation module is used to construct a multi-stage intelligent decision generation architecture for heavy-haul railway transportation, which divides the reasoning process of the heavy-haul railway transportation model into a knowledge retrieval stage, a parameter derivation stage, and a standard or scheme generation stage that are executed sequentially. The inference verification module is used to introduce an inference verification mechanism based on constraint rule feedback to verify the compliance and security of the candidate standards or candidate solutions output during the standard or solution generation stage. The inference recalculation module is used to output the candidate standard as the standard output result or the candidate scheme as the scheme output result when the verification passes; when the verification fails, it automatically generates a correction prompt and triggers secondary inference and parameter recalculation to obtain the standard correction result or the scheme correction result.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.