Rail transit field large model collaborative evolution training method and system

CN122819360APending Publication Date: 2026-09-25BEIJING HOLLYSYS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610722866.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-25
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]有鉴于此,本申请实施例提供了一种轨道交通领域大模型协同演化训练方法及系统,以解决现有技术存在的领域知识注入与指令保持失衡、小样本意图激活困难、适配器迁移失真的问题

Benefits of technology

通过获取轨道交通领域异构语料,对异构语料进行功能分桶,得到领域隐性知识数据、通用逻辑平衡数据和业务意图指令数据;基于预设标识符规则识别领域隐性知识数据中的轨道交通专业标识符,将轨道交通专业标识符加入分词器新增词元,并对领域隐性知识数据进行噪声清洗和重叠窗口切分;冻结基座模型主干权重,基于领域隐性知识数据和通用逻辑平衡数据的动态混合批次,对高秩领域适配矩阵进行持续预训练,生成领域知识适配器;将领域知识适配器挂载至指令模型对应的变换层旁路,冻结指令模型主干权重,并基于次词预测损失和输出分布散度约束对领域知识适配器进行对齐训练,生成融合模型;冻结融合模型中的领域知识适配器,挂载低秩任务适配矩阵,并基于业务意图指令数据和通用指令数据进行有监督训练,生成用于轨道交通业务指令处理的协同演化大模型。本申请能够提高领域知识注入稳定性、降低小样本激活成本、减少适配器迁移失真。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819360A_ABST
    Figure CN122819360A_ABST
Patent Text Reader

Abstract

The application provides a rail transit field large model collaborative evolution training method and system. The method comprises the following steps: freezing the base model backbone weight, continuously pre-training the high-rank field adaptation matrix based on the dynamic mixed batch of field implicit knowledge data and general logic balance data, and generating a field knowledge adapter; mounting the field knowledge adapter to the bypass of the transformation layer corresponding to the instruction model, freezing the instruction model backbone weight, and performing alignment training on the field knowledge adapter based on the next word prediction loss and the output distribution divergence constraint to generate a fusion model; freezing the field knowledge adapter in the fusion model, mounting a low-rank task adaptation matrix, and performing supervised training based on business intent instruction data and general instruction data to generate a collaborative evolution large model for rail transit business instruction processing. The application can improve the field knowledge injection stability, reduce the small sample activation cost, and reduce the adapter migration distortion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of large-scale rail transit model technology, and in particular to a method and system for collaborative evolution training of large-scale rail transit models. Background Technology

[0002] With the rapid development of artificial intelligence technology in the field of natural language processing, large language models have been gradually applied to scenarios such as knowledge question answering, text generation, and intelligent decision-making assistance. The rail transit field is characterized by rigorous business rules, dense professional terminology, complex equipment numbering, and a vast system of regulations. It involves a wide range of professional knowledge, including signaling equipment, turnout routes, train operation organization, fault handling, and maintenance operations, which places high demands on the model's professional understanding, logical reasoning ability, and instruction compliance capabilities.

[0003] Existing technologies typically employ methods such as full fine-tuning, efficient parameter fine-tuning, or continued pre-training to inject corpora from the rail transit domain into general-purpose large models, thereby improving the model's response capabilities in vertical domains. Some solutions also directly train domain data based on dialogue models that have already undergone instruction fine-tuning, or directly load the adaptation parameters obtained from training on the base model into the instruction model to reduce training costs.

[0004] However, existing technologies still have significant shortcomings: on the one hand, general-purpose word segmenters struggle to accurately handle special equipment numbers, train numbers, and asset identifiers in the rail transit field, easily leading to semantic segmentation errors and creating a professional knowledge illusion; on the other hand, directly injecting domain knowledge into the instruction model can easily disrupt the original dialogue space, resulting in a degradation of instruction compliance capabilities. Furthermore, acquiring high-quality business question-answering samples in rail transit is costly, and existing solutions heavily rely on large-scale labeled data. In addition, there are differences in parameter distribution between the base model and the instruction model, and direct transfer of domain-adapted parameters can easily lead to expression distortion and unnatural logical connections. Therefore, there is an urgent need for a large-scale model training method that can balance domain knowledge injection, dialogue capability preservation, and activation of business intent in small samples. Summary of the Invention

[0005] In view of this, embodiments of this application provide a method and system for collaborative evolution training of large models in the field of rail transit, in order to solve the problems of imbalance between domain knowledge injection and instruction retention, difficulty in activating intent in small samples, and adapter transfer distortion in the existing technology.

[0006] A first aspect of this application provides a method for collaborative evolution training of a large-scale model in the rail transit domain, comprising: acquiring heterogeneous corpora in the rail transit domain; functionally binning the heterogeneous corpora to obtain domain implicit knowledge data, general logical balance data, and business intent instruction data; identifying rail transit professional identifiers in the domain implicit knowledge data based on preset identifier rules; adding rail transit professional identifiers to a word segmenter to add new words; and performing noise cleaning and overlapping window segmentation on the domain implicit knowledge data; freezing the backbone weights of the base model; continuously pre-training the high-rank domain adaptation matrix based on the dynamic mixed batch of domain implicit knowledge data and general logical balance data to generate a domain knowledge adapter; attaching the domain knowledge adapter to the bypass of the transformation layer corresponding to the instruction model; freezing the backbone weights of the instruction model; and performing alignment training on the domain knowledge adapter based on the secondary word prediction loss and output distribution divergence constraints to generate a fusion model; freezing the domain knowledge adapter in the fusion model; attaching the low-rank task adaptation matrix; and performing supervised training based on the business intent instruction data and general instruction data to generate a collaborative evolution large-scale model for rail transit business instruction processing.

[0007] A second aspect of this application provides a large-scale model co-evolution training system for the rail transit domain, comprising: an acquisition module for acquiring heterogeneous corpora in the rail transit domain, performing functional bucketing on the heterogeneous corpora to obtain domain implicit knowledge data, general logical balance data, and business intent instruction data; an identification module for identifying rail transit professional identifiers in the domain implicit knowledge data based on preset identifier rules, adding the rail transit professional identifiers to the word segmenter to add new words, and performing noise cleaning and overlapping window segmentation on the domain implicit knowledge data; and a pre-training module for freezing the backbone weights of the base model, and based on the domain implicit knowledge data and... The dynamic mixing batch of general logical balance data is used to continuously pre-train the high-rank domain adaptation matrix to generate a domain knowledge adapter. The alignment training module is used to attach the domain knowledge adapter to the transformation layer bypass corresponding to the instruction model, freeze the backbone weights of the instruction model, and perform alignment training on the domain knowledge adapter based on the secondary word prediction loss and output distribution divergence constraints to generate a fusion model. The supervised training module is used to freeze the domain knowledge adapter in the fusion model, attach the low-rank task adaptation matrix, and perform supervised training based on business intent instruction data and general instruction data to generate a large-scale collaborative evolution model for rail transit business instruction processing.

[0008] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: By acquiring heterogeneous corpora in the rail transit domain, functional binning is performed on the heterogeneous corpora to obtain domain implicit knowledge data, general logical balance data, and business intent instruction data. Rail transit professional identifiers are identified in the domain implicit knowledge data based on preset identifier rules. These identifiers are added to the word segmenter to add new lexical units, and noise cleaning and overlapping window segmentation are performed on the domain implicit knowledge data. The backbone weights of the base model are frozen, and the high-rank domain adaptation matrix is ​​continuously pre-trained based on the dynamic mixed batches of domain implicit knowledge data and general logical balance data to generate a domain knowledge adapter. The domain knowledge adapter is attached to the bypass of the transformation layer corresponding to the instruction model, the backbone weights of the instruction model are frozen, and the domain knowledge adapter is aligned and trained based on the secondary word prediction loss and output distribution divergence constraints to generate a fusion model. The domain knowledge adapter in the fusion model is frozen, a low-rank task adaptation matrix is ​​attached, and supervised training is performed based on business intent instruction data and general instruction data to generate a large-scale co-evolutionary model for rail transit business instruction processing. This application can improve the stability of domain knowledge injection, reduce the cost of few-shot activation, and reduce adapter transfer distortion. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a flowchart illustrating the collaborative evolution training method for large models in the rail transit field provided in this application embodiment; Figure 2 This is a schematic diagram of the structure of the large-scale collaborative evolution training system for the rail transit field provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0011] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0012] With the development of artificial intelligence technology, generative preprocessing models, represented by Large Language Models (LLM), have made significant progress in the field of natural language processing. In scenarios such as general question answering, coding, and daily office assistance, general-purpose large models have demonstrated powerful semantic understanding and text generation capabilities. However, in vertical industrial sectors such as rail transportation, due to the highly rigorous business logic, dense technical terminology, and the existence of numerous non-public internal regulations and manuals, general-purpose models still face significant challenges in practical applications.

[0013] Currently, existing technologies for model adaptation in vertical domains mainly include full fine-tuning and efficient parameter fine-tuning (such as LoRA). However, in practical applications, these existing technologies have the following significant drawbacks: 1) Knowledge Illusion and Logical Collapse: The rail transit field involves a large number of signal numbers, turnout numbers, and complex interlocking logic (such as special identifiers like signal X123 and turnout No.5). General-purpose model tokenizers often fail to accurately recognize these special symbols, and the lack of such specialized knowledge in general pre-training corpora makes it highly susceptible to generating "false facts" (i.e., illusions) when processing related instructions.

[0014] 2) Degradation of Instruction Compliance (Catastrophic Forgetting): The industry-standard training approach typically involves large-scale domain knowledge injection (CPT) directly onto the instruction fine-tuning model (Chat / Instruction Model). Due to the significant differences in the distribution between domain data and general instruction data, this "hard injection" can easily disrupt the model's original established dialogue logic space. As a result, after learning specialized knowledge, the model loses its basic dialogue interaction capabilities and instruction compliance accuracy, leading to the so-called "becoming stupid after learning new knowledge" phenomenon.

[0015] 3) Scarcity of high-quality labeled data and cost bottlenecks: The cost of acquiring business communication (QA) data in the rail transit field is extremely high, requiring manual annotation by senior dispatchers or technical experts. Existing fine-tuning solutions often rely on large-scale, high-quality instruction pairs. When implemented in vertical fields, the insufficient sample size (e.g., only tens of thousands or even less) often leads to poor model generalization and makes it difficult to develop production-level expert capabilities.

[0016] 4) Heterogeneous architecture adaptation distortion: When attempting to mount a domain knowledge adapter trained on a base model to an aligned chat model, there is an inherent inconsistency in the parameter space distribution between the two. Existing mounting methods often ignore this distribution difference, resulting in a distorted sense of language and awkward logical connections in the merged model when providing professional answers.

[0017] In summary, existing technologies struggle to balance the conflict between "high density of professional knowledge injection" and "maintained instruction compliance capability," and they are heavily reliant on large-scale labeled data. Therefore, there is an urgent need for a large-scale model training solution in the rail transit field that can decouple knowledge and capabilities and achieve expert-level knowledge transfer at extremely low data costs.

[0018] In view of the problems existing in the prior art, this application proposes a three-stage training method for large models in the rail transit field based on LoRA co-evolution. This scheme realizes the smooth evolution from the basic model to the vertical domain expert through three stages: knowledge injection, distribution alignment, and intent activation.

[0019] This solution ensures the model's accurate understanding of rail transit-specific equipment numbers (such as signals and switches) and complex interlocking logic through domain identifier regularization and high-rank LoRA configuration, fundamentally eliminating the "knowledge illusion" of general models when processing rail-specific instructions.

[0020] Based on the ability to reuse cross-model grafting: The core advantage of this solution lies in its innovative architecture adaptation scheme, which "reuses" existing high-performance instruction fine-tuning (Chat / Instruct) model results. Even with limited local computing power and expert annotation resources, the model can acquire the dialogue feel and compliance capabilities of a top-tier industry platform without the need for an expensive full instruction alignment process.

[0021] A deep fusion of industrial-grade rigor and natural conversation: Through Stage 2 (Align CPT) distribution alignment, the model successfully solves the problems of "logical collapse" and "distorted language" that are prone to occur after knowledge injection into vertical domains. Thanks to the introduction of KL divergence constraints, the model can maintain smooth interaction capabilities while invoking rigorous knowledge such as the "Driving Organization Rules," achieving a synergistic evolution of professional depth and instruction compliance.

[0022] Rapid activation and cold start of key business intents: This solution proposes a path to "losslessly graft" the knowledge matrix produced by the Base model onto the Chat model, so that the model only needs a very small number (hundreds) of high-quality business QA to activate the execution capability of specific rail transit scenarios (such as fault emergency handling and dispatch command generation), which significantly reduces the cost threshold for industry expert annotation.

[0023] The technical solution of this application will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0024] Figure 1 This is a flowchart illustrating the collaborative evolution training method for large models in the rail transit field provided in this application embodiment. Figure 1 As shown, the method may specifically include: S101: Acquire heterogeneous corpus in the rail transit field, perform functional binning on the heterogeneous corpus, and obtain domain implicit knowledge data, general logical balance data, and business intent instruction data. S102, Based on preset identifier rules, identify rail transit professional identifiers in domain implicit knowledge data, add rail transit professional identifiers to the word segmenter to add new words, and perform noise cleaning and overlapping window segmentation on the domain implicit knowledge data. S103, freeze the backbone weights of the base model, and continuously pre-train the high-rank domain adaptation matrix based on the dynamic mixed batch of domain implicit knowledge data and general logical balance data to generate a domain knowledge adapter; S104, the domain knowledge adapter is attached to the bypass of the transformation layer corresponding to the instruction model, the backbone weight of the instruction model is frozen, and the domain knowledge adapter is aligned and trained based on the secondary word prediction loss and the output distribution divergence constraint to generate the fusion model; S105, the domain knowledge adapter in the frozen fusion model, is equipped with a low-rank task adaptation matrix and undergoes supervised training based on business intent instruction data and general instruction data to generate a large-scale collaborative evolution model for rail transit business instruction processing.

[0025] In some embodiments, heterogeneous corpora in the rail transit domain are acquired, and the heterogeneous corpora are functionally bucketed to obtain domain tacit knowledge data, general logical balance data, and business intent instruction data, including: Obtain raw corpus containing rail transit-specific texts, general Chinese texts, and business Q&A texts; Generate corpus function labels based on the source of the original corpus, text structure, and training purpose; Based on the functional tags of the corpus, the original corpus is divided into domain implicit knowledge data for continuous pre-training, general logical balance data for language logic constraints, and business intent instruction data for supervised training. Establish the calling relationship between various types of data and corresponding training stages.

[0026] Specifically, the training data management module gathers raw data from the rail transit database, general corpus, and business Q&A database, and generates source identifiers. The rail transit database can include professional texts such as line design specifications, train operation rules, signal equipment maintenance manuals, operation standards, and historical fault cases; the general corpus includes encyclopedic knowledge, news texts, general teaching materials, and logical corpora; the business Q&A database can include Q&A samples compiled by experts for scenarios such as emergency fault handling, regulation and clause retrieval, dispatch order generation, and equipment maintenance suggestions. After receiving the raw data, the system generates corpus function tags according to file source, document level, paragraph structure, field format, annotation status, and training stage.

[0027] When generating functional tags for the corpus, continuous regulatory texts, maintenance instructions, and fault case texts can be tagged as domain knowledge tags; general Chinese texts with standardized expressions and clear logical relationships can be tagged as logical balance tags; and structured question-and-answer texts containing task instructions, business inputs, and standard outputs can be tagged as business intent tags. For documents containing regulatory text, appendix explanations, and question-and-answer examples, the system splits them according to paragraph boundaries and field characteristics, assigning different functional tags to each to avoid misclassifying the entire document into a single data category. For example, continuous clauses in a section's regulations regarding the handling of live vehicle faults are entered into the domain implicit knowledge data for continuous pre-training to learn the sequential relationship between parking protection, fault reporting, and subsequent handling; question-and-answer samples on how to handle live vehicle faults, compiled by dispatchers, are entered into the business intent instruction data for supervised training to activate question-and-answer response capabilities.

[0028] After generating functional labels, the system performs binning based on the corpus functional labels. Corpus with domain knowledge labels is divided into domain implicit knowledge data, used for continuous pre-training in the first stage to establish semantic associations between rail transit terminology, equipment numbers, operating procedures, and interlocking logic; corpus with logic balance labels is divided into general logic balance data, used for mixed sampling with domain implicit knowledge data to constrain the language expression structure and general reasoning chain during training; corpus with business intent labels is divided into business intent instruction data, used for supervised training in the third stage to trigger the model's response patterns in tasks such as regulation retrieval, fault handling, and schedule generation.

[0029] Furthermore, the system establishes a call relationship between data bucketing and the training phase, binding domain tacit knowledge data and general logical balance data to the continuous pre-training and alignment training phases, and binding business intent instruction data to the business instruction activation phase, while configuring sampling weights, batch mixing rules, and quality status. After reading the call relationship, the training scheduler automatically determines the data buckets, sampling ratios, and filtering conditions to be called in the current phase. Heterogeneous corpora are transformed into data assets with clear training purposes, forming a manageable data collaboration relationship between domain knowledge injection, general logic preservation, and business intent activation, improving the stability of training data organization, and reducing the reliance on manual repetitive annotation for small-sample business instruction training.

[0030] In some embodiments, identifying rail transit specialty identifiers in domain tacit knowledge data based on preset identifier rules includes: Based on the coding structure and context boundaries of rail transit business objects, construct identifier matching rules corresponding to professional number texts; The domain tacit knowledge data is scanned segment by segment using identifier matching rules to extract candidate identifiers that conform to the complete coding structure; Based on the semantic dependency relationship and frequency of occurrence of candidate identifiers in adjacent business statements, the candidate identifiers are deduplicated, normalized, and filtered for validity to generate a set of rail transit professional identifiers.

[0031] Specifically, the identifier recognition module establishes identifier matching rules for specialized numbered texts with rail transit business implications within the domain's tacit knowledge data. These rules are generated based on the coding structure of the business object, the boundaries between preceding and following statements, and additional terms related to the number. Business objects can include signaling equipment, turnout routes, train numbers, rolling stock assets, and track sections. The system first segments the train operation rules, signaling equipment maintenance manuals, work standards, and historical fault cases, recording the document source, chapter position, and business topic for each segment. Then, it performs rule scanning within each segment to avoid misidentification due to cross-segment splicing.

[0032] When constructing identifier matching rules, the system configures a unified coding and recognition template based on common numbering formats in rail transit data. For signal equipment numbers, the recognition module combines equipment name, location description, and alphanumeric combinations to identify text such as "entry signal X123" and "departure signal S2"; for turnout numbers, it combines contextual terms such as "turnout," "route," "position," and "reverse position" to identify text such as "21# turnout" and "No.5 turnout"; for train number and vehicle asset number, it combines contextual terms such as "train," "train number," "vehicle," and "line" to identify text such as "G1234 train" and "Line 2, car 12A34." During the recognition process, the system not only determines whether the text meets the coding structure but also whether there is a corresponding business object description before and after the candidate number, thereby excluding non-professional numbers such as page numbers, clause numbers, and ordinary serial numbers.

[0033] During the segment-by-segment scanning process, the system takes the cleaned domain tacit knowledge data as input, extracts candidate identifiers according to paragraph order and sentence boundaries, and records the occurrence position, adjacent business statements, belonging business objects, and context fragments for each candidate identifier. For example, in a fault case, when there is a description of checking the positioning status of turnout #21 after X123 failed to open, the system extracts X123 and #21 as candidate identifiers and binds them to contextual relationships such as signal opening failure and turnout positioning check; when a maintenance manual describes the maintenance record of train G1234 passing through car 12A34 on Line 2, the system uses the train number and vehicle asset number as candidate identifiers and marks them with business categories.

[0034] During the candidate identifier screening phase, the system normalizes candidate identifiers based on their semantic dependencies in adjacent business statements, frequency of occurrence, and document source credibility. For multiple spellings of the same object due to differences in capitalization, full-width / half-width characters, spaces, or connectors, the system merges them into a unified expression. Candidate numbers appearing only once and lacking contextual support from business objects are downgraded or eliminated. Candidate numbers that repeatedly appear in regulations, fault cases, and maintenance records and have stable contextual relationships are retained as rail transit professional identifiers. The resulting set of rail transit professional identifiers can be used by the word segmenter to add new lexical tables, allowing signal numbers, turnout numbers, train numbers, and vehicle asset numbers to participate in subsequent training as complete semantic units, improving the stability of professional number recognition and reducing domain semantic bias caused by word segmentation fragmentation.

[0035] In some embodiments, rail transit professional identifiers are added to the word segmenter to create new words, and noise cleaning and overlapping window segmentation are performed on the domain implicit knowledge data, including: Generate a new word segmentation table based on the set of rail transit professional identifiers, and configure the overall word segmentation attributes of each professional identifier in the new word segmentation table; Based on the overall word segmentation attributes, domain implicit knowledge data is segmented, so that professional identifiers can participate in training as complete word units; The domain implicit knowledge data before and after word segmentation is cleaned up for layout noise, invalid text is removed, and character anomalies are repaired to generate clean text. The cleaned text is slid-segmented according to the preset window length and overlap ratio to generate domain knowledge training segments with context connection markers.

[0036] Specifically, after receiving the set of rail transit professional identifiers, the word segmentation module generates a new word segmentation table for the word segmenter according to the identifier category, standard spelling, and frequency of occurrence, and configures an overall word segmentation attribute for each professional identifier. The overall word segmentation attribute is used to instruct the word segmenter to treat the corresponding text as an indivisible complete word segment when processing domain implicit knowledge data, and not to split it according to general sub-word rules. For example, the entry signal X123, the exit signal S2, turnout #21, turnout No.5, train G1234, and car 12A34 of Line 2 can all be added to the new word segmentation table as complete words, and their association tags with business objects such as signal equipment, turnout routes, train operation, and vehicle assets are retained.

[0037] During word segmentation, the system first loads the newly added lexicon table and the original word segmenter lexicon table, establishing a priority matching relationship. When a professional identifier appears in the text before cleaning, the system prioritizes overall matching according to the newly added lexicon table. For a traffic organization rule containing a paragraph about checking the positioning status of turnout #21 after signal X123 fails to open, the system outputs X123 and #21 as complete lexicons, and retains adjacent context words such as "failed to open," "positioning status," and "route check" in the same training segment to prevent professional identifiers from being separated from business semantics after being broken down.

[0038] During the noise cleanup process, the system performs layout noise cleanup, invalid text removal, and character anomaly repair on domain tacit knowledge data. For text converted from regulations manuals or maintenance manuals, the system identifies and removes headers and footers, page numbers, table of contents indexes, duplicate headings, line breaks, and garbled characters; merges continuous business statements that have been broken by layout issues; and uniformly processes full-width and half-width characters, abnormal spaces, and number connectors. For long texts containing work processes and fault cases, the system retains chapter levels, clause numbers, and business topic identifiers, ensuring that subsequent segmentation can still trace the text's origin.

[0039] After generating the cleaned text, the system performs sliding segmentation according to a preset window length, which can be configured to 1024 or 2048 words, and sets an overlap area at a ratio of 10% to 20%. For fault emergency handling procedures, interlocking condition descriptions, and continuous scheduling operation rules, the system copies the key context from the end of the previous window to the beginning of the next window, generates a context connection identifier, and records the segment number, the relationship between preceding and following segments, and the source chapter. For example, in handling vehicle body electrical faults, parking protection, fault reporting, and subsequent inspection are continuous business chains, and the overlapping area is used to maintain the step-by-step relationship during segmentation.

[0040] Through the above processing, the domain knowledge training fragments can simultaneously preserve the integrity of professional identifiers, the cleanliness of regulatory texts, and the continuity of long text contexts, thereby improving the semantic modeling stability of the knowledge injection stage in the rail transit domain and reducing professional knowledge bias caused by word segmentation fragmentation and text noise.

[0041] In some embodiments, the backbone weights of the base model are frozen, and the high-rank domain adaptation matrix is ​​continuously pre-trained based on a dynamic mixed batch of domain implicit knowledge data and general logical balance data to generate a domain knowledge adapter, including: Keep the backbone weights of the base model frozen, and configure a high-rank neighborhood adaptation matrix in the transformation layer of the base model; Training segments are extracted from domain implicit knowledge data and general logical balance data according to preset sampling weights, and then mixed to generate dynamic mixed batches. Sub-word prediction training is performed using dynamic mixed batches as input, and only the high-rank domain adaptation matrix is ​​updated; After the training reaches the preset stopping condition, the updated high-rank domain adaptation matrix is ​​saved to obtain the domain knowledge adapter.

[0042] Specifically, the training scheduler uses the untuned base model as the knowledge injection target. While keeping the core parameters of the model frozen, it configures a high-rank domain adaptation matrix as a bypass in the attention mapping layer and the feedforward mapping layer. The high-rank domain adaptation matrix can use a rank value of at least 64, and a corresponding scaling factor is configured according to the rank value to ensure that the newly added matrix has the parameter capacity to accommodate the distribution of implicit knowledge in rail transit. In this embodiment, only the training gradient is applied to the high-rank domain adaptation matrix; the original lexical, syntactic, and general semantic representations of the base model are not updated.

[0043] During the training data organization phase, the training scheduler reads the call relationship between data buckets and the training phase. It extracts training fragments such as train operation rules, signal equipment maintenance manuals, operation standards, and historical fault cases from domain tacit knowledge data, and encyclopedic texts, news texts, general teaching materials, and logical corpora from general logical balance data. Sampling weights are configured according to the knowledge injection rounds; for example, domain tacit knowledge data and general logical balance data are mixed in a 4:1 ratio, or the proportion of domain tacit knowledge data is appropriately reduced in the later stages of training to form dynamic mixed batches. Each dynamic mixed batch includes both rail transit professional terminology, equipment numbers, interlocking conditions, and handling procedures, as well as general Chinese expressions and basic reasoning text.

[0044] During continuous pre-training, the system dynamically mixes batches of input to the base model, performs word prediction training, and updates the high-rank domain adaptation matrix inversely based on the cross-entropy loss between predicted and actual words. For training segments containing complete words such as station signal X123, turnout #21, train G1234, and car 12A34 on Line 2, the model learns the conditional probability relationship between professional numbers and business context through the high-rank domain adaptation matrix. For example, in the industry regulation text, the contextual relationship between the failure of signal X123 to open and route inspection, and the positioning status of turnout #21, is modeled as continuous semantics; in the case of a live fault in the car body, the sequential relationship between parking protection, fault reporting, and subsequent inspection is written into the parameter response.

[0045] Furthermore, after each training round, the system statistically analyzes the changes in perplexity on the implicit knowledge data in the statistical domain, the changes in language loss on the general logical balance data, and the gradient magnitude of the high-rank domain adaptation matrix. When the preset training rounds, loss convergence conditions, or stable prediction conditions of the specialized corpus are reached, continuous pre-training stops, and the updated high-rank domain adaptation matrix is ​​saved, forming a domain knowledge adapter. This domain knowledge adapter serves as a parameterized storage result of knowledge in the rail transit domain, which can be used in subsequent cross-model grafting and distribution alignment stages.

[0046] Through the above embodiments, rail transit regulations, equipment numbers, interlocking logic, and fault handling procedures can be injected into the model in the form of an incremental matrix. The general logic balancing data can constrain the expression drift during the training process, enabling the domain knowledge adapter to have both professional knowledge carrying capacity and general language structure preservation capacity, thereby improving the stability of domain knowledge injection and reducing the dependence of subsequent business instruction activation on large-scale expert-annotated samples.

[0047] In some embodiments, the domain knowledge adapter is attached to the bypass of the transformation layer corresponding to the instruction model, the backbone weights of the instruction model are frozen, and the domain knowledge adapter is aligned and trained based on the secondary word prediction loss and output distribution divergence constraints to generate a fusion model, including: Based on the hierarchical position and matrix dimension of the domain knowledge adapter in the base model, determine the corresponding transformation layer bypass in the instruction model, and mount the domain knowledge adapter to the transformation layer bypass layer by layer. The backbone weights of the instruction model are frozen, and the domain knowledge adapter is updated incrementally using a learning rate lower than that used during continuous pre-training. The professional continuation writing corpus is input into the mounted instruction model. The word prediction loss and the divergence loss between the output probability distributions before and after mounting are calculated respectively. The domain knowledge adapter is updated based on the word prediction loss and the divergence loss to generate a fusion model.

[0048] Specifically, the alignment training module reads the hierarchical index, target mapping sublayer, matrix rank, input dimension, and output dimension of the domain knowledge adapter in the base model, and establishes a mounting mapping based on the hierarchical structure of the corresponding transformation layer in the instruction model. For the high-rank domain adaptation matrix formed in the attention mapping layer and feedforward mapping layer of the base model, the system establishes a trainable mounting interface at the bypass position of the same semantic level in the instruction model and compensates for the matrix scaling factor, so that the rail transit knowledge response trained by the base model can enter the dialogue generation link of the instruction model. After mounting, the backbone weights of the instruction model remain frozen, allowing only the domain knowledge adapter to participate in minor updates, avoiding damage to the original dialogue space of the instruction model.

[0049] During the alignment training phase, the system selects a small-scale, high-quality professional continuation writing corpus as input. This corpus can be derived from bicycle organization rules, signal equipment maintenance manuals, fault case summaries, and operational standard fragments. For example, the professional continuation writing corpus includes a rule fragment on checking the positioning status of turnout #21 after signal X123 fails to open, a description of handling abnormalities in the route management of train G1234, and fault review content from the maintenance record of car 12A34 on Line 2. The system inputs the same professional continuation writing corpus into both the original instruction model without a domain knowledge adapter and the instruction model with a domain knowledge adapter, obtaining the output probability distribution of the original instruction model and the output probability distribution after the adapter is attached, and recording the differences between the two distributions on professional and ordinary lexical units.

[0050] During training, the system first calculates the sub-word prediction loss based on the prediction results of subsequent words by the model after mounting, ensuring that the domain knowledge adapter maintains its ability to call upon rail transit terminology, equipment numbers, and handling sequences. Then, it calculates the divergence loss based on the output probability distribution of the original instruction model and the output probability distribution after mounting, to constrain the expression distribution of the model after mounting to approximate the original dialogue distribution of the instruction model. The system synthesizes the sub-word prediction loss and divergence loss into an alignment loss using dynamic weights, and performs 1 to 3 rounds of minor updates to the domain knowledge adapter using a learning rate lower than that used in the continuous pre-training phase. For corpora related to handling vehicle electrical faults, the system learns the professional order of stopping protection, fault reporting, and subsequent inspection while suppressing abrupt changes in response sentence structure and logical breaks.

[0051] When the alignment loss reaches the convergence condition, and the lexical predictions on the professional continuation corpus are stable and the output distribution deviation is within the threshold range, the system stops training and saves the aligned domain knowledge adapter and its mounting relationship with the instruction model, generating a fusion model. Through the above embodiments, the rail transit domain knowledge formed in the basic large model stage can be smoothly transferred to the instruction model, reducing the linguistic distortion and probability jitter caused by differences in cross-model parameter spaces, and improving the stability of professional knowledge invocation and the consistency of dialogue expression.

[0052] In some embodiments, based on the hierarchical position and matrix dimension of the domain knowledge adapter in the base model, the corresponding transformation layer bypass in the instruction model is determined, and the domain knowledge adapter is mounted to the transformation layer bypass layer by layer, including: Read the hierarchical index, target mapping sub-layer, matrix rank, and input / output dimensions of the domain knowledge adapter to generate adapter mounting metadata; Based on the adapter mounting metadata, search for transformation layer bypasses with the same hierarchical semantics and compatible matrix dimensions in the instruction model; While keeping the domain knowledge adapter matrix structure unchanged, the mounting weights are scaled and compensated, and the domain knowledge adapter is mapped to the corresponding transformation layer bypass.

[0053] Specifically, before entering alignment training, the model grafting middleware first reads the parameter file and training configuration file of the domain knowledge adapter, parses the hierarchical index, target mapping sublayer, matrix rank, input dimension, output dimension, scaling factor, and parameter naming path of the domain knowledge adapter in the base model, and generates adapter mounting metadata. This adapter mounting metadata describes the correspondence between the domain knowledge adapter and the transformation layer of the base model. For example, a certain high-rank domain adaptation matrix comes from the 12th layer attention mapping bypass of the base model, with a matrix rank of 64, and the input and output dimensions match the linear transformation dimensions of the query mapping or value mapping of that layer, respectively.

[0054] After generating the adapter mounting metadata, the model grafting middleware reads the network structure description of the instruction model and searches for compatible transformation layer bypasses according to the hierarchical sequence, sub-layer function, linear mapping direction, and tensor dimension. For scenarios consisting of the same model family base model and instruction model, the system prioritizes mapping according to the same hierarchical semantics. For scenarios with different layer names but compatible structural dimensions, the system establishes equivalent mounting relationships based on attention mapping, feedforward mapping, normalized position, and residual connection relationships to avoid mounting misalignment caused by relying solely on parameter names. Taking the high-rank matrix in the domain knowledge adapter that carries professional semantics such as X123 signal, 21# turnout, and G1234 train as an example, the system maps it to the transformation layer bypass in the instruction model that performs the same semantic transformation function, enabling rail transit professional knowledge to enter the generation path of the instruction model.

[0055] When performing layer-by-layer mounting, the system maintains the matrix rank, matrix direction, and input / output dimensions of the domain knowledge adapter unchanged, without reinitializing the already formed rail transit knowledge parameters. For cases where there are differences in activation amplitude between the basic large model and the instruction model, the model grafting middleware calculates a scaling compensation value based on the scaling factor during the first-stage training, the current instruction model layer output norm, and the initial response strength of the adapter, and applies this scaling compensation value to the mounting weights or the adapter output branch. This approach reduces probability distribution jitter caused when knowledge responses formed in the parameter space of the basic large model are directly imported into the instruction model.

[0056] Furthermore, after completing the mounting of each layer, the system performs dimension verification and response verification. Dimension verification is used to determine whether the input tensor, output tensor, and bypass summation positions are consistent; response verification is used to input specialized segments, including handling of vehicle body electrical faults, signal opening failures, and turnout positioning checks, into the mounted instruction model to detect whether the adapter output amplitude is within the preset range. If a layer exhibits dimension incompatibility or abnormal response, the system marks that layer as needing adjustment and recalculates the scaling compensation or switches to an equivalent sublayer bypass.

[0057] Through the above embodiments, the domain knowledge adapter can smoothly migrate from the basic large model to the instruction model while keeping the matrix structure unchanged. It reduces cross-model grafting distortion through hierarchical semantic matching, dimensional compatibility verification and scaling compensation, so that the professional knowledge response of rail transit is stably connected with the original dialogue generation path of the instruction model, thereby improving the convergence stability and professional expression consistency of subsequent distributed alignment training.

[0058] In some embodiments, the domain knowledge adapter in the fusion model is frozen, a low-rank task adaptation matrix is ​​attached, and supervised training is performed based on business intent instruction data and general instruction data to generate a large-scale co-evolutionary model for rail transit business instruction processing, including: Keep the domain knowledge adapter in the fusion model frozen, and attach a low-rank task adaptation matrix to the target transformation layer of the fusion model according to the business scenario corresponding to the business intent instruction data. The business intent instruction data and general instruction data are mixed in a preset ratio to generate instruction training batches; Supervised training is performed in batches based on instructions, and only the low-rank task adaptation matrix is ​​updated. The trained low-rank task adaptation matrix and the domain knowledge adapter are co-loaded to generate a large co-evolutionary model.

[0059] Specifically, the business intent activation module uses the fusion model that has completed distribution alignment as the training object, and keeps the aligned domain knowledge adapter in the fusion model in a frozen state. This domain knowledge adapter is used to continue to carry the rail transit regulations knowledge, signal equipment number semantics, turnout route relationships, and fault handling order written in the previous stage, and is not updated again in the business intent activation stage. According to the scenario type corresponding to the business intent instruction data, the system bypasses the target transformation layer of the fusion model with a low-rank task adaptation matrix. The low-rank task adaptation matrix is ​​used to learn the mapping relationship between business instructions and output formats, avoiding the disturbance of the already formed domain knowledge parameters through large-scale parameter updates.

[0060] When constructing instruction training batches, the system reads task instructions, business inputs, and standard outputs from the business intent instruction data and organizes the samples according to scenario tags. Business intent instruction data can include samples such as emergency fault handling Q&A, train operation rule retrieval, dispatch command generation, and signal equipment maintenance suggestions. For example, for a business instruction on how to handle a red light band on a signal, the business input can include the status of signal X123, relevant routes, the location information of turnout #21, and current train operation conditions; the standard output is organized according to rule retrieval, status verification, handling steps, and reporting requirements. For a business instruction on handling a live fault in a train body, the standard output is organized in the order of stopping protection, fault reporting, personnel protection, and subsequent inspection. The system also extracts daily Q&A, text rewriting, translation, and basic logical reasoning samples from general instruction data and mixes them into the business intent instruction data according to a preset ratio, forming instruction training batches that combine professional tasks and general interactions.

[0061] During supervised training, the system inputs batches of instruction training data into the fusion model, calculates the supervised loss based on the standard output, and updates only the low-rank task adaptation matrix. During training, the domain knowledge adapter remains frozen and participates in forward inference, providing parameterized knowledge support for rail transit terminology, equipment numbers, and regulatory logic. The low-rank task adaptation matrix learns how to invoke this knowledge support and form a response structure that conforms to the business scenario based on the supervised loss. For the sample of anomaly handling for train route G1234, the low-rank task adaptation matrix learns to convert train number, route, signaling equipment, and turnout status into scheduling handling expressions; for the maintenance suggestion sample for car 12A34 on Line 2, the low-rank task adaptation matrix learns to convert vehicle asset number, fault records, and maintenance rules into maintenance suggestion outputs.

[0062] After training reaches a preset number of rounds or the validation set loss converges, the system saves the low-rank task adaptation matrix. During inference deployment, the low-rank task adaptation matrix is ​​co-loaded with the domain knowledge adapter to form a large-scale co-evolutionary model for rail transit business instruction processing. Through the above embodiments, rail transit domain knowledge storage and business intent triggering are respectively carried on different adaptation matrices, which can reduce the training cost of small-sample instructions, reduce interference with domain knowledge parameters, and improve the stability of instruction compliance in scenarios such as professional question answering, regulation retrieval, and scheduling generation.

[0063] In some embodiments, after generating the large co-evolutionary model, the method further includes: Based on domain tacit knowledge data, general logical balance data, and business intent instruction data, an evaluation sample covering professional knowledge retention, general instruction compliance, professional identifier stability, and business rule consistency is constructed. The evaluation samples are input into the co-evolutionary large model, the model output results are obtained, and the model output results are matched with the standard knowledge base and preset output constraints to generate model capability evaluation results. Based on the model capability evaluation results, determine the knowledge injection bias, alignment bias, or intent activation bias, and adjust the sampling weights of the dynamic mixed batch, the divergence constraint weights of the alignment training, or the instruction data ratio of the supervised training based on the bias.

[0064] Specifically, after the collaborative evolution model is generated, the evaluation and monitoring module extracts samples from domain tacit knowledge data, general logical balance data, and business intent instruction data to construct a hierarchical evaluation set. Professional knowledge retention samples come from bicycle organization rules, signal equipment maintenance manuals, operation standards, and historical fault cases to examine the model's retention of rail transit terminology, equipment relationships, and handling procedures. General instruction compliance samples come from general Chinese question-and-answer, formatted output, text rewriting, and basic logical reasoning tasks to examine whether the model retains its original interactive capabilities after completing domain-specific training. Professional identifier stability samples include complete words such as X123 for entry signals, S2 for exit signals, turnout #21, turnout No.5, train G1234, and car 12A34 of Line 2 to detect whether the numbers have been split, replaced, or formatted incorrectly. Business rule consistency samples come from fault emergency handling and dispatch command generation scenarios to verify the matching relationship between the output content and the standard knowledge base.

[0065] During the evaluation process, the system inputs the aforementioned evaluation samples into the co-evolutionary large model, obtains the model output results, and performs structured parsing on the output results. For professional knowledge preservation samples, the system calculates the semantic matching degree and key term coverage between the output text and standard regulation fragments; for general instruction compliance samples, the system checks whether the output format, number of steps, limiting conditions, and multi-round context responses meet the preset output constraints; for professional identifier stability samples, the system compares whether the number characters, connectors, line information, and object types in the output are consistent with the input; for business regulation consistency samples, the system uses the handling process in the standard knowledge base as a reference to check whether the sequence of stopping protection, fault reporting, personnel protection, and subsequent inspections in a vehicle body electrical fault is correctly expressed, or checks whether signal status verification, turnout positioning check, and route condition confirmation constitute a complete chain in a signal opening failure scenario.

[0066] After generating model capability evaluation results, the system identifies training biases based on different evaluation dimensions. If the perplexity increases or the coverage of key terms decreases in the professional knowledge retention samples, a knowledge injection bias is identified, and the sampling weight of domain tacit knowledge data in the dynamic mixing batch is increased. If formatting errors, abrupt changes in tone, or logical breaks occur in the general instruction compliance samples, an alignment bias is identified, and the divergence constraint weight in alignment training is increased or the micro-update learning rate is reduced. If misjudgments of handling scenarios, missing output structures, or incomplete intent triggering occur in the business regulation consistency samples, an intent activation bias is identified, and the mixing ratio of business intent instruction data and general instruction data is adjusted. Through the above evaluation and feedback adjustments, the training process can form a closed-loop monitoring mechanism for rail transit production scenarios, improving the stability of knowledge retention, instruction compliance, and regulation output consistency of the co-evolutionary large model.

[0067] In some embodiments, the parameter configurations of the high-rank neighborhood adaptation matrix and the low-rank task adaptation matrix include: Based on the density of professional identifiers, the length of regulatory texts, and the complexity of business logic associations in the domain tacit knowledge data, the matrix rank of the high-rank domain adaptation matrix is ​​configured to be no less than 64, and a scaling factor corresponding to the matrix rank is configured. Based on the sample size of the business intent instruction data and the number of business scenarios, the matrix rank of the low-rank task adaptation matrix is ​​configured to be between 8 and 32, and lower than the matrix rank of the high-rank domain adaptation matrix. During the continuous pre-training phase, only the high-rank domain adaptation matrix is ​​updated. During the supervised training phase, the domain knowledge adapter is frozen and only the low-rank task adaptation matrix is ​​updated.

[0068] Specifically, the parameter configuration module configures the RANK parameters for Lora1, the alignment phase adapter state, and Lora3 before the three-stage training. Lora1 is used for knowledge injection in the first stage, performing continuous pre-training on domain implicit knowledge data. Its input corpus comes from professional texts such as bicycle organization rules, signal equipment maintenance manuals, operation standards, and historical fault cases. The texts contain professional identifiers such as X123 signal, 21# turnout, G1234 train, and Line 2 12A34 car, as well as continuous business logic such as route processing, signal opening, turnout positioning, and live car handling. Since the corpus in the rail transit field needs to carry a large number of professional numbering relationships, regulatory clause relationships, and interlocking condition relationships, the parameter configuration module configures Lora1 as a high RANK structure, such as RANK 64, 96, or 128, and configures the corresponding scaling factor according to the RANK configuration, so that Lora1 has a higher parameter capacity than the conventional LORA instruction fine-tuning configuration.

[0069] In the first phase of training, the core parameters of the basic model remained frozen, with only Lora1 being updated. Training batches were dynamically formed by mixing domain tacit knowledge data and general logically balanced data, for example, extracting rail transit-specific segments and general Chinese segments at a 4:1 ratio. For the regulation segment on checking the positioning status of turnout #21 after the X123 signal failed to open, Lora1 learned the correlation between signal equipment, turnout status, and route conditions; for the case of a live fault in the car body, Lora1 learned the order of handling between stopping protection, fault reporting, personnel protection, and subsequent inspection. This phase did not rely on large-scale business question-and-answer annotation, but instead completed the accumulation of rail transit tacit knowledge with a high-ranking Lora1 and a relatively small industry corpus.

[0070] In the second phase, the system grafts Lora1 across models to the instruction model and aligns it using low learning rates and output distribution divergence constraints, adapting the rail transit domain knowledge accumulated in Lora1 to the original dialogue distribution of the instruction model. After alignment, the system enters the third phase, where the business intent activation module keeps Lora1 frozen and attaches Lora3 to the target transformation layer of the fusion model. Lora3 is used for SFT activation, with its RANK configured as a medium-low RANK, such as 8, 16, or 32, lower than the RANK of Lora1. The business intent instruction data can contain only 100 to 1000 high-quality question-and-answer samples, mixed with 10% to 20% of general instruction samples. For samples such as how to handle red light bands on signal lights, how to deal with electrified car bodies, and how to generate dispatch commands for abnormal routes of train G1234, Lora3 mainly learns instruction triggering, output format, and business scenario response methods, without undertaking the function of rewriting domain knowledge.

[0071] In the above embodiments, the high-rank Lora1 is used for knowledge accumulation, and the medium-to-low-rank Lora3 is used for SFT activation. They respectively handle domain knowledge storage and business intent triggering functions. This parameter configuration can expand the model's rail transit domain knowledge with far less data than the industry's standard training data, and improve the stability of professional knowledge retrieval and the ability to activate small-sample business instructions without updating the core parameters of the instruction model or compromising the original Chat capabilities.

[0072] The specific content of the embodiments of this application has been described in detail above. The following will illustrate the specific implementation process of the three-stage training method for a large model in the rail transit field based on LoRA co-evolution, using examples from real-world scenarios. Specifically, it may include the following: 1. Heterogeneous data preprocessing and feature binning strategy 1.1 Heterogeneous Data Bucketing Process This approach employs a dynamic weighted sampling data binning process to ensure that the model maintains its general expressive power while learning domain knowledge.

[0073] 1.2 Definition of Heterogeneous Data Bucketing To balance the need for specialized data in the rail transit sector with general logical requirements, this solution divides the original corpus into three logical buckets with different functional attributes: 1) Bucket A (Domain Implicit Knowledge Pool): Data composition: It covers rail transit industry standards (such as the "Railway Line Design Specification"), train operation rules (industry regulations) of various sections, equipment maintenance manuals, operation standards and compilation of historical failure cases.

[0074] Technical intent: As the core input of the first stage of LoRA CPT, it aims to enable the model to establish the conditional probability distribution between domain terms (such as interlocking) through large-scale plain text injection.

[0075] 2) Bucket B (General Logic Balanced Pool): Data composition: extracted from high-quality general Chinese corpora (such as encyclopedias, news, and general textbooks).

[0076] Technical intent: To act as an "anchor" in incremental pre-training to prevent the model from suffering catastrophic forgetting due to overfitting to industry terminology, and to ensure that the model maintains basic syntactic structure and logical reasoning ability.

[0077] 3) C-bucket (Business Intent Instruction Set): Data structure: Vertical domain QA pairs organized using the {"instruction", "input", "output"} structure.

[0078] Technical intent: Used for the third phase of LoRA SFT to activate the model's intent-triggered capabilities in specific business scenarios (such as fault emergency handling procedures and regulatory clause retrieval).

[0079] 1.3 Domain Identifier Regularization Protection Strategy Because the rail transit sector contains a large number of specific codes (such as signals, switches, and vehicle numbers), general-purpose tokenizers often break them down into meaningless sub-words, leading to semantic breaks. This solution uses preset regular expressions to scan the text and identify the following key identifiers: 1) Signal equipment: ^[X|S|D|L][AZ]?\d+$ (e.g., entrance signal X123, exit signal S2).

[0080] 2) Routes and turnouts: ^\d{1,3}#$ or ^No\.\d+$ (e.g., 21# turnout).

[0081] 3) Train number and formation type: ^[G|D|C|K]\d{1,4}$ (e.g., train G1234).

[0082] 4) Production asset number: Line 1 [0-9A-Z] [5] cars (e.g., Line 2, car 12A34).

[0083] The identified high-frequency specific identifiers are added to the tokenizer's `added_tokens` list, forcing them to be treated as a single token and not participating in word splitting.

[0084] 1.4 Data Cleaning and Long Text Processing Workflow 1) Noise Removal from Basic Corpus Features: For PDF-exported industry manuals, implement automated header and footer noise removal, table of contents index removal, and garbled character repair to prevent noise in unstructured text from interfering with the model's knowledge absorption.

[0085] 2) Sliding window segmentation of CPT corpus For lengthy rules and regulations in bucket A, use a sliding window of fixed length (such as 1024 or 2048 tokens) to divide them.

[0086] Overlap mechanism: settings Text overlap ensures that semantic logic across windows (such as consecutive processing steps) does not lose context due to physical splitting.

[0087] 1.5 Dynamic Mixed Sampling in Training Batches In implementation, this scheme does not use single-bucket data in isolation, but achieves co-evolution through feature-matching sampling: 1) Core data allocation and recommended thresholds: Bucket A (Domain Knowledge): To ensure the coverage density of domain knowledge, it is recommended to select a professional corpus with a size of at least [missing information]. (In practical applications, the larger the corpus, the more accurate the domain knowledge representation, with no upper limit).

[0088] Bucket B (General Logic): Serving as the "skeleton support," its initial sampling ratio with Bucket A. Set as (That is, A:B = 4:1).

[0089] C-bucket (Business Intent): The minimum order of magnitude for activating an intent is... A high-quality QA is achieved. As the complexity of business scenarios increases, the granularity of instruction compliance can be further improved by expanding the sample size of the C bucket.

[0090] 2) Sampling formula: For each training batch Depend on , Two buckets of data by weight ratio Mixed together: 3) Data allocation for each stage: Phase 1 (Injection Period): Set a relatively high proportion To enhance the infusion of knowledge in the field of rail transit.

[0091] Phase Two (Alignment Period): Reduction to (i.e., A:B=1:1), and a low learning rate is matched to ensure that while the model absorbs professional knowledge, its probability distribution is aligned with the instruction base to prevent distortion of language sense.

[0092] Phase Three (Activation Period): Adopting Bucket the full data and mix in General instruction data This enables a cold start for business capabilities.

[0093] 2. Implicit Knowledge Parameterized Storage (LoRA CPT) 2.1 Selection of Operating Object and Base Base Model Selection: The object of operation in this stage is the base model that has not been fine-tuned by instructions.

[0094] Technical logic: The reason for choosing the base model instead of the instruction model is to achieve "full coverage" of domain knowledge in the purest semantic space, and to avoid interference from the existing RLHF preferences of the instruction model (Chat / Instruct) on knowledge injection.

[0095] 2.2 High-Rank LoRA Parameter Configuration To address the high-density nature of rail transit knowledge, this solution performs specific optimizations on the hyperparameters of LoRA: Rank Setting: When implementing Domain Knowledge Injection (CPT), a high-rank adapter is set to accommodate the massive tacit knowledge distribution unique to the orbital discipline, thus supporting the high-density knowledge distribution. This differs from traditional instruction fine-tuning (SFT) tasks (which typically only require...). or Compared to traditional methods, incremental pre-training (CPT) requires building the underlying semantic relationships and complex interlocking logic of the rail transit system from scratch (e.g., approach control, signal degradation modes). Low-rank matrices struggle to support such a high density of implicit knowledge distribution. Therefore, this approach significantly improves the Rank, increasing... The parameter capacity ensures knowledge storage density, thereby enabling in-depth representation of industry knowledge.

[0096] Scaling factor (Alpha): Setting This enhances the domain adapter's ability to correct the original pre-trained weights.

[0097] 2.3 Training Strategy Based on Hybrid Bucketing This stage utilizes the A bucket (specialized knowledge) and B bucket (general logic) partitioned during data preprocessing for joint training: 1) Knowledge injection mechanism: The Next Token Prediction (NTP) task is adopted to enable the model to learn the underlying semantics of rail transit in the A bucket corpus.

[0098] 2) Logistic drift suppression: By introducing [a feature] in each training batch The data from bucket B is used to perform gradient backpropagation on a general corpus to constrain the LoRA weights so that they do not deviate from natural language expression habits.

[0099] 2.4 Stage Deliverable: Domain Adapter 1) Storage format: After this stage, the base model parameters remain frozen, and only the incremental matrix obtained from training is saved. .

[0100] 2) Technical characteristics: This matrix is ​​essentially a "parameterized compressed package" of knowledge in the field of rail transit. It carries the professional logic extracted from bucket A, but the model at this time does not yet have the ability to have dialogue and interaction, and only serves as a knowledge template for subsequent stages.

[0101] 3. Cross-model architecture adaptation and distribution alignment (Align CPT) The core technical objective of this phase is to address the parameter space distortion problem between the "knowledge adapter" and the "dialogue base," and through refined distribution alignment, enable the model to regain the original dialogue feel while possessing professional knowledge.

[0102] 3.1 Cross-model architecture integration Target: Select an instruction model (Chat / InstructionModel) that has been fine-tuned and aligned with human preferences as the target foundation. .

[0103] Mounting mechanism: The pedestal-based model in the first phase (LoRA CPT) Domain adapter matrix obtained during training Extract and mount it to the bypass of the Transformer layer corresponding to the instruction model.

[0104] Technical challenges: due to It is optimized under the activation value distribution of the Base model. Direct activation on the Chat model will cause drastic fluctuations in the probability distribution of the output layer, which manifests as the model's answers containing professional vocabulary, but with chaotic sentence structure and logical breaks.

[0105] 3.2 Formal Description and Objective Function To quantify and eliminate this distribution shift, this scheme introduces an alignment objective function based on Kullback-Leibler Divergence (KL) constraints. Let the model fitted with LoRA have the input... The predicted probability distribution is as follows The probability distribution of the original instruction model is: .

[0106] Optimize the objective function: in: Cross-entropy loss for word prediction on micro-samples in a vertical domain.

[0107] The divergence term, which measures the difference between the old and new distributions, is used to force the LoRA weights to maintain the original dialogue logic during updates, without deviating from the expression distribution of the Chat model.

[0108] To dynamically adjust the coefficients and ensure a balance between "knowledge retrieval" and "language sense alignment".

[0109] 3.3 Parameter Space Smoothing and Weight Scaling Compensation To eliminate the aforementioned "rejection reaction," this solution implements weight scaling compensation and gradient constraint strategies through model grafting middleware. 1) LoRA state and dimension consistency: Inherited mounting: Maintains the mounting obtained from the first phase of CPT training. The high-rank LoRA matrix structure remains unchanged and is directly mounted to the instruction model. The corresponding Transformer bypass.

[0110] Freeze control: Freeze command model The full backbone weight, only the high-rank Set it to trainable state.

[0111] 2) Extremely low learning rate configuration: Use an extremely low learning rate for minor weight adjustments.

[0112] 3) Micro-sample continuation: Data volume: Only a very small amount of high-quality rail transit-related corpus is required.

[0113] Training Epoch: Limited to One epoch is used to prevent overfitting.

[0114] 4) Training Intent: This step is not aimed at “learning new knowledge”, but rather at correcting the weight responses in the LoRA matrix through backpropagation to ensure the non-conflicting coexistence of knowledge access and conversational fluency.

[0115] 3.4 Distribution Alignment Targets and Effects Alignment objective: Eliminate the spatial distribution differences between the Base adapter weights and the Chat base weights.

[0116] Expected outcome: To resolve the issue of "distorted speech after knowledge injection." After this alignment stage, the model will be able to accurately invoke speech with the fluent tone characteristic of the Chat model. The stored rail transit terminology enables a shift from rote memorization to natural expression.

[0117] 4. Business Intent Triggering and Instruction Activation (LoRA SFT) This phase utilizes high-quality business alignment data on a very small scale to activate the model's intent recognition and task execution capabilities in specific rail transit scenarios, ultimately transforming it from a "general-purpose platform" to an "industry expert."

[0118] 4.1 Alignment of Business Intent 1) Target of operation: The fusion model after the second stage (Align CPT) processing .

[0119] 2) LoRA rank strategy: Knowledge adapter maintenance: Maintaining the high rank of the second-phase mount ( ) freeze.

[0120] Task Adapter Injection: Adds low-rank tasks (LoRAs) to the model or attaches them to specific layers. or ).

[0121] Training objective: To leverage the implicit knowledge of rail transit stored in a high-rank matrix in a parameterized manner, and to quickly activate the ability to follow specific business instructions through a low-rank task adapter, thereby achieving the effect of "heavy knowledge, fast instructions".

[0122] 3) Core Data: Supervised fine-tuning (SFT) is performed using C-buckets (business scenario instruction pairs). During the SFT process, 10%-20% of general conversational instructions (such as casual conversation, translation, and basic logic questions) are forcibly mixed in.

[0123] 4) Technical features: Minimalist Sample Strategy: Since the model has absorbed ample domain knowledge through the LoRA matrix in the first two stages, only 100 to 1000 high-quality QA statements are needed to trigger intent in this stage.

[0124] Instruction pair design: Covers core business scenarios of rail transit, such as: fault handling rule query, train dispatching command generation, equipment maintenance suggestions, etc.

[0125] 5) Technical Intent: Degradation of defensive capabilities: Addressing the issue of "loss of general interaction capabilities" that may occur after the model becomes highly verticalized.

[0126] Co-evolution: Ensure that when the model answers technical questions such as "how to handle the red light band of the signal horn", it can still maintain the affinity and multi-turn dialogue logic similar to the native Chat model.

[0127] 4.2 Deliverables: Business Adapter Matrix 1) Final form: Acquiring both in-depth industry knowledge (from...) The final business adapter matrix with instruction compliance capabilities (activated by the C bucket). .

[0128] 2) Technical advantages: This adapter decouples knowledge and capabilities, allowing users to quickly switch between different Mini SFT adapters based on different business segments (such as scheduling and maintenance versions) without having to go through the time-consuming CPT process again.

[0129] 5. Verification and Monitoring System To ensure the reliability of the model after three-stage training in the rail transit production environment, this solution constructs a comprehensive monitoring system covering "basic capabilities, industry knowledge, business logic, and safety red lines": 5.1 Automated Dynamic Evaluation Pipeline 1) Knowledge Retention Calculation (PPL-Base Line): Calculate perplexity on the A-bucket test set and compare the changes in values ​​before and after CPT. It is required that after injecting domain knowledge, the prediction probability of specialized terms must increase to a significant threshold (e.g., a PPL decrease of more than 20%).

[0130] 2) Instruction Compliance Rate Test (IF-Eval): Construct a mixed test set containing 500 general instructions and 500 rail transit instructions to verify the accuracy of the model when executing "formatted output" and "multi-step logical reasoning" and monitor whether there is a decline in instruction comprehension due to verticalization.

[0131] 5.2 Evaluation Indicators Specific to the Rail Transit Industry 1) Terminology Recognition Accuracy (TER): For regularized protected objects such as signal numbers and turnout numbers, verify whether the model has numbering errors or format breaks when generating responses.

[0132] 2) Rule-Consistency Score: Establish a "standard answer knowledge base" and use semantic similarity algorithms (such as BERTScore) or long text comparison technology to quantify the consistency between the handling process output by the model and the original "Traffic Organization Rules".

[0133] 5.3 Expert-level alignment evaluation based on LLM-as-a-Judge 1) High-order model dimensional alignment: Call upon more powerful models (such as GPT-4 or Qwen-Max) as "digital reviewers" to give a sliding score of 1-10 points from two dimensions: professionalism and safety.

[0134] 2) Logic Chain Verification: Specifically for "Fault Handling" type outputs, verify whether the logic chain is complete. For example, when handling "vehicle body electrification" faults, does the model mention "stop protection" before mentioning "fault reporting"? The logical order must not be reversed.

[0135] The following are system embodiments of this application, which can be used to execute the method embodiments of this application. For details not disclosed in the system embodiments of this application, please refer to the method embodiments of this application.

[0136] Figure 2 This is a schematic diagram of the structure of the large-scale collaborative evolution training system for the rail transit field provided in this application embodiment. For example... Figure 2 As shown, the system includes: The acquisition module 201 is used to acquire heterogeneous corpus in the rail transit field, perform functional bucketing on the heterogeneous corpus, and obtain domain implicit knowledge data, general logical balance data, and business intent instruction data. The identification module 202 is used to identify rail transit professional identifiers in domain implicit knowledge data based on preset identifier rules, add rail transit professional identifiers to the word segmenter to add new words, and perform noise cleaning and overlapping window segmentation on the domain implicit knowledge data. The pre-training module 203 is used to freeze the backbone weights of the base model and continuously pre-train the high-rank domain adaptation matrix based on a dynamic mixed batch of domain implicit knowledge data and general logical balance data to generate a domain knowledge adapter. The alignment training module 204 is used to attach the domain knowledge adapter to the bypass of the transformation layer corresponding to the instruction model, freeze the backbone weights of the instruction model, and perform alignment training on the domain knowledge adapter based on the secondary word prediction loss and output distribution divergence constraints to generate a fusion model. The supervised training module 205 is used to freeze the domain knowledge adapter in the fusion model, attach a low-rank task adaptation matrix, and perform supervised training based on business intent instruction data and general instruction data to generate a large-scale collaborative evolution model for rail transit business instruction processing.

[0137] In some embodiments, Figure 2 The acquisition module 201 acquires the original corpus containing rail transit professional text, general Chinese text, and business question and answer text; generates corpus function tags according to the corpus source, text structure, and training purpose of the original corpus; based on the corpus function tags, the original corpus is divided into domain implicit knowledge data for continuous pre-training, general logical balance data for language logic constraints, and business intent instruction data for supervised training; and establishes the calling relationship between various types of data and corresponding training stages.

[0138] In some embodiments, Figure 2 The identification module 202 constructs identifier matching rules corresponding to the professional number text based on the coding structure and context boundary of the rail transit business object; it uses the identifier matching rules to scan the domain implicit knowledge data segment by segment to extract candidate identifiers that conform to the complete coding structure; based on the semantic dependency relationship and frequency of occurrence of the candidate identifiers in adjacent business statements, it performs deduplication, normalization and validity screening on the candidate identifiers to generate a set of rail transit professional identifiers.

[0139] In some embodiments, Figure 2 The recognition module 202 generates a new word segmentation table for the word segmenter based on the set of professional identifiers for rail transit, and configures the overall word segmentation attributes of each professional identifier in the new word segmentation table; it performs word segmentation processing on the domain implicit knowledge data based on the overall word segmentation attributes, so that the professional identifiers participate in training as complete words; it performs layout noise cleaning, invalid text removal and character anomaly repair on the domain implicit knowledge data before and after word segmentation processing to generate cleaned text; it performs sliding segmentation of the cleaned text according to the preset window length and overlap ratio to generate domain knowledge training segments with context connection markers.

[0140] In some embodiments, Figure 2The pre-training module 203 keeps the backbone weights of the base model frozen and configures a high-rank domain adaptation matrix in the transformation layer of the base model. Training segments are extracted from the domain implicit knowledge data and the general logical balance data according to the preset sampling weights, and dynamic hybrid batches are generated. Sub-word prediction training is performed with the dynamic hybrid batches as input, and only the high-rank domain adaptation matrix is ​​updated. After the training reaches the preset stopping condition, the updated high-rank domain adaptation matrix is ​​saved to obtain the domain knowledge adapter.

[0141] In some embodiments, Figure 2 The alignment training module 204 determines the corresponding transformation layer bypass in the instruction model based on the hierarchical position and matrix dimension of the domain knowledge adapter in the base model, and attaches the domain knowledge adapter to the transformation layer bypass layer by layer; freezes the backbone weights of the instruction model, and performs a small update on the domain knowledge adapter using a learning rate lower than that of continuous pre-training; inputs the professional continuation corpus into the attached instruction model, calculates the word prediction loss and the divergence loss between the output probability distributions before and after attachment, and updates the domain knowledge adapter based on the word prediction loss and the divergence loss to generate the fusion model.

[0142] In some embodiments, Figure 2 The alignment training module 204 reads the hierarchical index, target mapping sub-layer, matrix rank, and input / output dimensions of the domain knowledge adapter to generate adapter mounting metadata. Based on the adapter mounting metadata, it searches for transformation layer bypasses with the same hierarchical semantics and compatible matrix dimensions in the instruction model. While keeping the domain knowledge adapter matrix structure unchanged, it performs scaling compensation on the mounting weights and maps the domain knowledge adapter to the corresponding transformation layer bypass.

[0143] In some embodiments, Figure 2 The supervised training module 205 keeps the domain knowledge adapter in the fusion model frozen. Based on the business scenario corresponding to the business intent instruction data, it attaches a low-rank task adaptation matrix to the target transformation layer of the fusion model. It mixes the business intent instruction data and general instruction data in a preset ratio to generate instruction training batches. It performs supervised training based on the instruction training batches, updating only the low-rank task adaptation matrix. It loads the trained low-rank task adaptation matrix and the domain knowledge adapter together to generate a co-evolutionary large model.

[0144] In some embodiments, Figure 2After generating the co-evolutionary large model, the matching module 206 constructs evaluation samples covering professional knowledge retention, general instruction compliance, professional identifier stability, and business rule consistency based on domain tacit knowledge data, general logical balance data, and business intent instruction data. The evaluation samples are input into the co-evolutionary large model to obtain the model output results, and the model output results are matched with the standard knowledge base and preset output constraints to generate model capability evaluation results. Based on the model capability evaluation results, knowledge injection bias, alignment bias, or intent activation bias are determined, and the sampling weights of dynamic mixed batches, the divergence constraint weights of alignment training, or the instruction data ratio of supervised training are adjusted based on the biases.

[0145] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0146] Figure 3 This is a schematic diagram of the electronic device 3 provided in an embodiment of this application. Figure 3 As shown, the electronic device 3 of this embodiment includes: a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program 303, it implements the steps in the various method embodiments described above. Alternatively, when the processor 301 executes the computer program 303, it implements the functions of each module / unit in the various system embodiments described above.

[0147] Electronic device 3 can be a desktop computer, laptop, handheld computer, cloud server, or other electronic device. Electronic device 3 may include, but is not limited to, processor 301 and memory 302. Those skilled in the art will understand that... Figure 3 This is merely an example of electronic device 3 and does not constitute a limitation on electronic device 3. It may include more or fewer components than shown, or different components.

[0148] The processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0149] The memory 302 can be an internal storage unit of the electronic device 3, such as a hard disk or memory of the electronic device 3. The memory 302 can also be an external storage device of the electronic device 3, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 3. The memory 302 can also include both internal and external storage units of the electronic device 3. The memory 302 is used to store computer programs and other programs and data required by the electronic device.

[0150] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0151] If integrated modules / units are implemented as software functional units and sold or used as independent products, they can be stored in a readable storage medium (e.g., a computer-readable storage medium). Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program may include computer program code, which may be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable storage medium may include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.

[0152] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for collaborative evolution training of large-scale models in the field of rail transit, characterized in that, include: Obtain heterogeneous corpus in the rail transit field, and perform functional bucketing on the heterogeneous corpus to obtain domain implicit knowledge data, general logical balance data, and business intent instruction data; Based on preset identifier rules, the rail transit professional identifier in the domain implicit knowledge data is identified, the rail transit professional identifier is added to the word segmenter to add new words, and the domain implicit knowledge data is cleaned of noise and segmented by overlapping windows. Freeze the backbone weights of the base model, and continuously pre-train the high-rank domain adaptation matrix based on the dynamic mixed batch of the domain implicit knowledge data and the general logical balance data to generate a domain knowledge adapter; The domain knowledge adapter is attached to the bypass of the transformation layer corresponding to the instruction model, the backbone weights of the instruction model are frozen, and the domain knowledge adapter is aligned and trained based on the secondary word prediction loss and output distribution divergence constraint to generate a fusion model. The domain knowledge adapter in the fusion model is frozen, a low-rank task adaptation matrix is ​​attached, and supervised training is performed based on the business intent instruction data and general instruction data to generate a large-scale collaborative evolution model for rail transit business instruction processing.

2. The method according to claim 1, characterized in that, The process of acquiring heterogeneous corpora in the rail transit field, performing functional bucketing on the heterogeneous corpora, and obtaining domain implicit knowledge data, general logical balance data, and business intent instruction data includes: Obtain raw corpus containing rail transit-specific texts, general Chinese texts, and business Q&A texts; Generate corpus function labels based on the corpus source, text structure, and training purpose of the original corpus; Based on the functional tags of the corpus, the original corpus is divided into domain implicit knowledge data for continuous pre-training, general logical balance data for language logic constraints, and business intent instruction data for supervised training. Establish the calling relationship between various types of data and corresponding training stages.

3. The method according to claim 1, characterized in that, The method of identifying rail transit specialty identifiers in the domain tacit knowledge data based on preset identifier rules includes: Based on the coding structure and context boundaries of rail transit business objects, construct identifier matching rules corresponding to professional number texts; The domain implicit knowledge data is scanned segment by segment using the identifier matching rules to extract candidate identifiers that conform to the complete coding structure; Based on the semantic dependency relationship and frequency of occurrence of the candidate identifiers in adjacent business statements, the candidate identifiers are deduplicated, normalized, and filtered for validity to generate a set of rail transit professional identifiers.

4. The method according to claim 3, characterized in that, The step of adding the rail transit professional identifier to the word segmenter to create new words, and performing noise cleaning and overlapping window segmentation on the domain implicit knowledge data, includes: The word segmenter generates a new word element table based on the set of rail transit professional identifiers, and configures the overall word segmentation attributes of each professional identifier in the new word element table; Based on the overall word segmentation attributes, the domain implicit knowledge data is segmented, so that the professional identifiers participate in training as complete word units; The domain implicit knowledge data before and after word segmentation is cleaned up for layout noise, invalid text is removed, and character anomalies are repaired to generate clean text. The cleaned text is slid-segmented according to a preset window length and overlap ratio to generate domain knowledge training segments with context connection markers.

5. The method according to claim 1, characterized in that, The frozen base model backbone weights, based on a dynamic mixed batch of the domain implicit knowledge data and the general logical balance data, continuously pre-train the high-rank domain adaptation matrix to generate a domain knowledge adapter, including: Keep the backbone weights of the base model frozen, and configure the high-rank neighborhood adaptation matrix in the transformation layer of the base model; Training segments are extracted from the domain implicit knowledge data and the general logical balance data according to preset sampling weights, and then mixed to generate dynamic mixed batches. The dynamic mixed batch is used as input to perform word prediction training, and only the high-rank domain adaptation matrix is ​​updated. After the training reaches the preset stopping condition, the updated high-rank domain adaptation matrix is ​​saved to obtain the domain knowledge adapter.

6. The method according to claim 1, characterized in that, The process of attaching the domain knowledge adapter to the bypass of the transformation layer corresponding to the instruction model, freezing the backbone weights of the instruction model, and aligning and training the domain knowledge adapter based on the secondary word prediction loss and output distribution divergence constraints to generate a fusion model includes: Based on the hierarchical position and matrix dimension of the domain knowledge adapter in the base model, the corresponding transformation layer bypass in the instruction model is determined, and the domain knowledge adapter is mounted to the transformation layer bypass layer by layer. Freeze the backbone weights of the instruction model and perform minor updates on the domain knowledge adapter using a learning rate lower than that used in continuous pre-training; The professional continuation writing corpus is input into the mounted instruction model, and the word prediction loss and the divergence loss between the output probability distributions before and after mounting are calculated respectively. The domain knowledge adapter is then updated based on the word prediction loss and the divergence loss to generate the fusion model.

7. The method according to claim 6, characterized in that, The step of determining the corresponding transformation layer bypass in the instruction model based on the hierarchical position and matrix dimension of the domain knowledge adapter in the base model, and attaching the domain knowledge adapter to the transformation layer bypass layer by layer, includes: Read the hierarchical index, target mapping sub-layer, matrix rank, and input / output dimensions of the domain knowledge adapter to generate adapter mounting metadata; Based on the adapter mounting metadata, find a transformation layer bypass with the same hierarchical semantics and compatible matrix dimension in the instruction model; While keeping the structure of the domain knowledge adapter matrix unchanged, the mounting weights are scaled and compensated, and the domain knowledge adapter is mapped to the corresponding transformation layer bypass.

8. The method according to claim 1, characterized in that, The domain knowledge adapter in the frozen fusion model is attached with a low-rank task adaptation matrix, and supervised training is performed based on the business intent instruction data and general instruction data to generate a large-scale collaborative evolution model for rail transit business instruction processing, including: Keep the domain knowledge adapter in the fusion model in a frozen state, and attach a low-rank task adaptation matrix to the target transformation layer of the fusion model according to the business scenario corresponding to the business intent instruction data. The business intent instruction data and general instruction data are mixed in a preset ratio to generate instruction training batches; Supervised training is performed in batches based on the instructions, and only the low-rank task adaptation matrix is ​​updated. The trained low-rank task adaptation matrix is ​​co-loaded with the domain knowledge adapter to generate the co-evolutionary large model.

9. The method according to claim 8, characterized in that, After generating the aforementioned large-scale co-evolutionary model, the following is also included: Based on the domain tacit knowledge data, the general logical balance data, and the business intent instruction data, an evaluation sample covering professional knowledge retention, general instruction compliance, professional identifier stability, and business rule consistency is constructed. The evaluation samples are input into the co-evolutionary large model to obtain the model output results. The model output results are then matched with the standard knowledge base and preset output constraints to generate model capability evaluation results. Based on the model capability evaluation results, knowledge injection bias, alignment bias, or intent activation bias are determined, and the sampling weights of dynamic mixed batches, the divergence constraint weights of alignment training, or the instruction data ratios of supervised training are adjusted based on the biases.

10. A large-scale model collaborative evolution training system for the field of rail transit, characterized in that, include: The acquisition module is used to acquire heterogeneous corpus in the rail transit field, and to perform functional bucketing on the heterogeneous corpus to obtain domain implicit knowledge data, general logical balance data and business intent instruction data. The identification module is used to identify rail transit professional identifiers in the domain implicit knowledge data based on preset identifier rules, add the rail transit professional identifiers to the word segmenter to add new words, and perform noise cleaning and overlapping window segmentation on the domain implicit knowledge data. The pre-training module is used to freeze the backbone weights of the base model and continuously pre-train the high-rank domain adaptation matrix based on the dynamic mixed batch of the domain implicit knowledge data and the general logical balance data to generate a domain knowledge adapter. The alignment training module is used to attach the domain knowledge adapter to the bypass of the transformation layer corresponding to the instruction model, freeze the backbone weights of the instruction model, and perform alignment training on the domain knowledge adapter based on the secondary word prediction loss and output distribution divergence constraints to generate a fusion model. The supervised training module is used to freeze the domain knowledge adapter in the fusion model, attach a low-rank task adaptation matrix, and perform supervised training based on the business intent instruction data and general instruction data to generate a large-scale collaborative evolution model for rail transit business instruction processing.