Signal field large model training method, service processing method, device and equipment
By acquiring general knowledge and business scenario datasets in the signaling field, we trained an open-source large model and established an adapted large model for the signaling field. This solved the problems of poor data quality and low adaptability, and enabled efficient business processing in the field of urban rail transit signaling.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-04-07
AI Technical Summary
The application and practical use of large language models in the field of urban rail transit signaling faces problems such as poor data quality and low adaptability, which makes it impossible to effectively apply vertical domain models.
By acquiring general knowledge datasets and business scenario datasets in the signal domain, we train an open-source general-purpose large model to establish a basic model with fundamental cognitive capabilities in the signal domain. We then further train the model based on business scenario data to form a scenario-adapted large model in the signal domain, which can be used for code generation, test case generation, and engineering data configuration.
It has achieved precise matching between large-scale signaling models and actual business scenarios of enterprises, improved business processing efficiency, solved the problem that vertical domain models cannot be implemented, and provided a feasible technical solution for the digital transformation of urban rail transit signaling.
Smart Images

Figure CN121809628A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of urban rail transit signaling technology, and in particular to a large-scale model training method, service processing method, apparatus and equipment for the signaling field. Background Technology
[0002] With the deep penetration of artificial intelligence technology in the transportation sector, the cross-integration of AI and rail transit has become a core direction for driving the industry's digital transformation. While large language models, with their powerful text understanding and generation capabilities, have achieved large-scale application in general fields, their implementation and practical application in the specialized and highly scenario-driven vertical field of urban rail transit signaling still face significant technical bottlenecks. Summary of the Invention
[0003] This invention provides a training method, business processing method, apparatus, and device for a large-scale signaling model. By acquiring a general knowledge dataset and a business scenario dataset for the signaling domain, it not only provides accurate fundamental theoretical knowledge and industry technical standards for the large-scale signaling model, but also ensures that the training data is accurately matched with the actual business scenarios of enterprises. This solves the problems of poor data quality and low adaptability in traditional vertical domain model training from the dual dimensions of comprehensive knowledge coverage and accurate scenario matching. Ultimately, the trained large-scale signaling model can accurately complete tasks such as signaling domain code-assisted generation, test case generation, and engineering data configuration, significantly improving the efficiency of signaling domain business processing. It effectively solves the problem that vertical domain models cannot be implemented, and provides a feasible technical solution for the digital transformation of the urban rail transit signaling field.
[0004] In a first aspect, the present invention provides a method for training a large model in the signal domain, comprising the following steps: Acquire general knowledge datasets and business scenario datasets in the signal domain; The open-source general large model is trained based on the aforementioned signal domain general knowledge dataset to obtain a basic signal domain large model with basic cognitive capabilities in the signal domain; The basic signal domain large model is trained based on the signal domain business scenario dataset to obtain a scenario-adapted signal domain large model; the scenario-adapted signal domain large model is used for signal domain code auxiliary generation, test case generation, and engineering data configuration.
[0005] According to the signal domain large model training method provided by the present invention, the step of obtaining the signal domain general knowledge dataset includes: We acquire professional literature, technical standards, and public business data in the field of urban rail transit signaling, and clean and label them to obtain an initial general knowledge dataset in the field of signaling. Based on the initial signal domain general knowledge dataset, a preset lightweight model is trained, and the signal domain general knowledge dataset is obtained based on the accuracy of the trained lightweight model in the basic question answering task in the signal domain.
[0006] According to the present invention, a method for training a large model in the signal domain is provided, wherein the signal domain general knowledge dataset is obtained based on the accuracy of the trained lightweight model in a basic question-answering task in the signal domain, and includes: If the accuracy of the lightweight model in the basic question-answering task of the signal domain is greater than or equal to a threshold, the initial signal domain general knowledge dataset will be used as the signal domain general knowledge dataset.
[0007] According to the present invention, a method for training a large model in the signal domain is provided, wherein the signal domain general knowledge dataset is obtained based on the accuracy of the trained lightweight model in a basic question-answering task in the signal domain, and includes: If the accuracy of the lightweight model in the basic question-answering task of the signal domain is less than a threshold, the initial signal domain general knowledge dataset is supplemented and corrected based on the erroneous question-answering results of the lightweight model in the basic question-answering task of the signal domain. The training and validation process of the lightweight model is repeated until the accuracy of the lightweight model in the basic question-answering task of the signal domain is greater than or equal to the threshold. The supplemented and corrected initial signal domain general knowledge dataset is then used as the signal domain general knowledge dataset.
[0008] According to the present invention, a large model training method for the signal domain is provided, which obtains the signal domain business scenario dataset based on the following method: Acquire internal business scenario data of the enterprise; the business scenario data includes at least one of the following: historical cases of code generation in the signal domain, test case writing records, and practical logs of engineering data configuration; The business scenario data is matched with the preset high-value business scenario requirement standards to obtain the signal domain business scenario dataset; the preset high-value business scenario requirement standards include at least one of the following: code syntax compliance rate, coverage of core fault points of signal equipment in test cases, and accuracy of engineering data configuration.
[0009] Secondly, the present invention provides a business processing method, comprising the following steps: Obtain unprocessed business requirements in the signal domain, which include at least one of the following: code generation assistance requirements in the signal domain, test case generation requirements, and engineering data configuration requirements; The signal domain business requirements to be processed are input into a scenario-adapted large signal domain model, and the business processing results are output; the scenario-adapted large signal domain model is trained based on the signal domain large model training method described in the first aspect.
[0010] Thirdly, the present invention also provides a large-scale signal domain model training device, comprising the following modules: The acquisition module is used to acquire general knowledge datasets and business scenario datasets in the signal domain. The first training module is used to train the open-source general large model based on the signal domain general knowledge dataset to obtain a basic signal domain large model with basic cognitive ability in the signal domain. The second training module trains the basic signal domain large model based on the signal domain business scenario dataset to obtain a scenario-adapted signal domain large model; the scenario-adapted signal domain large model is used for signal domain code auxiliary generation, test case generation, and engineering data configuration.
[0011] Fourthly, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the signal domain large model training method as described above.
[0012] Fifthly, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the signal domain large model training method as described above.
[0013] In a sixth aspect, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the signal domain large model training method as described above.
[0014] The signal domain large-scale model training method, business processing method, apparatus, and equipment provided by this invention, by acquiring signal domain general knowledge datasets and signal domain business scenario datasets, not only provide the signal domain large-scale model with accurate signal domain basic theoretical and industry technical standard knowledge, but also enable the training data to accurately match the actual business scenarios of enterprises. This solves the problems of poor data quality and low adaptability in traditional vertical domain model training from the dual dimensions of comprehensive knowledge coverage and accurate scenario matching. Ultimately, the trained signal domain large-scale model can accurately complete signal domain code-assisted generation, test case generation, and engineering data configuration tasks, significantly improving the efficiency of signal domain business processing, effectively solving the problem of vertical domain models being unable to be implemented, and providing a feasible technical solution for the digital transformation of the urban rail transit signaling field. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0016] Figure 1 This is a schematic diagram illustrating the current state of AI ToB applications provided by this invention.
[0017] Figure 2 This is one of the flowcharts illustrating the large model training method for the signal domain provided by this invention.
[0018] Figure 3 This is the second flowchart of the large model training method for the signal domain provided by this invention.
[0019] Figure 4 This is one of the flowcharts illustrating the method for obtaining model training datasets provided by this invention.
[0020] Figure 5 This is the second flowchart illustrating the method for obtaining the model training dataset provided by this invention.
[0021] Figure 6 This is a schematic diagram of the structure of the large model training device for the signal domain provided by the present invention.
[0022] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0024] The following is combined with Figures 1-7 This invention describes a large-scale signal domain model training method, service processing method, apparatus, and device.
[0025] To facilitate a clearer understanding of the technical solutions of the various embodiments of this application, some technical content related to the various embodiments of this application will be introduced first.
[0026] The development of artificial intelligence (AI) technology has greatly boosted productivity. From AI-generated office solutions to smart manufacturing and the digital infrastructure of industries, large-scale language models are rapidly being implemented in enterprises. With its foundational computing power, embodied intelligence, and iterative updates to trillion-parameter models, AI has passed conceptual validation and entered the industrial cycle. The practical application of AI is based on the deep integration of large-scale models across vertical domains with enterprise production scenarios and application data, reshaping industrial scale and innovation paradigms.
[0027] The intelligence of a Large Language Model (LLM) comes from multimodal information, primarily semantic textual information, labeled with human language. An open-source Large Language Model can perceive and manipulate most of the information within the realm of natural language. If a task or job is to be automated by a Large Language Model, then that task must be able to be described and processed in detail by language.
[0028] Therefore, the implementation of large language models involves two issues: ① Can the task be described in detail in language? ② Can the evaluation of the task performance be described in detail in language? Even in the information domain, large language models do not yet possess full automation capabilities. Large models exist in a gray area between rules and ambiguity. If a domain is easily automated, traditional programs can achieve automation based on rule-based scripts. However, other domains are difficult to fully automate due to the ambiguity and complexity of multiple processing steps, thus requiring manual handling. Within an enterprise, certain tasks or work often encounter difficult decision-making problems due to insufficient information sources. When multiple choices are involved, automated rule scripts fail to execute smoothly, and large language models also perform less than ideally at the decision-making stage. Regarding the performance of large language models, their mature applications mainly focus on the following aspects: ① Transform precise descriptive language into functional code / or templated official documents; ② Generate texts with high tolerance based on specific prompts, such as copywriting, novels, and scripts; ③ Simulate language style, convert language style, or perform language tasks such as translation; ④ In fields where language precision is high, complete tasks with significant workflows, such as business agents.
[0029] ⑤ Based on open domain data, complete multimodal recognition and classification work in general domains, such as the most common symptom diagnosis and AI legal consultation in the field of people's livelihood.
[0030] ⑥ Complete encyclopedia knowledge Q&A to assist in knowledge retrieval.
[0031] In summary, the primary application scenarios for large language models are in the language domain rather than the decision-making domain. Benefiting from the exponential growth of data in the mobile internet era, large language models have access to abundant data samples and a data annotation industry, leading to numerous application scenarios in open domains. According to an IDC representative at the World Artificial Intelligence Conference, most tasks in enterprises require interaction with the physical world, yet multimodal task services based on large language models are relatively rare. Furthermore, they also need to handle ambiguous and polysemous tasks that are difficult to express linguistically, and coordinate various relationships and schedules. For most people, their understanding of large language models is limited to a search tool without embedded ads.
[0032] Large-scale models are rapidly being adopted by enterprises, but the large-scale implementation of enterprise-level AI-driven transformation faces significant challenges, such as... Figure 1 As shown, both sides in the AI ToB sector (computing power providers and enterprises) possess knowledge barriers and industry moats. Enterprises urgently need AI to enhance efficiency and empower businesses, while computing power providers are even more eager to commercialize their technologies and monetize them. However, the main difficulty in implementing AI ToB lies in the creation of large model datasets for application scenarios. Computing power providers cannot directly empower enterprises because they do not understand the internal industry knowledge and data composition; similarly, application enterprises are not willing to easily open up their internal data and core production processes. Therefore, enterprises need to build data centers that serve their internal applications. As an intermediary node connecting the enterprise's internal and external computing power providers, the data center is responsible for the visualization of specific application scenario requirements, computing power planning, and data planning and management internally. Externally, it collaborates with computing power providers to jointly develop and deploy application-end applications for data cleaning and data governance toolchains. In addition, the training of large-scale models in vertical domains has extremely high requirements for computing power resources, while the pre-training and fine-tuning of non-sensitive application models can leverage the cloud computing power resources of computing power providers.
[0033] As shown in Table 1, the current implementation of large-scale models within enterprises must focus on obtaining feedback from business engineers. Feedback from engineers with limited business understanding and a lack of a systemic perspective lacks technical depth and cannot generate effective, high-quality training data. Senior engineers essentially act as "data labelers." Therefore, creating high-quality, high-value business datasets is the core of large-scale model implementation and also the primary task of the data center.
[0034] Table 1 aspect ToC (Consumer-Oriented) ToB (Business-oriented) Data volume We need large-scale user behavior data (such as clicks, purchase records, etc.). We need industry-specific data for businesses; the amount of data may be small, but the quality requirements are high. Data labeling Some labeling is required, such as for speech recognition and image classification. High-quality professional data annotation typically requires the participation of industry experts. technical complexity The technical requirements are moderate, and they are usually based on mature algorithms (such as collaborative filtering and rule-based recommendation). It is technically complex, involving advanced technologies such as deep learning and natural language processing. Intelligence level Highly personalized, real-time responsive systems are required (e.g., recommendation systems, voice assistants). High precision, customization, and specialization are required, such as intelligent decision support and industry analysis. Fault tolerance and interpretability It has high fault tolerance, and a certain amount of error has a small impact on the user experience. It has low fault tolerance and must provide explainability so that companies can make the right decisions. Figure 2 This is one of the flowcharts illustrating the large-scale signal domain model training method provided by the present invention. The method includes the following: Step 201: Obtain the signal domain general knowledge dataset and the signal domain business scenario dataset.
[0035] Specifically, in this embodiment, a general knowledge dataset and a business scenario dataset for the signaling domain are first acquired. Optionally, the general knowledge dataset can come from professional literature, national / industry technical standards, and public business data in the field of urban rail transit signaling, such as operation and maintenance records of rail transit signaling systems and general fault handling logs from the past 5 years. Optionally, the business scenario dataset can come from business data from different scenarios within an enterprise, such as historical cases of signaling code generation, test case writing records, and practical logs of engineering data configuration. By acquiring the general knowledge dataset and the business scenario dataset for the signaling domain, this application can provide accurate basic theoretical and industry technical standard knowledge in the signaling domain for large-scale signaling models. On the other hand, it enables precise matching of training data with actual business scenarios of enterprises, providing efficient support for model training from both knowledge coverage and scenario adaptation dimensions, and effectively improving the training effect of large-scale signaling models.
[0036] Step 202: Train the open-source general large model based on the signal domain general knowledge dataset to obtain a basic signal domain large model with basic cognitive ability in the signal domain.
[0037] Specifically, it should be noted that related technologies directly apply open-source general-purpose models to the signal domain. Because these models lack signal domain-specific knowledge, misunderstandings of industry terminology, technical principles, or standards are common. This application trains the open-source general-purpose model using a signal domain general knowledge dataset, enabling the model to systematically learn and establish a signal domain knowledge system. This effectively corrects the cognitive biases of the general-purpose model and significantly improves the accuracy and effectiveness of the model in handling signal domain business. Optionally, a pre-training fine-tuning mode can be used, employing the AdamW optimizer, setting the initial learning rate and training iterations, and evaluating model performance.
[0038] Step 203: Train the basic signal domain large model based on the signal domain business scenario dataset to obtain a scenario-adapted signal domain large model; the scenario-adapted signal domain large model is used for signal domain code auxiliary generation, test case generation, and engineering data configuration.
[0039] Specifically, this application trains an open-source general-purpose large model based on a general dataset in the signal domain to obtain a basic signal domain large model with basic cognitive capabilities in the signal domain. Then, it further trains the basic signal domain large model based on a business scenario dataset in the signal domain to obtain a scenario-adapted signal domain large model. This enables the trained signal domain large model to not only understand the basic knowledge of the signal domain, but also to efficiently and accurately respond to the business needs of various scenarios in the signal domain, realize code generation, test case generation, and accurate configuration of engineering data in the signal domain, and effectively solve the problem that vertical domain models cannot be implemented in related technologies.
[0040] The method described above, by acquiring a general knowledge dataset and a business scenario dataset for the signaling domain, not only provides accurate fundamental theoretical knowledge and industry technical standards for the signaling domain's large-scale model, but also ensures that the training data is accurately matched with the actual business scenarios of enterprises. This solves the problems of poor data quality and low adaptability in traditional vertical domain model training from the dual dimensions of comprehensive knowledge coverage and accurate scenario matching. Ultimately, the trained large-scale signaling domain model can accurately complete tasks such as signaling domain code-assisted generation, test case generation, and engineering data configuration, significantly improving the efficiency of signaling domain business processing and effectively solving the problem that vertical domain models cannot be implemented. This provides a feasible technical solution for the digital transformation of the urban rail transit signaling field.
[0041] In some embodiments, obtaining a signal domain general knowledge dataset includes: We acquire professional literature, technical standards, and public business data in the field of urban rail transit signaling, and clean and label them to obtain an initial general knowledge dataset in the field of signaling. Based on the initial general knowledge dataset for the signal domain, a pre-defined lightweight model is trained, and the general knowledge dataset for the signal domain is obtained based on the accuracy of the trained lightweight model in the basic question-answering task in the signal domain.
[0042] Specifically, in this embodiment, professional literature, technical standards, and public business data in the field of urban rail transit signaling are first acquired, and then data cleaning is performed on them. Optionally, duplicate content can be removed, semantically confusing text generated by machine translation, redundant information irrelevant to the signaling field, and records with incomplete data formats can be deleted, and the data format can be unified to obtain an initial signaling field general knowledge dataset, ensuring the signaling field relevance, information validity, and knowledge structure of the dataset from the source.
[0043] Optionally, after obtaining the initial signal domain general knowledge dataset, this application further verifies the quality of the dataset. Optionally, a lightweight pre-trained model, such as BERT-base, can be used. The initial signal domain general knowledge dataset is divided into a training set and a test set in an 8:2 ratio. The test set is input into the trained lightweight model, and the accuracy of the model in the basic question-answering task is statistically analyzed. This quickly verifies whether the dataset meets the training requirements, ensures the accuracy and completeness of the knowledge in the general knowledge dataset, avoids large model training failure or poor performance due to data quality issues, effectively establishes a quantitative verification mechanism for dataset quality, solves the pain point of not being able to judge the quality of general knowledge data, and improves the training effect of the model.
[0044] It should be noted that traditional techniques typically use raw data directly for training large models, lacking effective verification of data quality. If the data contains knowledge errors or omissions, it can lead to training failure or poor performance of the large model. This application, through the accuracy feedback of a lightweight model, transforms dataset quality into a quantifiable metric. This allows for rapid verification of whether the data meets training requirements, precise identification and optimization of data defects, and ensures the accuracy and completeness of knowledge in the general dataset. Furthermore, the computational cost of the lightweight model is only 1 / 10 to 1 / 5 of that of open-source general-purpose large models. By using it to pre-verify dataset quality, the risk of poor data quality can be eliminated before formal training of the large model. Alternatively, directly training a large model with defective initial data not only wastes significant computational power and time but also requires data readjustment and retraining, significantly increasing trial-and-error costs. This application, through lightweight model pre-verification, avoids data problems in advance, ensuring that the large model is trained on high-quality data, reducing unnecessary retraining, significantly improving overall training efficiency, and lowering the trial-and-error costs of large model training.
[0045] The method described above cleans and annotates professional literature, technical standards, and public business data in the field of urban rail transit signaling to ensure the relevance of the dataset in the signaling field, the validity of information, and the structure of knowledge from the source. By using a lightweight pre-trained model, data quality risks are eliminated in advance before the formal training of the large model, avoiding repeated training of the large model and wasting computing power and time due to data defects. This significantly reduces the trial and error cost of large model training, improves the overall training efficiency, and effectively solves the problems of unreliable data quality, difficulty in locating data defects, and high trial and error cost of large model training in traditional technologies.
[0046] In some embodiments, a signal domain general knowledge dataset is obtained based on the accuracy of the trained lightweight model in a basic question-answering task in the signal domain, including: If the accuracy of the lightweight model in the basic question-answering task of the signal domain is greater than or equal to the threshold, the initial signal domain general knowledge dataset will be used as the signal domain general knowledge dataset.
[0047] Specifically, in this embodiment, a preset lightweight model is trained using an initial signal domain general knowledge dataset. The accuracy of the trained lightweight model in basic question-answering tasks within the signal domain is used to determine whether the dataset meets the requirements for training a large model. Optionally, the basic question-answering accuracy of the trained lightweight model on the test set can be compared with a preset threshold. If the accuracy is greater than or equal to the threshold, it indicates that the initial dataset contains complete and accurate coverage of basic signal domain theories, technical standards, and general terminology, and that the data format and annotations meet the model's learning requirements, requiring no additional optimization. Therefore, this initial signal domain general knowledge dataset can be directly identified as the final signal domain general knowledge dataset used for training open-source general-purpose large models. This achieves efficient and accurate quantitative verification of dataset quality, provides reliable data support for model training, and effectively improves model training performance.
[0048] The method described above, when the accuracy of the lightweight model in the basic question-answering task in the signal domain is greater than or equal to a threshold, uses the initial signal domain general knowledge dataset as the signal domain general knowledge dataset. This achieves efficient and accurate verification of the quality of the signal domain general knowledge dataset, provides reliable data support for model training, effectively avoids model cognitive bias caused by data quality issues, and effectively improves model training effect and efficiency.
[0049] In some embodiments, a signal domain general knowledge dataset is obtained based on the accuracy of the trained lightweight model in a basic question-answering task in the signal domain, including: If the accuracy of the lightweight model in the basic question-answering task of the signal domain is less than the threshold, the initial signal domain general knowledge dataset is supplemented and corrected based on the erroneous question-answering results of the lightweight model in the basic question-answering task of the signal domain. The training and validation process of the lightweight model is repeated until the accuracy of the lightweight model in the basic question-answering task of the signal domain is greater than or equal to the threshold. The supplemented and corrected initial signal domain general knowledge dataset is then used as the signal domain general knowledge dataset.
[0050] Specifically, in this embodiment of the application, if the accuracy of the lightweight model in the basic question-answering task in the signal domain is less than a threshold, it indicates that the knowledge coverage of basic theories, technical standards, and general terms in the signal domain in the initial dataset is incomplete or incorrect. In this case, the incorrect question-answering results of the model can be analyzed to locate the root cause of the problem, check whether there are omissions or inaccuracies in the initial dataset, and accurately determine the specific content that needs to be supplemented or corrected in the initial dataset through the precise mapping between the incorrect results and data defects. This achieves precise optimization of the dataset, greatly improves the efficiency and quality of dataset optimization, and ensures that there is no missing knowledge and no errors in the description in the general knowledge dataset.
[0051] Optionally, after supplementing and correcting the initial signal domain general knowledge dataset, the supplemented and corrected initial signal domain general knowledge dataset can be re-divided into training set and test set. Repeat the lightweight model training and basic question answering task accuracy evaluation process. If the accuracy is still less than the threshold, repeat the above steps again until the accuracy of the lightweight model is greater than or equal to the threshold. At this time, the final supplemented and corrected initial dataset is determined as the signal domain general knowledge dataset for model training. This effectively makes up for the knowledge gaps and errors in the dataset, effectively improves the accuracy and completeness of the dataset, and provides high-quality data support for model training.
[0052] The method described above, when the accuracy of the lightweight model in the basic question-answering task in the signal domain is less than a threshold, supplements and corrects the initial general knowledge dataset in the signal domain based on the erroneous question-answering results of the lightweight model in the basic question-answering task in the signal domain. This achieves accurate location and correction of knowledge gaps in the dataset, effectively improves the accuracy and completeness of the dataset, provides high-quality data support for model training, and improves the training efficiency and reliability of the model.
[0053] In some embodiments, signal domain business scenario datasets are obtained based on the following methods: Acquire internal business scenario data from the enterprise; business scenario data includes at least one of the following: historical cases of code generation in the signal domain, test case writing records, and practical logs of engineering data configuration. The business scenario data is matched with the preset high-value business scenario requirement standards to obtain the signal domain business scenario dataset; the preset high-value business scenario requirement standards include at least one of the following: code syntax compliance rate, coverage of core fault points of signal equipment in test cases, and accuracy of engineering data configuration.
[0054] Specifically, this application acquires a dataset of business scenarios in the signaling domain, enabling precise matching of training data with actual enterprise business scenarios. This provides efficient support for model training from the perspective of scenario adaptation, effectively improving the training performance of large-scale signaling domain models. Optionally, historical operational data directly related to core business scenarios in the signaling domain can be collected from within the enterprise, such as historical code generation cases, test case writing records, and engineering data configuration logs. It should be noted that business scenario data contains real business logic and operational details. After learning from this type of data, the model can better adapt to the actual business needs of the enterprise, improving the model's processing performance for actual business operations.
[0055] Optionally, after acquiring internal business scenario data, this application matches the business scenario data with pre-defined high-value business scenario requirement standards to filter signal domain business scenario data, ensuring the dataset's high value for model training. Optionally, signal domain business scenario data can be filtered based on code syntax compliance rate, coverage of core fault points in signal equipment in test cases, and accuracy of engineering data configuration. This addresses potential issues such as syntax errors, incomplete fault point coverage, and inaccurate configurations in the business scenario data, eliminating low-quality data and retaining high-value business scenario data to provide high-quality data support for model training.
[0056] The method described above acquires core scenario-based practical data, such as historical code generation cases, test case writing records, and engineering data configuration operation logs within the enterprise's signal domain. This ensures that the training data carries real business logic and practical details, avoids interference from generalized data, and allows the model to directly learn knowledge that aligns with the enterprise's actual business, thereby improving the model's adaptability to the enterprise's real business from the source. In addition, based on quantitative requirements standards such as code syntax compliance rate, coverage of core fault points of signal equipment, and accuracy of engineering data configuration, the collected data is precisely screened. This effectively eliminates low-quality data with syntax errors, incomplete fault point coverage, and inaccurate configuration, while retaining high-value data. This not only solves the pain point of unreliable business data quality but also provides high-quality data support for the model's scenario adaptation training, significantly improving the training effect and practical value of large-scale signal domain models.
[0057] In some embodiments, this application also provides a business processing method, as follows: Obtain pending business requirements in the signal domain, which include at least one of the following: code generation assistance requirements in the signal domain, test case generation requirements, and engineering data configuration requirements; The business requirements to be processed in the signal domain are input into the large signal domain model adapted to the scenario, and the business processing results are output. The large signal domain model adapted to the scenario is trained based on the above-mentioned large signal domain model training method.
[0058] Specifically, after training the model based on the signal domain general knowledge dataset and the signal domain business scenario dataset to obtain a large signal domain model, the business requirements to be processed in the signal domain can be input into the trained large signal domain model, thereby efficiently and accurately outputting business processing results, meeting the usage requirements under different business scenarios, significantly improving the efficiency and accuracy of enterprise business processing, and providing a cost-effective solution for the digital transformation of the signal domain.
[0059] For example, such as Figure 3 As shown, this application provides a method for training large models in the signal domain, and the specific process is as follows: For low-altitude and rail scenarios and signaling services, datasets such as signaling general knowledge text data and signaling general knowledge SFT are constructed through methods such as targeted web crawling, signaling service data processing, and data synthesis. Based on open-source large models, and deeply integrated with general data and knowledge of the urban rail transit industry, and addressing the production technology requirements of industry applications, a preliminary L1 vertical domain large model, namely a basic signaling domain large model, is constructed using methods such as pre-training of the L0 basic model. Further, the feasibility of AI models in application modes such as code-assisted generation, test case generation, and automatic configuration of engineering data is explored. Based on the L1 vertical domain model, a business L2 scenario model for the signaling domain is constructed using methods such as fine-tuning (LORA, full parameter fine-tuning, etc.) and RLFH, thus obtaining a scenario-adapted signaling domain large model.
[0060] For example, such as Figure 4 The present application provides a method for obtaining a general knowledge dataset in the field of signal processing, and the specific steps are as follows: General-purpose large-scale models often perform poorly and lack professionalism in solving vertical domain problems. This is due to factors such as the high cost of constructing vertical domain data, the scarcity of open-source data, and the increasing scarcity of specialized data. Furthermore, training data includes national standards, regulations, books, domain websites, and general corpora, requiring the construction of single-turn and multi-turn dialogue data. Pre-training requires significant computational resources and high dataset quality requirements, necessitating data cleaning strategies and related configuration tools, leading to a lack of high-quality data. Existing files contain a large amount of non-textual information such as figures and tables, requiring the filtering of low-quality text, such as machine translation and duplicate content. Creating high-quality datasets involves large-scale data cleaning; otherwise, it can severely impact model performance. This application improves the relevant data management system and compiles the "Guideline for the Construction of High-Quality Datasets in Urban Transportation" based on scenario requirements. This includes data classification standards, dataset element composition, dataset construction requirements, quality evaluation dimensions, and a full lifecycle guide covering data preparation, labeling, application, and security. The evaluation of high-quality datasets employs objective indicators (rule detection), subjective indicators (manual sampling detection), and application indicators (model performance). The application indicators utilize rapid testing with small models to assess performance. Optionally, the industry general knowledge dataset required for pre-training can be collected and organized from public channels, including professional literature such as papers and reports, and annotated by personnel with relevant industry experience; the industry specialized knowledge dataset used for fine-tuning comes from high-value scenario data provided by subsidiaries, including documents and drawings from internal organizations, and annotated by domain experts, thus achieving the accuracy of high-quality datasets, that is, achieving efficient and accurate acquisition of signal domain general knowledge datasets.
[0061] For example, such as Figure 5The present application provides a method for obtaining a dataset of signal domain service scenarios, and the specific steps are as follows: Enterprises identify high-value scenarios for efficiency improvement. After matching needs and producing raw datasets, these datasets are processed using rule bases and manually labeled to generate AI datasets. Quality assessment and feedback are essential processes for creating high-quality datasets, and this process is also the core area of vertical domain model building. The difference between high-quality datasets and high-value datasets—that is, the difference between general datasets in the signal processing domain and business scenario datasets in the signal processing domain—lies in their suitability for the application scenario and their ability to efficiently empower that scenario. Future large-scale model application development must start with application to unlock the potential of data and complete the enterprise's AI transformation.
[0062] Among the methods described above, the AI-enabled platform for the signal domain based on large language models aims to create high-quality datasets, algorithm libraries, and toolchains for "AI + Transportation," providing a technical foundation for building an intelligent integrated three-dimensional transportation network. It will become a core productivity tool for enterprise digital transformation, realizing a closed-loop logic of "data foundation building - scenario-driven - technology adaptation - application feedback." Through industry knowledge distillation, lightweight deployment, and ecosystem collaborative construction, it can effectively overcome obstacles to industrialization and complete digital transformation.
[0063] The signal domain large model training apparatus provided by the present invention is described below. The signal domain large model training apparatus described below can be referred to in correspondence with the signal domain large model training method described above. The signal domain large model training apparatus of the embodiments of this application is as follows: Figure 6 As shown, it includes: The acquisition module 610 is used to acquire the signal domain general knowledge dataset and the signal domain business scenario dataset; The first training module 620 is used to train the open-source general large model based on the signal domain general knowledge dataset to obtain a basic signal domain large model with basic cognitive ability in the signal domain. The second training module 630 trains the basic signal domain large model based on the signal domain business scenario dataset to obtain a scenario-adapted signal domain large model; the scenario-adapted signal domain large model is used for signal domain code auxiliary generation, test case generation, and engineering data configuration.
[0064] Figure 7This example illustrates the physical structure of an electronic device, which may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740. The processor 710, communications interface 720, and memory 730 communicate with each other via the communication bus 740. The processor 710 can invoke logical instructions in the memory 730 to execute a signal domain large model training method. This method includes: acquiring a signal domain general knowledge dataset and a signal domain business scenario dataset; training an open-source general large model based on the signal domain general knowledge dataset to obtain a basic signal domain large model with fundamental signal domain cognitive capabilities; training the basic signal domain large model based on the signal domain business scenario dataset to obtain a scenario-adapted signal domain large model; and using the scenario-adapted signal domain large model for signal domain code auxiliary generation, test case generation, and engineering data configuration.
[0065] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0066] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the signal domain large model training method provided by the above methods. The method includes: acquiring a signal domain general knowledge dataset and a signal domain business scenario dataset; training an open-source general large model based on the signal domain general knowledge dataset to obtain a basic signal domain large model with basic signal domain cognitive capabilities; training the basic signal domain large model based on the signal domain business scenario dataset to obtain a scenario-adapted signal domain large model; the scenario-adapted signal domain large model is used for signal domain code auxiliary generation, test case generation, and engineering data configuration.
[0067] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements a signal domain large model training method provided by the above methods. This method includes: acquiring a signal domain general knowledge dataset and a signal domain business scenario dataset; training an open-source general large model based on the signal domain general knowledge dataset to obtain a basic signal domain large model with basic signal domain cognitive capabilities; training the basic signal domain large model based on the signal domain business scenario dataset to obtain a scenario-adapted signal domain large model; the scenario-adapted signal domain large model is used for signal domain code auxiliary generation, test case generation, and engineering data configuration.
[0068] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0069] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for training large-scale models in the signal domain, characterized in that, include: Acquire general knowledge datasets and business scenario datasets in the signal domain; The open-source general large model is trained based on the aforementioned signal domain general knowledge dataset to obtain a basic signal domain large model with basic cognitive capabilities in the signal domain; The basic signal domain large model is trained based on the signal domain business scenario dataset to obtain a scenario-adapted signal domain large model. The large signal domain model adapted to the scenario is used for signal domain code generation, test case generation, and engineering data configuration.
2. The signal domain large model training method according to claim 1, characterized in that, The acquisition of the signal domain general knowledge dataset includes: We acquire professional literature, technical standards, and public business data in the field of urban rail transit signaling, and clean and label them to obtain an initial general knowledge dataset in the field of signaling. Based on the initial signal domain general knowledge dataset, a preset lightweight model is trained, and the signal domain general knowledge dataset is obtained based on the accuracy of the trained lightweight model in the basic question answering task in the signal domain.
3. The signal domain large model training method according to claim 2, characterized in that, The general knowledge dataset for the signal domain is obtained based on the accuracy of the trained lightweight model in a basic question-answering task in the signal domain, including: If the accuracy of the lightweight model in the basic question-answering task in the signal domain is greater than or equal to a threshold, the initial signal domain general knowledge dataset will be used as the signal domain general knowledge dataset.
4. The signal domain large model training method according to claim 2, characterized in that, The general knowledge dataset for the signal domain is obtained based on the accuracy of the trained lightweight model in a basic question-answering task in the signal domain, including: If the accuracy of the lightweight model in the basic question-answering task of the signal domain is less than a threshold, the initial signal domain general knowledge dataset is supplemented and corrected based on the erroneous question-answering results of the lightweight model in the basic question-answering task of the signal domain. The training and validation process of the lightweight model is repeated until the accuracy of the lightweight model in the basic question-answering task of the signal domain is greater than or equal to the threshold. The supplemented and corrected initial signal domain general knowledge dataset is then used as the signal domain general knowledge dataset.
5. The signal domain large model training method according to any one of claims 1-4, characterized in that, The signal domain service scenario dataset was obtained using the following method: Acquire internal business scenario data of the enterprise; the business scenario data includes at least one of the following: historical cases of code generation in the signal domain, test case writing records, and practical logs of engineering data configuration; The business scenario data is matched with the preset high-value business scenario requirement standards to obtain the signal domain business scenario dataset; the preset high-value business scenario requirement standards include at least one of the following: code syntax compliance rate, coverage of core fault points of signal equipment in test cases, and accuracy of engineering data configuration.
6. A business processing method, characterized in that, include: Obtain unprocessed business requirements in the signal domain, which include at least one of the following: code generation assistance requirements in the signal domain, test case generation requirements, and engineering data configuration requirements; The signal domain business requirements to be processed are input into a scenario-adapted large signal domain model, and the business processing results are output; the scenario-adapted large signal domain model is trained based on the large signal domain model training method of any one of claims 1-5.
7. A large-scale model training device for the signal domain, characterized in that, include: The acquisition module is used to acquire general knowledge datasets and business scenario datasets in the signal domain. The first training module is used to train the open-source general large model based on the signal domain general knowledge dataset to obtain a basic signal domain large model with basic cognitive ability in the signal domain. The second training module trains the basic signal domain large model based on the signal domain business scenario dataset to obtain a scenario-adapted signal domain large model. The large signal domain model adapted to the scenario is used for signal domain code generation, test case generation, and engineering data configuration.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the signal domain large model training method as described in any one of claims 1 to 5, or the service processing method as described in claim 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the signal domain large model training method as described in any one of claims 1 to 5, or the service processing method as described in claim 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the signal domain large model training method as described in any one of claims 1 to 5, or the service processing method as described in claim 6.