Railway traffic field semantic recognition method and device based on few sample data and medium

By preprocessing railway data and performing multi-scale model legitimacy judgment, and combining large language models and RAG architecture for semantic rewriting, the problem of high-precision speech semantic recognition under limited sample data in railway scheduling was solved, achieving high accuracy and structured output of railway scheduling instructions.

CN121747577APending Publication Date: 2026-03-27CASCO SIGNAL LTD

Patent Information

Application Number
CN202512039071.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In railway dispatching and traffic safety monitoring, existing technologies struggle to achieve high-precision, robust, and easily maintainable speech and semantic recognition with limited sample data, especially in the recognition and understanding of professional dispatching instructions, where there are problems such as high error rates and high maintenance costs.

Method used

By preprocessing railway data, constructing a dedicated dictionary and rule base, fine-tuning the ASR model in conjunction with enhanced speech, and using a multi-scale model to perform legality judgment and semantic rewriting on the initial ASR transcribed text, including sentence-level, local word order-level, and semantic tag-level multi-scale detection, and combining a large language model and RAG architecture for domain semantic rewriting, the final output is a structured instruction text.

Benefits of technology

It improves the transcription accuracy and semantic recognition professionalism of railway dispatch voice commands, adapts to the business scenario needs of the railway transportation field, reduces the cost of manual annotation and rule maintenance, and enhances the comprehensiveness and accuracy of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121747577A_ABST
    Figure CN121747577A_ABST
Patent Text Reader

Abstract

The invention relates to a railway traffic field semantic recognition method and device based on few sample data and a medium, and the method comprises the steps: carrying out the preprocessing of the obtained railway traffic field data, and constructing a field dictionary and a rule base; converting the preprocessed domain data into enhanced voice, forming a training pair by the enhanced voice and a corresponding text, and performing domain fine tuning on the ASR model; inputting a voice instruction in the railway traffic field into the field fine-tuned ASR model, and outputting an ASR initial transcription text; on the three levels of sentence level, local word order level and semantic tag level, word legality judgment is performed on the ASR initial transliteration text through a multi-scale model, and a high-credibility anchor point set and uncertain content are recognized; and field semantic rewriting is carried out based on the legitimacy judgment result of the words, and finally a structured instruction text is output. Compared with the prior art, the method has the advantages of realizing accuracy and adaptability of semantic recognition in the field of railway traffic and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of railway traffic, and in particular to a railway traffic field semantic recognition method based on few-shot data, equipment and medium. BACKGROUND

[0002] In core businesses such as railway dispatching command and train operation safety monitoring, voice is the most direct and efficient communication method, but general voice semantic recognition technology faces multiple challenges in this scenario: first, high-quality field data is scarce, and it is difficult to collect on a large scale due to safety and privacy restrictions, and the transcription and annotation cost is high and the cycle is long; second, the railway field terminology is highly professional and has multiple fixed expressions, and general models lack relevant knowledge, resulting in high recognition error rate; third, dispatching and train operation instructions have very low tolerance for station names, track numbers, time, and numbers, and any keyword or number error can cause a safety accident; fourth, railway regulations and business rules change continuously with line equipment upgrades, train operation diagram adjustments, etc., and if the system knowledge cannot be updated quickly, rule lag may occur, affecting long-term stability and applicability.

[0003] Existing solutions also have obvious deficiencies: traditional statistical and deep learning methods that rely on large-scale labeled corpus data are difficult to ensure both accuracy and generalization ability under data scarcity conditions; directly using or slightly adjusting general large language models can easily produce illegal or illogical "hallucinations" on professional dispatching instructions, and even damage their general capabilities; keyword and template matching based on artificial rules may work well in fixed language scenarios, but it is difficult to deal with spoken language expressions, cross-sentence references, and new terminology, and the maintenance cost is high and the scalability is poor. In summary, current technology is still difficult to build a railway voice semantic recognition system that is both highly accurate and robust, and easy to maintain under the constraints of "few-shot field data, high safety requirements, and dynamic rule updates". How to significantly reduce the cost of artificial annotation and rule maintenance while accurately recognizing and understanding railway professional voice instructions has become a key problem that needs to be solved.

[0004] Through retrieval, it is found that Chinese invention patent application publication No. CN111144127A proposes a text semantic recognition method, a model acquisition method and related devices. The text semantic recognition model is obtained by training a preset neural network using multiple training texts annotated with text semantics and their keyword sequences. Chinese invention patent application publication No. CN115983282A proposes an efficient small sample dialogue semantic understanding method based on prompts for small sample dialogue semantic understanding prediction. The present application provides an efficient small sample dialogue semantic understanding method based on prompts, which predicts slot values by stating slot types in prompts, reduces the required forward propagation times of prediction, and improves the efficiency of the model. Chinese invention patent application publication No. CN117292680A discloses a power transmission inspection voice recognition method based on small sample synthesis. A large number of power transmission inspection professional text corpora are extracted to form a power transmission inspection professional semantic recognition model, and the acoustic models of multiple power transmission inspection personnel are established by sampling and voice input to realize accurate recognition of professional voice data of power transmission inspection.

[0005] In summary, although existing patents and academic solutions have made achievements in small sample synthesis, prompt understanding and text semantic mapping, they often do not focus on the railway transportation field or rely on high-quality data samples. In the context of high safety threshold, few samples, dynamic rules and real-time requirements of railway dispatching, there is still a lack of an overall solution that can balance high precision, strong controllability, strong universality, high scalability and low maintenance cost. In order to truly meet the landing needs of railway business, in the case of few samples, it is necessary to improve the accuracy and efficiency of semantic recognition, which is a technical problem to be solved.

[0006] A search revealed Chinese invention patent application publication number CN117875304A, which discloses a method, system, and storage medium for constructing a corpus for the subway domain. The method includes the following steps: collecting subway-related speech data; preprocessing the subway-related speech data by cleaning, labeling, and segmenting; converting the preprocessed subway-related speech data into text, and adding the converted text data to AISHELL's proprietary speech library to form an AISHELL dataset; training a speech recognition model using data from the AISHELL dataset, and using the trained speech recognition model to recognize audio files integrated from the AISHELL dataset to generate text data; constructing a business keyword matching model and a label classification model for the subway domain; structuring the text data generated by the speech recognition model, and classifying the structured data according to the business keyword matching model and the label classification model to obtain a subway-related corpus. This existing patent application lacks a legality check for the converted text, resulting in low speech recognition accuracy.

[0007] Achieving high-precision recognition and robust semantic understanding in the railway transportation field under limited sample conditions has become a technical problem that needs to be solved. Summary of the Invention

[0008] The purpose of this invention is to overcome the shortcomings of the existing technology and provide a semantic recognition method, device and medium for railway transportation based on few sample data.

[0009] The objective of this invention can be achieved through the following technical solutions: According to one aspect of the present invention, a semantic recognition method for the railway transportation domain based on few-sample data is provided, the method comprising: The acquired railway transportation data is preprocessed to construct a domain dictionary and rule base; The preprocessed domain data is converted into enhanced speech, and the enhanced speech and corresponding text are used to form training pairs to fine-tune the ASR model in the domain. Input the voice commands from the railway transportation field into the domain-fine-tuned ASR model, and output the initial ASR transcribed text. At the sentence level, local word order level, and semantic tag level, a multi-scale model is used to determine the word legitimacy of the initial ASR transcribed text, and identify the set of high-confidence anchor points and uncertain content. Domain semantic rewriting is performed based on the word legality judgment results, and finally a structured instruction text is output.

[0010] As a preferred technical solution, the word legitimacy judgment of the initial ASR transcribed text is performed through a multi-scale model, including: performing perplexity analysis and detection at the sentence level and local word order level, and outputting a set of high-confidence anchor points and a set of uncertain content.

[0011] As a preferred technical solution, word legality determination of the initial ASR transcribed text using a multi-scale model also includes: Local legality is determined based on n-grams, and the set of words that are illegal in the context is output. Based on CRF, sequence validity annotation is performed, and a set of candidate error words is output. Based on the rule base and domain dictionary, the results judged as legitimate by CRF are subjected to secondary verification to identify a set of suspicious words; By integrating the set of invalid words in the context, the set of incorrectly selected words, and the set of suspicious words, a total set of candidate incorrect words is obtained.

[0012] As a preferred technical solution, the CRF-based sequence validity annotation specifically involves: when words If a word is deemed invalid, or if the confidence level of its best-case legal category label is below a threshold, then the word will be removed from the list. Add the candidate error word set from the CRF perspective.

[0013] As a preferred technical solution, the secondary verification of the CRF-determined results specifically involves: based on the constructed domain dictionary and rule base, for the words to be detected... With dictionary entries Define a string or phonetic similarity function, if words If a word belongs to the set of suspicious words or the set of words with invalid context, and the similarity function falls between the defined lower and upper similarity limits, then the word will be considered a suspicious word. Add a set of suspicious words from a rule-based perspective.

[0014] As a preferred technical solution, the domain semantic rewriting based on word legality judgment results includes: using high-confidence anchor points as core search keywords for domain knowledge retrieval, introducing RAG architecture and large language model, and performing locally controlled semantic rewriting of uncertain content.

[0015] As a preferred technical solution, using high-reliability anchor points as core search keywords for domain knowledge retrieval includes: High-confidence anchors and their sentence context are encoded to obtain vectorized representations, and a knowledge base for the railway transportation domain is constructed. Search query construction and execution: High-confidence anchors are used as core search keywords, and combined with the local context information of the sentences containing the high-confidence anchors, a search vector is constructed. The top K documents with the highest similarity are selected to form a candidate knowledge set.

[0016] As a preferred technical solution, local semantic rewriting of uncertain content with local control includes: inputting the initial ASR transcription text, uncertain content and the location of candidate error words and their context windows, a set of high-confidence anchor points and corresponding business entity information, and a set of candidate knowledge into a large language model; setting explicit rewriting constraint rules in the prompt words of the large language model based on rewriting constraint rules; and outputting the final instruction text for local correction.

[0017] As a preferred technical solution, the rewrite constraint rules include: (1) Correction scope constraint: Only local corrections are allowed for content marked as uncertain or candidate error words; (2) Semantic compliance constraints: The rewriting results must meet the constraints of the retrieved candidate knowledge set; (3) Terminology adaptation constraint: Under the premise that the correction result meets the semantic correctness of railway traffic business, legal professional terms that are similar to the original recognition result in terms of pronunciation or character shape shall be given priority.

[0018] As a preferred technical solution, the speech similarity between the final instruction text and the initial ASR transcription text should not exceed the maximum allowable amount of modification.

[0019] As a preferred technical solution, preprocessing the acquired railway transportation data includes: Collect text data related to train dispatching to form the original text set in the railway transportation field; Each original collection of railway transportation texts was sequentially segmented into sentences and words. Each word is tagged with part-of-speech and grammatical information to obtain a tagged word sequence; The key fields of the tagged word sequences are normalized to obtain a normalized sentence set.

[0020] According to another aspect of the present invention, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement the method described thereon.

[0021] According to a third aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described thereon.

[0022] Compared with the prior art, the present invention has the following beneficial effects: 1) This invention is a semantic recognition method for the railway transportation field based on few sample data. It constructs a dedicated dictionary rule base by preprocessing railway field data, combines enhanced speech to achieve domain fine-tuning of the ASR model, and then uses a multi-scale model to perform legality judgment and semantic rewriting on the initial ASR transcribed text. This not only solves the problem of poor adaptability of ASR models under few sample data in the railway field, but also accurately identifies the reasonable and uncertain content of the transcribed text, and finally outputs structured instruction text. This effectively improves the accuracy, professionalism and structuring of railway dispatch voice instruction transcription and semantic recognition, and adapts to the business scenario requirements of the railway transportation field.

[0023] 2) The multi-scale model of this invention uses sentence-level and local word order-level perplexity analysis for detection. By measuring the perplexity of the text, it can efficiently screen out highly credible text anchors to ensure the reliability of the core content, and accurately locate uncertain content, providing a clear direction for subsequent semantic rewriting. This improves the pertinence and efficiency of the legality judgment of the transcribed text and reduces the resource consumption of invalid analysis.

[0024] 3) The legality judgment process of the multi-scale model in this invention uses multi-layer verification of n-gram, CRF model and rule dictionary to identify problematic words from three dimensions: contextual rationality, sequence labeling and domain rule matching. Then, the multi-dimensional results are integrated to obtain a candidate error word set, realizing the double backstop of statistical model and domain rules. It not only covers different types of text error scenarios, but also makes up for the judgment blind spots of single model. It significantly improves the comprehensiveness and accuracy of error recognition in railway transcribed text, and provides a more complete error reference for subsequent semantic rewriting.

[0025] 4) This invention is based on high-confidence anchors and precise retrieval constraints for uncertain content. It not only uses high-confidence anchors to accurately retrieve domain knowledge, solving the problem of accurate matching of railway professional information, but also introduces a domain knowledge base through the RAG architecture to enhance the professional semantic capabilities of the large language model, avoiding the domain knowledge bias of the general model. At the same time, it corrects uncertain content through local controlled rewriting, which not only ensures the stability of the core information of the original text, but also combines business entity information and knowledge sets to generate compliant railway instruction text. Attached Figure Description

[0026] Figure 1 This is a flowchart illustrating the semantic recognition method in the railway transportation field in an embodiment of the present invention; Figure 2 This is a knowledge distribution structure diagram in the field of railway transportation in an embodiment of the present invention; Figure 3 This is a schematic diagram of the semantic recognition system in the railway transportation field in an embodiment of the present invention; Figure 4This is a detailed flowchart illustrating the semantic recognition process in the railway transportation field in an embodiment of the present invention. Detailed Implementation

[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0028] This embodiment relates to a semantic recognition method for railway transportation based on few sample data. Through a multi-stage pipeline, it organically combines domain data synthesis, acoustic model adaptation, multi-scale legality judgment, large language model and retrieval enhancement generation, and templated structured output, and completes highly reliable semantic recognition of railway voice commands under the premise of limited domain knowledge and safety regulations.

[0029] like Figure 1 As shown, the method includes the following steps: Step S1: Construct and preprocess knowledge training data in the field of railway transportation; Step S2: Based on knowledge in the railway transportation field, data is synthesized, and the synthesized data is used to train a speech recognition model (ASR: Automatic Speech Recognition), in which the ASR model automatically converts audio into corresponding text content; Step S3: Word legitimacy determination based on a multi-scale model; Step S4: Domain semantic rewriting combining Large Language Model (LLM) and RAG (Retrieval-Augmented Generation); Step S5: Generate formatting instructions based on RAG.

[0030] In step S1, the system first constructs a training corpus based on texts from railway dispatching orders, traffic regulations, and historical traffic records. The domain texts undergo word segmentation, part-of-speech tagging, and regularization preprocessing to filter out non-standard terms verified manually or by rules, resulting in a set of standardized statements within the railway transportation domain. This set serves as the training sample for subsequent statistical models, forming the data foundation for all subsequent steps. Specifically, this includes: Step S1-1, Collect raw corpus: Collect text data related to train dispatching from existing railway information systems and document resources to form the original railway transportation domain text set. Let the original railway transportation domain text set be... Each of them A single original railway traffic domain text (hereinafter referred to as a domain text) includes railway dispatch orders, traffic regulations, or historical traffic records.

[0031] Step S1-2, Sentence segmentation and word segmentation processing: For each original railway transportation domain text Segment by sentence to obtain a set of sentences. For each sentence Perform word segmentation to obtain a word sequence. Where N is the total number of sentences, L j For the j-th sentence The number of word sequences.

[0032] Steps S1-3, Part-of-Speech Tagging and Syntax Information Annotation: Using existing Chinese word segmentation and part-of-speech tagging tools, or a self-trained annotation model based on railway corpus, for each word... Assigning part-of-speech tags This yields a labeled word sequence.

[0033] Steps S1-4, Regularization and Standardized Expression Filtering: The word sequence involving key fields such as train number, time, and track number is normalized. Let the normalization function be... Then the normalized words are: The normalized sentence set is .

[0034] Steps S1-5: Domain dictionary and rule base construction: Based on the standardized statements, construct the domain dictionary, including: station name dictionary, track name dictionary, train number pattern library and other railway domain entity dictionaries, and combine it with the description of command format in the train operation regulations to form the initial rule base. Finally, it is incorporated into the task-level domain knowledge base through metadata. This part corresponds to... Figure 2 The left-side branch structure.

[0035] In step S2, firstly, based on the collected normalized sentence set in the railway transportation field, synthetic speech data containing professional terminology and instructional phrases is synthesized using text-to-speech conversion technology. During this process, to make the generated data closer to the real-world distribution, noise information including track vibration noise, equipment noise, residual noise from in-vehicle broadcasts, and crowd noise is injected. Subsequently, based on the synthesized speech data, the speech recognition model is adaptively fine-tuned for the railway transportation field. This step enables the ASR model to initially possess acoustic perception capabilities for railway transportation professional vocabulary, providing higher-quality initial recognition results for subsequent processing.

[0036] Includes the following steps: Step S2-1, text subset selection, from Select a subset of business-related technical terms and typical instruction phrases: , Step S2-2, Text-to-Speech Synthesis: The system synthesizes each sentence from a subset of the text. Convert to speech waveform .

[0037] Steps S2-3, Noise Injection and Acoustic Enhancement: Considering the acoustic environment of the actual dispatch console, station, and driver's cab, this invention employs a multi-source noise superposition model to achieve typical environmental noise injection for synthesized speech. The multi-source noise superposition model includes, but is not limited to, the following noises: Track vibration noise 50–200 Hz bandpass noise, center frequency relative to train speed The correlation can be modeled as follows: , , in, (·) represents a bandpass filter. The center frequency is noise signal, For train speed, , These are the relevant linear parameters.

[0038] Equipment noise It contains a low-frequency broadband band mainly consisting of 50–500 Hz and mechanical harmonics, which are modeled as noisy sinusoidal noise in this invention. With low-frequency pink noise Superimposed form: , , , in The number of sine components. For the first The frequency of a sinusoidal component , The average amplitude and phase of the corresponding components, respectively. and These are random processes with amplitude noise and phase noise, respectively. Zero-mean white noise, Impulse response of a linear time-invariant filter that achieves pink noise characteristics.

[0039] Residual noise from in-car radio The speech signal is roughly in the same frequency band as the target speech signal, but with low amplitude and low clarity. Residual noise from the in-car radio. With the target speech signal Meet the target signal-to-noise ratio .

[0040] Crowd noise Low-intelligibility human voice mixed noise at 300–3000 Hz.

[0041] Therefore, the total noise expression is: , in, The above types of noise are represented by adjustment coefficients. It can simulate noise characteristics in different scenarios.

[0042] Noise is added to each synthesized speech to obtain enhanced speech. : , Among them, coefficient , Determined based on target SNR and noise ratio. At this point, the training pairs are obtained: , in, This is the k-th sentence.

[0043] Step S2-4: Domain fine-tuning of the ASR model. A general Chinese speech recognition model is selected as the baseline model, and the enhanced speech and its corresponding text are used as training data to perform domain-adaptive fine-tuning.

[0044] Let the ASR model parameters be Acoustic feature sequence Mapping to word sequences .

[0045] Define the following training loss function: , Where K is the number of training samples; L k The length of the word sequence corresponding to the k-th sample; For the i-th word in the k-th sample, This refers to the context preceding the i-th word in the k-th sample.

[0046] The fine-tuned ASR model has a higher recognition accuracy for railway-related terms, and outputs the initial recognition results: ,in This represents the initial recognition result for the T-th word.

[0047] Step S3: Considering that large amounts of training data are often unavailable in actual engineering projects, this invention uses a multi-scale model for word legitimacy discrimination to achieve efficient and accurate training semantic legitimacy discrimination under such circumstances. To effectively review and locate errors in the ASR model output under limited sample conditions, this invention performs multi-scale legitimacy detection at three levels: sentence level, local word order level, and semantic label level. A word legitimacy discrimination module is constructed using multiple lightweight models and rules, including: a perplexity analysis and detection submodule; an ngram-based local legitimacy discrimination submodule; a CRF-based sequence legitimacy annotation submodule; and a rule- and dictionary-based auxiliary checking submodule. Figure 4 Specifically, it includes: Step S3-1, Perplexity Analysis and Detection: Calculate sentence-level and word-level perplexity for the initial ASR transcription results. High-confidence words and phrases are selected as candidates for "high-confidence anchors"; words with perplexity significantly higher than the sentence average are considered as uncertain content. Sentence-level perplexity is defined as follows: , Where s represents a sentence, PPL(s) is the perplexity of sentence s, and T is the number of words in sentence s, i.e., the sentence length; The word at the t-th position in the sentence; The model predicts words based on the context preceding the t-th position in the sentence. The probability of.

[0048] Define word-level local perplexity score : , Let the average local score of the whole sentence be for: , Given threshold Make a judgment: like Then Marked as "Uncertain Content"; like Then Marked as a candidate for "high credibility anchor".

[0049] Finally, a set of highly reliable anchor points is obtained. and a collection of uncertain content : .

[0050] Sentence-level perplexity (PPL) is used to judge the fluency and rationality of the entire sentence (the higher the PPL, the less fluent the sentence); word-level local perplexity (PPL) is used to locate which specific word in the sentence causes the illogicality. The higher the score, the more likely the word is to be an outlier. First, look at the whole sentence (PPL), then break it down into parts (word-level local perplexity score), together to achieve sentence plausibility assessment and outlier word location.

[0051] Step S3-2, Local Legality Judgment Based on n-grams: Based on the standardized sentences constructed in step S1, an n-order n-gram language model is trained to estimate the conditional probability of a word appearing in a given context and to identify extremely low-frequency word orders. For vocabulary... Given a relative threshold or absolute threshold If one of the following two conditions is met: , in, For vocabulary The conditional probability of occurrence in a given context; The average conditional probability of the context.

[0052] That is, if vocabulary If the conditional probability of a word appearing in a given context is much lower than the average conditional probability or lower than the absolute threshold, then the word will be... These are labeled as "contextually invalid words" from an n-gram perspective, forming a set of contextually invalid words: .

[0053] Step S3-3: Based on CRF (Conditional Random Field, a sequence labeling model), the valid statements are labeled as a sequence of tags containing train numbers, station names, operational verbs, and illegal / suspicious words. This is used to train the CRF model. During detection, when a certain word... Labeled (Illegal), or the confidence level of its best legal category label is below the threshold. , , Then add it to the candidate error word set from the CRF perspective: .

[0054] Steps S3-4 involve auxiliary checks based on a rule base and domain dictionary. This invention utilizes a constructed domain dictionary and rule base to perform secondary verification on results deemed valid by CRF. For words to be detected... With dictionary entries Define a string or phonetic similarity function. If a word belongs to a set However, there are certain typical matching scenarios. Taking the station dictionary as an example, such as... Make: , in, , These are the lower and upper limits of similarity, respectively. Then add the word to the set of suspicious words from the rule perspective:

[0055] Step S3-5: Integrate the four categories of discrimination results from steps S3-1 to S3-4 to form a set of high-confidence anchor points. Uncertainty content collection With the total set of candidate error words .

[0056] After obtaining the judgment result, the system will perform domain knowledge retrieval based on the set of high-confidence anchor points, as well as error correction and rewriting of low-confidence content (including the set of uncertain content and the set of candidate error words), and generate results based on the rewritten content, corresponding to steps S4 and S5.

[0057] Step S4: Domain semantic rewriting combining large language models and RAG: Building upon the multi-scale discrimination in step S3, and based on high-confidence anchors, semantic approximation constraints, and historical railway traffic information, uncertain content and illegal words are precisely corrected. Using high-confidence anchors as constraints for retrieval and rewriting, a RAG architecture and a large language model are introduced to perform controlled local semantic rewriting of uncertain content, and results are generated based on the rewritten content. For example... Figure 4 ,include: Step S4-1: Domain knowledge retrieval based on high-confidence anchor points; Step S4-2: Domain semantic rewriting based on a large language model.

[0058] Step S4-1 includes the following steps: Step S4-1-1: Encode the high-confidence anchor points and their sentence context to obtain vectorized representations, and construct a knowledge base for the railway transportation field, including but not limited to railway dispatching regulations, train operation organization methods and operation instructions, standard dispatching command templates, typical terms and historical dispatching cases and their corresponding structured records.

[0059] Step S4-1-2, Query Construction and Execution: The high-confidence anchor points integrated in Step S3-5 are used as core search keywords, and combined with the local context information of the current sentence, a search vector is constructed. Compared with traditional methods that search based on the entire sentence text, this invention significantly reduces search bias by using only "anchor points" consistently confirmed by multiple models. Through different similarity search methods, the top similarity values ​​are selected... Large documents form a candidate knowledge set: , in, For the Kth candidate documents, sorted from high to low similarity, This is the candidate knowledge set.

[0060] Step S4-2 includes the following steps: Step S4-2-1, Input Information Construction: Organize the following information into the input of the large language model: (1) Initial ASR transcription text ; (2) The location of uncertain content and candidate error words and their context window; (3) A set of high-credibility anchor points and corresponding business entity information; (4) The set of candidate knowledge retrieved from the knowledge base (i.e., relevant regulations and transfer order templates).

[0061] Step S4-2-2: Rewrite the constraint rules and explicitly constrain them in the prompt words of the large language model: (1) Only local corrections are allowed for text segments marked as “uncertain content / candidate error words”; other text segments and high-confidence anchors must not be changed. (2) The rewriting results must meet the semantic constraints of the retrieved candidate knowledge set and must not generate content that violates the procedures or is logically inconsistent; (3) Under the premise that multiple candidate words meet the business semantic correctness, priority should be given to selecting legal terms that are similar to the original recognition results in terms of pronunciation or character shape, so as to maintain the correspondence with the original speech as much as possible.

[0062] Step S4-2-3: Output the corrected instruction text. Outside of the constraints mentioned above, the large language model performs locally controlled rewriting of the original sentence. This invention includes, but is not limited to, the use of formal speech similarity constraints. in, The maximum amount of modification allowed; This is the revised instruction text; The initial transcribed text for ASR is provided; Dist is used to measure the distance between the two speech similarities. In other words, the LLM generation result must meet the above constraints to be considered a valid output. Afterwards, the final instruction text with correct semantics and standardized terminology is obtained, expressed by the following formula: , in This is an LLM model.

[0063] The retrieval process selects the most relevant knowledge fragments to the current instruction context using a similarity metric to form a candidate knowledge set. Because the retrieval uses a combination of "high-confidence anchors + instruction context," it reduces the impact of noise compared to traditional whole-sentence retrieval. Furthermore, the rewritten result must possess the following characteristics: 1) Maintain a certain degree of similarity to the original expression at the phonetic level; 2) Strictly adhere to the retrieved candidate knowledge set at the semantic level of railway traffic operations to ensure clear meaning and accurate terminology in instructions; 3) The entity information corresponding to the high credibility anchors in the original sentence is preserved and not tampered with.

[0064] In step S5, after the previous steps, semantically correct instruction text has been obtained. This step is responsible for converting the instruction text into machine-readable structured instructions. A domain-fine-tuned lightweight entity extraction model or rule system is invoked to locate and extract corresponding entities from the final instruction text based on the key field list of the retrieved dispatch template (such as instruction type, train number, station name, track number, time, etc.). Furthermore, the final instruction text, the selected dispatch template description, and field descriptions are input as prompts into the large language model, ultimately generating a structured output for downstream systems to directly execute or record. Figure 4 Specifically, it includes: Step S5-1, Construction and Retrieval of Dispatch Order Template Vector Library: Different types of dispatch order templates are pre-defined according to railway dispatching operations and vectorized to form a template vector library. During the detection process, the final instruction text is first matched with the template vector library to lock the major instruction types, that is, to lock the instruction type corresponding to this instruction.

[0065] Step S5-2, under template constraints, parameter extraction and mapping: based on the list of key fields defined in the retrieved dispatch order template (such as train number, time, track number, etc.), locate and extract the corresponding entities one by one in the final instruction text to obtain the output that conforms to the structural constraints of the dispatch order template.

[0066] Through the above steps, natural language commands can be automatically converted into standardized data structures, which facilitates integration with existing dispatch automation systems and vehicle monitoring systems, enabling automatic verification, automatic recording, and partial automatic execution.

[0067] This embodiment also relates to a semantic recognition system for the railway transportation field based on few-sample data, such as... Figure 3 The system includes: Corpus Construction and Preprocessing Module: The system first constructs a corpus based on texts from fields such as railway dispatching orders, train operation regulations, and historical train operation records. For each field's text, preprocessing is performed including word segmentation, part-of-speech tagging, and regularization. Non-standard terms, verified manually or by rules, are filtered out, resulting in a set of "standard statements" within the field, which serve as training samples for subsequent statistical models. This step forms the data foundation for all subsequent steps.

[0068] The domain-knowledge-based data synthesis module synthesizes speech data containing technical terms and command phrases based on collected professional text corpora in the railway field using text-to-speech conversion technology. During this process, to make the generated data more closely resemble real-world distributions, noise information including track vibration noise, equipment noise, residual noise from in-vehicle broadcasts, and crowd noise is injected.

[0069] Speech Recognition Model: The Automatic Speech (ASR) model was fine-tuned for domain adaptation based on synthesized speech data. This enabled the ASR model to initially acquire acoustic perception capabilities for railway-related terminology, providing higher-quality initial recognition results for subsequent processing.

[0070] A word legitimacy discrimination module based on a multi-scale model: Considering that large amounts of training data are often unavailable in real-world engineering projects, this solution employs a multi-statistical model fusion architecture to construct a word legitimacy discrimination module for efficient and accurate training of semantic legitimacy discrimination under such circumstances. It performs multi-scale legitimacy detection at three levels: sentence level, local word order level, and semantic label level. This module mainly includes the following four types of modules: a perplexity analysis and detection submodule; an ngram-based local legitimacy discrimination submodule; a CRF-based sequence legitimacy annotation submodule; and a rule-based and dictionary-assisted checking submodule.

[0071] The domain semantic rewriting module, combining LLM and RAG, leverages the detection results of the preceding module to precisely correct "uncertain content" and "illegal words" based on "high-confidence anchors," semantic approximation constraints, and historical railway traffic information. The retrieval submodule selects the most relevant knowledge fragments to the current instruction context using similarity metrics to form a candidate knowledge set. Because the retrieval uses a combination of "high-confidence anchors + instruction context," it reduces noise impact compared to traditional whole-sentence retrieval. Furthermore, the rewritten result must possess the following characteristics: 1) Maintain a certain degree of similarity to the original expression at the phonetic level; 2) Strictly adhere to the retrieved regulations and templates at the semantic level of railway operations to ensure that instructions are clear and terminology is accurate; 3) The entity information corresponding to the high credibility anchors in the original sentence is preserved and not tampered with.

[0072] The RAG-based formatted instruction generation module transforms semantically correct instruction text into machine-readable structured instructions. It invokes a domain-tuned lightweight entity extraction model or rule system to locate and extract relevant entities from the text based on a list of key fields from the retrieved dispatch template (e.g., instruction type, train number, station name, track number, time). Furthermore, the final instruction text, the selected dispatch template description, and field specifications are fed into the large language model as prompts, ultimately generating a structured output for downstream systems to execute or record.

[0073] Figure 2 This diagram illustrates the knowledge distribution structure in the railway transportation domain of this invention, explaining how domain data drives the various modules within the system. The rule and dictionary extension interfaces indicate that the system can continuously access and expand with new domain rules and dictionary content, such as adding stations, adjusting routes, or implementing new operating procedures. The expanded rules and dictionary can provide more targeted and up-to-date references for subsequent auxiliary inspection modules, ensuring the system has good adaptability and scalability to business changes.

[0074] Figure 4 This is a framework diagram of the semantic recognition system in this invention, illustrating the process from user input, speech recognition, legality judgment, intelligent rewriting and error correction to the large model outputting structured text. First, the user speaks railway instructions in scenarios such as the dispatch console / driver's cab. The system then calls a domain-adjusted speech recognition model, which can recognize train numbers, station names, track numbers, dispatching terms, etc., and outputs the initial ASR transcribed text.

[0075] After this, the initial ASR transcribed text enters Figure 4 The three discriminant sub-modules and the fusion module in the middle perform a word legitimacy discrimination module based on a multi-scale model to judge the legitimacy of each word and each segment, identify highly credible content and potential errors, and provide precise positioning for subsequent LLM rewriting.

[0076] The implementation steps of the word validity determination module based on the local validity determination submodule of n-gram include: 3-2-1, n-gram language model training Using a normalized sentence set Train an nth-order language model for any sentence The n-gram model gives: in, , For the current word In the context of the first n-1 words Probability under given conditions; This represents the number of times the n-1-th word and the current word, forming an n-gram sequence, appear in the training data. This represents the number of times an n-1 gram sequence consisting of the first n-1 words appears in the training data.

[0077] 3-2-2: Calculation of Local Logarithmic Probability during the Detection Phase Sentences to be tested Each word Calculate the word in context Local logarithmic probability in: , in, This represents the Kneser-Ney smoothed n-gram probability, used to solve the probability estimation problem for low-frequency sequences.

[0078] Calculate the mean log probability within the sentence: , The mean log probability within a sentence can be used to measure the contextual coherence of the entire sentence: the lower the mean log probability (the more negative), the less likely the words in the sentence are to appear in the context, and the more likely the sentence contains illogical content.

[0079] S3-2-3: Determination of Illegal Words in Context Given a relative threshold or absolute threshold If satisfied or , Then Marked as "contextually invalid words" from the perspective of n-grams, forming a set: .

[0080] S3-3: Implementation steps of the CRF-based sequence validity labeling submodule S3-3-1: Labeling System Definition Define a tag set: . 3-3-2: Feature Extraction For each word Extracting feature vectors This includes, but is not limited to: word features themselves, part-of-speech tags, contextual word features, and position in the sentence.

[0081] 3-3-3: Training and Reasoning Maximize log-likelihood during training: , in, The features of the nth sample are the feature vectors of the word. For the corresponding tags; During detection, for each input sentence x and its corresponding label y, the Viterbi algorithm is used to find the optimal label sequence, i.e., the label with the highest probability. : It can also calculate the label confidence level for each location. .

[0082] Step S3-3-4: Determining Candidate Error Words Set threshold If a certain word Labeled , or the confidence level of its best legitimate category label: , Then add it to the candidate error word set from the CRF perspective: .

[0083] Step S3-4: Implementation steps of the rule-based and dictionary-assisted inspection submodule Step S3-4-1: Constructing the Domain Dictionary and Rule Base This step is used to build different domain dictionaries, including but not limited to: station name dictionary. Stock Road Name Dictionary ; Set of regular expressions for train number patterns Other domain entity dictionaries .

[0084] Step S3-4-2: Rule Check For the word to be detected With dictionary entries Define a string or phonetic similarity function. If a word belongs to a set However, there are certain typical matching scenarios. Taking the station dictionary as an example, such as... Make: , Then add the word to the set of suspicious words from the rule perspective: .

[0085] Step S3-5: Multi-view fusion The results from the four perspectives of PPL, n-gram, CRF, and rules are summarized as follows: High-credibility anchor set ; Uncertainty Content Collection ; Total set of candidate error words: .

[0086] The word legality discrimination module based on a multi-scale model performs simultaneous legality detection of instructions at three levels: sentence level, local word order level, and semantic label level. It combines the characteristics of each model; for example, the n-gram model has high sensitivity to extremely low-frequency word order combinations, while the CRF model performs better in terms of semantic category errors and new word generalization. Compared to existing solutions based on only a single model, it significantly improves the breadth and accuracy of illegal word detection and reduces the system's requirements for training data. Furthermore, the application of a rule-based and dictionary-assisted checking submodule provides an entry point for manually adding sample categories when data sample diversity is insufficient, increasing the system's scalability and generality. The precise retrieval constraint based on high-confidence anchors differs from traditional RAGs that directly use whole sentences or keywords for retrieval. The word legality discrimination module only uses high-confidence anchors that have been consistently confirmed by multiple models as the retrieval core. Even when the instruction as a whole has recognition noise or multiple errors, it still ensures that the retrieval remains stable within the correct business scenario, significantly reducing retrieval drift and noise interference.

[0087] After obtaining the judgment result, the system will perform domain knowledge retrieval based on the set of high-confidence anchor points, as well as error correction and rewriting of low-confidence content, and generate results based on the rewritten content, which is completed in the domain semantic rewriting module.

[0088] The implementation steps of the domain semantic rewriting module combining LLM and RAG include: 4-1: Implementation Steps for Domain Knowledge Retrieval Based on High-Confidence Anchor Points 4-1-1: Anchors and Context Vectorization For a set of high-confidence anchor points It is encoded with its sentence context to obtain a vector representation.

[0089] 4-1-2: Similarity Search By using different similarity retrieval methods, the top similarity results were selected. Large documents form a candidate knowledge set: 4-2: Implementation Steps for Domain Semantic Rewriting Based on LLM 4-2-1: Input Information Construction. Organize the following information into an LLM. Input Prompts: 1) Original or preliminary error-corrected sentences ; 2) Location of uncertain content and its local window context; 3) High-credibility anchor points The corresponding business entities and scenario descriptions; 4) Candidate knowledge set .

[0090] 4-2-2: Constraint Design. Clearly define constraints in the prompts, including: 1) Maintaining a certain "voice similarity" with the original distance can be achieved by limiting the editing distance or by using natural language to describe "minimizing changes to non-erroneous parts"; 2) Strictly adhere to regulations and templates; do not issue instructions that violate the rules. 3) Entities in high-confidence anchors must not be modified.

[0091] Let LLM be The output is the corrected final instruction text: .

[0092] The constraint of "voice similarity" can be further formalized using "editing cost": , in This represents the maximum amount of modification allowed.

[0093] Local semantic rewriting for uncertain content, because the scope of LLM rewriting is strictly limited to the marked uncertain content, the rest of the confirmed correct text fragments and anchors must remain unchanged; compared with the free generation of whole sentences, this local controlled rewriting greatly reduces the risk of LLM generating illusions or unintentionally changing the business meaning in security scenarios.

[0094] In addition, the dual-constraint rewriting strategy of phonetic similarity and domain semantic correctness requires LLM to prioritize legal terms that are similar to the original expression in phonetics or glossary, while ensuring the semantic correctness of railway business. This dual constraint is particularly suitable for handling near-phonetic misspellings caused by ASR.

[0095] The system supports lightweight deployment. LLM is only used to extract and fill fields from domain text and works under strict template constraints, making the invention suitable for both high-computing cloud environments and edge or station-level servers.

[0096] The electronic device of this invention includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) or loaded from a storage unit into random access memory (RAM). The RAM may also store various programs and data required for device operation. The CPU, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0097] Multiple components in the device are connected to the I / O interface, including: input units such as keyboards and mice; output units such as various types of displays and speakers; storage units such as disks and optical discs; and communication units such as network interface cards (NICs), modems, and wireless transceivers. The communication unit allows the device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0098] The processing unit performs the various methods and processes described above. For example, in some embodiments, the methods may be implemented as computer software programs tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed on the device via ROM and / or a communication unit. When the computer program is loaded into RAM and executed by the CPU, one or more steps of the methods described above may be performed. Alternatively, in other embodiments, the CPU may be configured to execute the methods by any other suitable means (e.g., by means of firmware).

[0099] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0100] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0101] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0102] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for railway traffic field semantic recognition based on small sample data, characterized in that, The method comprises: preprocessing the acquired railway traffic field data, constructing a field dictionary and a rule base; converting the preprocessed field data into enhanced speech, forming a training pair with the enhanced speech and the corresponding text, and fine-tuning an ASR model in the field; inputting the voice command in the railway traffic field into the ASR model fine-tuned in the field, and outputting an ASR initial transcription text; performing word legality discrimination on the ASR initial transcription text through a multi-scale model at three levels of sentence level, local word sequence level and semantic label level, identifying a high-confidence anchor point set and uncertain content; based on the word legality discrimination result, rewriting the field semantics, and finally outputting a structured instruction text.

2. The method according to claim 1, wherein, The word legality discrimination on the ASR initial transcription text through the multi-scale model comprises: confusion degree analysis and detection at the sentence level and the local word sequence level, and outputting a high-confidence anchor point set and an uncertain content set.

3. The method according to claim 2, wherein, The word legality discrimination on the ASR initial transcription text through the multi-scale model further comprises: performing local legality discrimination based on n-gram, and outputting a context-illegal-word set; performing sequence legality labeling based on CRF, and outputting a candidate error word set; based on the rule base and the field dictionary, performing secondary verification on the result determined as legal by CRF, and identifying a suspicious word set; integrating the context-illegal-word set, the candidate error word set and the suspicious word set to obtain a total candidate error word set.

4. The method according to claim 3, wherein, The CRF-based sequence validity labeling is specifically: when the label of the word is illegal, or the confidence of the best legal category label is lower than a threshold, the word is added to the candidate error word set under the CRF perspective.

5. The method according to claim 3, wherein, The secondary verification of the result determined as legal by the CRF is specifically: based on the constructed domain dictionary and rule base, for the to-be-detected words and the dictionary entries , a string definition or a phonetic similarity function is defined, if the words belong to the suspicious word set or the context illegal word set, if the similarity function is between the defined similarity lower limit and the similarity upper limit, the words are added to the suspicious word set under the rule perspective.

6. The method of claim 1, wherein the method is based on a small sample of data in the field of railway traffic. The rewriting of the field semantics based on the word legality discrimination result comprises: taking the high-confidence anchor point as a retrieval core keyword to perform domain knowledge retrieval, introducing a RAG architecture and a large language model, and performing local controlled local semantic rewriting on the uncertain content.

7. The method according to claim 6, wherein, Taking the high-confidence anchor point as a retrieval core keyword to perform domain knowledge retrieval comprises: encoding the high-confidence anchor point and the context of the sentence where the high-confidence anchor point is located to obtain a vectorized representation, and constructing a railway traffic domain knowledge base; retrieval query construction and execution: taking the high-confidence anchor point as a retrieval core keyword, combining the local context information of the sentence where the high-confidence anchor point is located, constructing a retrieval vector, and selecting the top K documents in descending order of similarity to form a candidate knowledge set.

8. The method according to claim 6, wherein, Performing local controlled local semantic rewriting on the uncertain content comprises: inputting the ASR initial transcription text, the uncertain content and the position of the candidate error word and its context window, the high-confidence anchor point set and the corresponding business entity information, and the candidate knowledge set into a large language model, setting explicit rewriting constraint rules in the large language model prompt words based on rewriting constraint rules, and outputting a locally modified final instruction text.

9. The method according to claim 8, wherein, The rewriting constraint rules comprise: (1) modification range constraint: only allow modification of the local part marked as uncertain content or candidate error word; (2) semantic compliance constraint: the rewriting result must satisfy the constraint of the retrieved candidate knowledge set; (3) term adaptation constraint: under the premise that the modified result satisfies the correct railway traffic business semantics, preferentially select legal professional terms similar in voice or form to the original recognition result.

10. The method of claim 8, wherein the method is based on a small sample of data in the field of railway traffic. The voice similarity between the final instruction text and the ASR initial transcription text should satisfy that it is not greater than the allowed maximum modification amount.

11. The method of claim 1, wherein the method is based on a small sample of data in the field of railway traffic. The preprocessing of the obtained railway traffic field data comprises: collecting text data related to train dispatching to form an original railway traffic field text set; sequentially performing sentence segmentation and word segmentation on each original railway traffic field text set; performing part-of-speech tagging and syntax information tagging on each word to obtain a word sequence with labels; performing format normalization processing on key fields of the word sequence with labels to obtain a normalized sentence set.

12. An electronic device comprising a memory and a processor, said memory having stored thereon a computer program, characterized in that, The processor implements the method of any one of claims 1-11 when executing the program.

13. A computer readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the method of any one of claims 1-11.

Citation Information

Patent Citations

  • Text semantic recognition method, acquisition method of model text semantic recognition method and related device

    CN111144127A

  • Prompt-based high-efficiency small sample dialogue semantic understanding method

    CN115983282A

  • Speech recognition method for power transmission operation inspection based on small sample synthesis

    CN117292680A

  • Corpus construction method and system for subway field and storage medium

    CN117875304A

Cited By

  • Hybrid architecture semantic restoration method and device based on DFA and large language model

    CN121960501A

  • Dfa and large language model based hybrid architecture semantic repair method and device

    CN121960501B