Railway alarm processing information text recognition method based on large language model

Through the railway alarm processing information text recognition method based on the large language model, the railway alarm processing information text is generated and classified, and the problems of data category imbalance and insufficient classification accuracy are solved, efficient and accurate data processing is achieved, and labor investment is reduced.

CN120067883AActive Publication Date: 2025-05-30SIGNAL & COMM RES INST OF CHINA ACAD OF RAILWAY SCI +3
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510562952.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-05-30
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

There is a problem of category imbalance in the railway alarm processing information text data, which leads to insufficient accuracy of data classification, requiring manual secondary sorting, which consumes huge labor, and may cause incorrect questions, affecting the quality of the data.

Method used

The railway alarm processing information text recognition method based on a large language model is adopted to perform data generation or classification tasks by analyzing the user intention of the input text. Use the first major language model to generate railway alarm processing information text, and use retrieval enhancement generation technology to combine knowledge base to classify data to improve the accuracy and reliability of data classification.

Benefits of technology

It significantly reduces the labor investment required to process railway alarm processing information text data, improves the accuracy and reliability of data classification, and solves the problem of imbalance in the data categories of railway alarm processing information text data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067883A_ABST
    Figure CN120067883A_ABST
Patent Text Reader

Abstract

The invention discloses a railway alarm processing information text recognition method based on a large language model, and provides a scientific, reasonable, high-applicability and good-effect scheme for solving the problems existing in the railway alarm processing information text recognition process so as to deeply mine and analyze effective information in text data. According to the method, the framework is easy to establish, the method is easy to implement, the railway alarm processing information texts of different categories are expanded and enhanced by utilizing the powerful generation capability of the large language model, and the problem of sample imbalance is solved. And in combination with a retrieval enhancement generation technology, input prompts are fused with a knowledge base retrieval result, a stronger reasoning basis is obtained, the credibility of a user to a data classification model and the accuracy of railway alarm processing information text classification are improved, and generally, the method has the advantages of being scientific, reasonable, high in applicability, good in effect and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of railway systems, and particularly to a method for identifying text of railway alarm processing information based on a large language model. Background Art

[0002] The analysis of the causes of railway signal equipment failures has always been a key concern in the signal field. In particular, text mining of data such as alarm records and alarm handling opinions is particularly important and there is a wide and urgent need. At present, the signal centralized monitoring system has accumulated a large amount of text-based fault data such as alarm records and alarm handling opinions. The text information recorded by signal maintenance personnel is subject to the expression habits and cultural background differences of the recorders, and there is a problem of ambiguity in the description content. This forces signal maintenance personnel to manually re-organize this data when performing tasks such as fault cause classification, consuming a huge amount of labor. At the same time, there may also be a phenomenon where the text does not match the topic due to subjective or objective reasons, seriously affecting the data quality and resulting in inaccurate processing, mining, and analysis results based on this data. Generally speaking, the existing technology has problems of class imbalance in the text data of railway alarm processing information, and the accuracy of data classification also needs to be improved.

[0003] In existing research, although some improvements have been made, the effects are not good. For example: (1) The article "Using Signal Centralized Monitoring to Process Alarm Information" published by Xie in the 5th issue of "Railway Signalling & Communication Engineering" in 2019 (i.e., May 2019) takes the railway signal centralized monitoring system as the research object and proposes a three-level classification method for alarm information. However, this scheme can only be applied to the classification of relatively simple alarm information, and the classification accuracy also needs to be improved; (2) The Chinese patent application for invention "Railway Signal Monitoring Data Generation Method and System" with the publication number CN119128802A, published on December 13, 2024. Although this scheme can generate railway signal monitoring data, it generates by collecting a series of external information and cannot be applied to the generation of railway alarm processing information text, thus unable to solve the problem of class imbalance in railway alarm processing information text data.

[0004] In view of this, the present invention is specifically proposed. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for identifying text of railway alarm processing information based on a large language model, which can achieve effective expansion and enhancement of the text of railway alarm processing information, as well as accurate and reliable classification of the text data of railway alarm processing information, significantly reducing the labor input required in the process of processing the text data of railway alarm processing information.

[0006] The purpose of the present invention is achieved through the following technical solutions: A method for identifying railway alarm processing information text based on a large language model, including: Analyze the input text to identify the user intention contained in the input text; the user intention includes: performing a data generation task or performing a data classification task; If the user intention is to perform a data generation task, according to the data category information contained in the input text, obtain multiple keywords from the database, generate a railway alarm processing information text through the first large language model, and then select the finally output railway alarm processing information text in combination with the set screening method for updating the database and knowledge base; If the user intention is to perform a data classification task, obtain the updated knowledge base, use the retrieval-augmented generation technology and combine it with the input text to generate input prompt information, and use the second large language model to classify the input text according to the input prompt information.

[0007] It can be seen from the technical solution provided by the present invention above that the powerful generation ability of the large language model is used to expand and enhance the railway alarm processing information text data of different categories, solving the problem of imbalance in various categories of railway alarm processing information text data. And combined with the retrieval-augmented generation technology, the input prompt is fused with the retrieval results of the knowledge base, obtaining a stronger reasoning basis, improving the user's confidence in the data classification model and the accuracy of railway alarm processing information text data classification. Generally speaking, it has the advantages of being scientific and reasonable, strong applicability, good effect, etc. Description of the Drawings

[0008] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0009] Figure 1 It is a flowchart of a method for identifying railway alarm processing information text based on a large language model provided by an embodiment of the present invention; Figure 2 It is an overall framework schematic diagram of a method for identifying railway alarm processing information text based on a large language model provided by an embodiment of the present invention. Detailed Embodiments

[0010] The following combines the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0011] First, the terms that may be used in this article are explained as follows: Descriptions with semantic meanings such as "including", "comprising", "containing", "having", or other similar ones should be interpreted as non-exclusive inclusion. For example, including a certain technical feature element (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, processes, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products, or articles, etc.) should be interpreted as not only including the explicitly listed certain technical feature element, but also including other technical feature elements well-known in the art that are not explicitly listed.

[0012] The term "consisting of..." means excluding any technical feature elements that are not explicitly listed. If this term is used in a claim, this term will make the claim a closed type, making it not contain technical feature elements other than the explicitly listed ones, except for the related conventional impurities. If this term only appears in a certain clause of the claim, then it only limits the elements explicitly listed in that clause, and the elements recorded in other clauses are not excluded from the overall claim.

[0013] Next, a method for identifying railway alarm processing information text based on a large language model provided by the present invention will be described in detail. The content not described in detail in the embodiments of the present invention belongs to the prior art well-known to those skilled in the art. For the conditions not specified in the embodiments of the present invention, they are carried out according to the conventional conditions in the art or the conditions recommended by the manufacturer. For the reagents or instruments not specified in the production manufacturers in the embodiments of the present invention, they are all conventional products that can be obtained by purchasing in the market.

[0014] As Figure 1 shown, a method for identifying railway alarm processing information text based on a large language model provided by the embodiments of the present invention mainly includes the following steps: Step 1: Analyze the input text to identify the user intention included in the input text.

[0015] In the embodiments of the present invention, the input text also belongs to the railway alarm processing information text; the user intention includes: performing a data generation task or performing a data classification task; the essence of identifying the user intention is to classify to determine the task to be performed subsequently.

[0016] In the embodiments of the present invention, the data generation task is mainly responsible for generating railway alarm processing information text to update the knowledge base, and the data classification task is to perform data classification in combination with the updated knowledge base.

[0017] Preferably, this step can be implemented by a pre-trained first adapter; the first adapter includes: a text vectorization operation layer, a first fully connected layer, a first non-linear activation layer, and a first Softmax layer arranged in sequence; where Softmax is a normalized exponential function.

[0018] Preferably, an execution agent can also be configured to call a corresponding model according to the user's intention to execute relevant tasks.

[0019] Step 2: Execute the data generation task, generate the railway alarm handling information text, and update the database and knowledge base.

[0020] In the embodiment of the present invention, the data generation task is executed by a data generation model. The main work content of the data generation model includes: keyword extraction, text generation based on the first large language model, and text screening, etc. Specifically: according to the data category information included in the input text, multiple keywords are obtained from the database, and the first large language model is used to generate the railway alarm handling information text by using multiple keywords, and then combined with a set screening method, the finally output railway alarm handling information text is selected to update the database and knowledge base.

[0021] Preferably, the obtaining of multiple keywords from the database according to the data category information included in the input text includes: obtaining all the railway alarm handling information texts of the corresponding data category from the database according to the data category information included in the input text; extracting keywords from each railway alarm handling information text respectively through an end-to-end sequence labeling model, and finally obtaining multiple keywords.

[0022] Preferably, a second adapter is configured at the output end of the first large language model, and the first large language model is fine-tuned by fine-tuning the second adapter; the second adapter includes: an information bottleneck layer, a second non-linear activation layer, and an expansion layer arranged in sequence.

[0023] Preferably, the combining of the set screening method to select the finally output railway alarm handling information text includes: respectively performing vectorization processing on each railway alarm handling information text in the database through a text embedding model, and the vectorization result is called the original data vector, and finally obtaining multiple original data vectors; clustering all the original data vectors to obtain the cluster centers; performing vectorization processing on each generated railway alarm handling information text through the text embedding model, and the obtained result is called the generated data vector, and finally obtaining multiple generated data vectors; respectively measuring the similarity between each generated data vector and the cluster centers through a similarity function; if the similarity is higher than the set threshold, then retain the corresponding generated railway alarm handling information text; otherwise, discard it; all the retained generated railway alarm handling information texts are the finally output railway alarm handling information texts.

[0024] Preferably, the updated knowledge base includes: splitting each railway alarm handling information text of the final output into multiple text blocks according to a preset text block splitting size and overlapping size; performing vectorization processing on each text block through a text embedding model, and then uploading it to the knowledge base.

[0025] Step 3, perform a data classification task, and classify the input text in combination with the updated knowledge base.

[0026] In the embodiment of the present invention, the data classification task is performed by a data classification model. The main work content of the data classification model includes: generating input prompt information based on the retrieval-augmented generation (RAG) technology, and classifying the input text based on the second large language model, etc. Specifically: obtaining the updated knowledge base, using the retrieval-augmented generation technology and combining with the input text to generate input prompt information, and using the second large language model to classify the input text according to the input prompt information. The knowledge base is continuously updated by repeatedly performing the data generation task to balance the data categories; moreover, the knowledge base used when performing the data classification task is also the knowledge base updated after repeatedly performing the data generation task. Thus, the data classification model can more accurately complete the data classification task.

[0027] Preferably, obtaining the updated knowledge base and using the retrieval-augmented generation technology and the input text to generate input prompt information includes: after obtaining the updated knowledge base, using the retrieval-augmented generation technology to retrieve the k text blocks with the highest similarity to the input text from the updated knowledge base; where k is a set value, for example, set k = 4; combining the retrieved k text blocks with the input text to generate input prompt information.

[0028] Preferably, a third adapter is configured at the output end of the second large language model, and the second large language model is fine-tuned by fine-tuning the third adapter; the third adapter includes: a second fully connected layer, a third non-linear activation layer, and a second Softmax layer arranged in sequence; where Softmax is a normalized exponential function.

[0029] In the embodiment of the present invention, the data generation task serves the data classification task. Considering that there is currently less existing data and there is an extremely unbalanced phenomenon in categories, it is impossible to train a good data classification model. Therefore, the data generation task is introduced. On the one hand, the data volume is increased through the data generation task; on the other hand, when new type of data appears, when there is less data of a certain type or there is an unbalanced problem in categories in the existing data, through the data generation task, the data category distribution can be balanced, and then these data are used for training to obtain a better data classification model.

[0030] In the embodiments of the present invention, the fine-tuning principles of the two large language models are similar. Taking the second large language model as an example, its fine-tuning is achieved by combining the parameters of the second large language model with the third adapter, that is, inserting the third adapter after the second large language model to fine-tune the second large language model. In addition, both of the two large language models in the embodiments of the present invention are pre-trained models, and users can select the specific types of the two large language models according to actual situations, which are not limited in the present invention. At the same time, during the fine-tuning process of the two large language models, users can select the corresponding data set for fine-tuning according to the task type of the large language model, which is not limited in the present invention.

[0031] Those skilled in the art can understand that fine-tuning is a general technical term in this field, which refers to slightly adjusting the parameters of a pre-trained model in combination with the requirements of new tasks (for example, the data generation task and data classification task in the present invention) so that it can be applied to new tasks.

[0032] Those skilled in the art can understand that the outputs of the above first Softmax layer and second Softmax layer are both classification probabilities, and the category with the highest probability value is the classification result.

[0033] Preferably, it further includes the step of pre-constructing a knowledge base, and the steps include: obtaining the text of railway alarm handling information in the database, and splitting each text of railway alarm handling information into multiple text blocks according to the preset text block splitting size and overlapping size; performing vectorization processing on each text block respectively through a text embedding model to obtain the corresponding text block vectors; constructing a knowledge base by using all the obtained text block vectors. Specifically, before the first execution of the data classification task, it is necessary to construct a knowledge base according to the above solution. During the application process, new texts of railway alarm handling information will be continuously generated and the knowledge base will be updated as the data generation task is executed. On the one hand, it can solve the problem of unbalanced categories of railway alarm handling information texts, and on the other hand, it can also achieve accurate and reliable classification of railway alarm handling information texts, significantly reducing the labor input required in the process of processing railway alarm handling information text data.

[0034] In the embodiments of the present invention, the first and second mainly play an identifying role. For example, the main purpose of the first large language model and the second large language model is to indicate that two large language models are involved here. In actual applications, the two large language models can be of the same type or different types, and the specific types can be determined by those skilled in the art according to actual situations or experience, which are not limited in the present invention.

[0035] In order to more clearly show the technical solutions provided by the present invention and the technical effects produced, the following uses specific embodiments to describe in detail the methods provided by the embodiments of the present invention.

[0036] In recent years, the progress of deep learning and natural language processing technologies has brought new prospects to the field of artificial intelligence. For example, the large language model ChatGPT (Chat Generative Pretrained Transformer) released by OpenAI can focus on generative dialogue question and answer. In addition to being able to understand complex semantic questions raised by humans more effectively and accurately, the large language model also supports the expression of various heterogeneous knowledge and multi-round semantic interactions, thus realizing general artificial intelligence in the true sense. In addition, by training in a self-supervised or semi-supervised manner on a large amount of unlabeled text, the large language model can have general capabilities and can perform a wide range of natural language processing tasks. In view of this, the present invention utilizes the powerful capabilities of the large language model in text processing to achieve effective mining and analysis of text data of railway alarm handling information. As Figure 2 shown, it is the overall framework of the present invention, and the following will mainly introduce it from two parts: user intention recognition and model invocation.

[0037] I. User intention recognition.

[0038] The input text is first classified by the first adapter to identify the user's intention, that is, to perform a data generation task or a data classification task. Usually, in order to ensure the application effect, it is necessary to execute the data generation task multiple times according to the data category situation in the database to make the data categories balanced (that is, the number of data in each category is roughly the same). This process can be guided by the input text. For example, when the data volume of a certain type (assumed to be the skylight operation category) is small, the input text can be set as "generate / enlarge the text data of railway alarm handling information in the skylight operation category and classify it". At this time, after analysis by the first adapter, it is obtained that the data generation task of the skylight operation category needs to be done at this time, and then the data generation model will be called through the execution agent to generate the text of railway alarm handling information in the skylight operation category, which can increase the data volume and thus help the data classification model to better classify the data. If the first adapter identifies that the user wants to classify a certain piece of railway alarm handling information text, the data classification model will be called through the execution agent to classify and display the input text.

[0039] In the embodiment of the present invention, the first adapter mainly includes: a text vectorization operation layer, a first fully connected layer, a first non-linear activation layer, and a first Softmax (normalized exponential function) layer arranged in sequence. The text vectorization operation layer here mainly refers to converting the input text into a vector form. The first fully connected layer mainly performs a linear transformation on the vector output by the text vectorization operation layer. The first non-linear activation layer mainly performs non-linear activation on the processing result of the first fully connected layer to introduce non-linear factors. The first Softmax outputs the prediction probability to reflect the user's intention; considering that the specific processing scheme involved in this part can be implemented with reference to conventional technologies, it will not be elaborated here.

[0040] II. Knowledge Base Construction

[0041] Under normal circumstances, before calling the data generation model and the data classification model, it is necessary to complete the construction of the knowledge base. Specifically, under the langchain framework (which is a development framework driven by large language models), load the text of railway alarm handling information in the database and split it into multiple text blocks according to the preset text segmentation size and overlap size; then use a text embedding model (for example, the text2vec model mentioned above) to vectorize each text block respectively; finally, use the vectors of each text block to construct the knowledge base.

[0042] Those skilled in the art can understand that the text segmentation size can be understood as the number of words contained in each text block, and the overlap size represents the number of overlapping words contained in adjacent text blocks. Exemplarily, the text segmentation size can be set to 64 and the overlap size can be set to 32.

[0043] III. Model Call

[0044] In the embodiments of the present invention, the model call mainly calls the corresponding model in combination with the user's intention. Specifically, if the user's intention is a data generation task, the data generation model is called; if the user's intention is a data classification task, the data classification model is called.

[0045] The working processes of the two types of models will be introduced in detail below.

[0046] 1. Working Process of the Data Generation Model

[0047] When the user's intention is a data generation task, the data generation model is called through the execution agent. As Figure 2 shown, the main working contents of the data generation model include: keyword extraction, text generation based on the first large language model, and text screening, etc. The main steps are as follows.

[0048] (1.1) Keyword Extraction: According to the data category information contained in the input text, extract all the text of railway alarm handling information of the corresponding category from the database, and extract the keywords in each text of railway alarm handling information through an end-to-end sequence labeling model.

[0049] Exemplarily, the extracted keywords are stored in the following structure: {#keyword1#keyword2#keyword3…}, where the symbol # is mainly used to separate different keywords, and the values involved in keyword1, 2, 3, etc. are the codes of different keywords. In the application, keyword1, 2, 3 are the actual keywords. In addition, the number of extracted keywords is set according to the actual situation, and the present invention does not make any restrictions.

[0050] Exemplarily, an end-to-end sequence labeling model can adopt an LSTM-CNNs-CRF model, which is a hybrid model formed by combining LSTM (Long Short-Term Memory Neural Network), CNNs (Convolutional Neural Network), and CRF (Discriminative Model).

[0051] (1.2) Text data generation: Input the extracted multiple keywords into the first large language model fine-tuned by the second adapter to generate the text of railway alarm handling information.

[0052] Use the keywords extracted in the above step (1.1) as the input to the first large language model to generate the text of railway alarm handling information. The first large language model is fine-tuned in an efficient parameter way, that is, by fine-tuning the second adapter and inserting it at the end of the first large language model to achieve the fine-tuning of the first large language model. The specific structure of the second adapter includes: an information bottleneck layer to compress the data; a second non-linear activation layer to improve the non-linear ability of the model through a non-linear activation function, and an expansion layer to expand the compressed data to the same dimension as the original features. Considering that the specific processing scheme involved here can be implemented with reference to conventional technologies, it will not be elaborated.

[0053] (1.3) Generated data filtering: Use a filter to screen the data that meets the requirements in the generated text of railway alarm handling information and store it in the database and knowledge base.

[0054] In the embodiment of the present invention, the effectiveness of the generated text of railway alarm handling information is ensured by screening. The main steps include: (1.3.1) Vectorization of existing data: Use a text embedding model to vectorize the existing text of railway alarm handling information in the database.

[0055] Exemplarily, the text embedding model can select the text2vec model (text vectorization model).

[0056] (1.3.2) Calculation of the clustering center: Calculate the mean value of all vectors obtained in the above step (1.3.1) as the clustering center.

[0057] (1.3.3) Vectorization of generated data: Use a text embedding model to vectorize the text of railway alarm handling information generated in the above step (1.2).

[0058] (1.3.4) Similarity measurement: The similarity (e.g., cosine similarity) between each vector obtained in the above step (1.3.3) and the clustering center.

[0059] (1.3.2) Text screening: If the calculated similarity is greater than the set threshold λ, retain the corresponding generated text of railway alarm handling information and store it in the database and knowledge base, otherwise, discard it.

[0060] Specifically, denote the clustering center as A, and denote the vector of the generated single railway alarm processing information text as B. Calculate the cosine similarity S through the following formula:

[0061] where the symbol represents the dot product operation, and the symbol is the norm symbol.

[0062] If the cosine similarity S is greater than the set threshold , that is , then the generated railway alarm processing information text is retained. Otherwise, it means that the generated railway alarm processing information text does not meet the requirements and is discarded.

[0063] (1.4) Generation data saving.

[0064] Store the railway alarm processing information text filtered by the filter into the database and upload to update the knowledge base. When updating the knowledge base, for each railway alarm processing information text filtered by the filter, the following processing is performed: split it into multiple text blocks according to the preset text block splitting size and overlapping size; then use a text embedding model (for example, the text2vec model described above) to vectorize each text block respectively; finally, upload it to the knowledge base.

[0065] So far, the call of the data generation model is completed. In practical applications, calling the data generation model multiple times can effectively expand and enhance the railway alarm processing information text.

[0066] 2. Working process of the data classification model.

[0067] (2.1) Input prompt information generation: According to the input text, retrieve the top k text blocks (in the present invention, k is preset to 4) with the highest similarity to the input text from the updated knowledge base, and combine them with the input text to generate input prompt information, so as to improve the classification credibility of the data classification model.

[0068] (2.2) Text classification and recognition: Input the input prompt information generated in the above step (2.1) into the fine-tuned second large language model to complete the classification of the input text.

[0069] In the embodiment of the present invention, a third adapter is configured at the output end of the second large language model, and the second large language model is fine-tuned by fine-tuning the third adapter. The third adapter includes: a second fully connected layer, a third non-linear activation layer, and a second Softmax layer arranged in sequence; to realize the classification of the input text.

[0070] The framework of the method provided by the embodiments of the present invention is easy to establish, and the method is easy to implement. By utilizing the powerful generation ability of the large language model to augment and enhance different categories of text data, the problem of imbalance in various categories of railway alarm processing information text is solved. And combined with retrieval-augmented generation, the input prompt is fused with the retrieval results of the knowledge base, obtaining a stronger reasoning basis, improving the credibility of the data classification model and the accuracy of alarm text classification. It has the advantages of being scientific and reasonable, having strong applicability, and good effects.

[0071] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software or by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0072] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims. The information disclosed in the background art part of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or any form of suggestion that this information constitutes the prior art already known to those skilled in the art.

Claims

1. A method for recognizing railway alarm processing information text based on a large language model, characterized in that: include: Analyze the input text and identify the user intent contained in the input text; The user intention includes: performing a data generation task or performing a data classification task; If the user intends to perform a data generation task, multiple keywords are obtained from the database according to the data category information contained in the input text, and the railway alarm processing information text is generated through the first language model. Then, the final output railway alarm processing information text is selected in combination with the set screening method to update the database and knowledge base; If the user intends to perform a data classification task, an updated knowledge base is obtained, and input prompt information is generated by using retrieval enhancement generation technology and combining with the input text. The input text is classified according to the input prompt information using the second largest language model.

2. The method for railway alarm processing information text recognition based on a large language model according to claim 1 is characterized in that: Analyze the input text and identify the user intent contained in the input text through a pre-trained first adapter; The first adapter includes: a text vectorization operation layer, a first fully connected layer, a first nonlinear activation layer and a first Softmax layer arranged in sequence; wherein Softmax is a normalized exponential function.

3. The method for railway alarm processing information text recognition based on a large language model according to claim 1 is characterized in that: The acquiring of multiple keywords from a database according to data category information contained in the input text comprises: According to the data category information contained in the input text, all railway alarm processing information texts of the corresponding data category are obtained from the database; Keywords are extracted from each railway alarm processing information text through an end-to-end sequence labeling model, and finally multiple keywords are obtained.

4. The method for railway alarm processing information text recognition based on a large language model according to claim 1 is characterized in that: The output end of the first large language model is configured with a second adapter, and the first large language model is fine-tuned by fine-tuning the second adapter; The second adapter includes: an information bottleneck layer, a second nonlinear activation layer and an expansion layer which are arranged in sequence.

5. The method for railway alarm processing information text recognition based on a large language model according to claim 1 is characterized in that: The railway alarm processing information text that is finally outputted by combining the set screening method includes: Each railway alarm processing information text in the database is vectorized by a text embedding model. The vectorization result is called a raw data vector, and finally multiple raw data vectors are obtained. Cluster all original data vectors to obtain cluster centers; Each generated railway alarm processing information text is vectorized by a text embedding model, and the result obtained is called a generated data vector, and finally multiple generated data vectors are obtained; Through the similarity function, the similarity between each generated data vector and the cluster center is measured respectively; if the similarity is higher than the set threshold, the corresponding generated railway alarm processing information text is retained; otherwise, it is discarded; all retained generated railway alarm processing information texts are the final output railway alarm processing information texts.

6. The method for railway alarm processing information text recognition based on a large language model according to claim 1 is characterized in that: Updated knowledge base includes: Each railway alarm processing information text that is finally output is divided into multiple text blocks according to a preset text block segmentation size and overlap size; Each text block is vectorized through the text embedding model and then uploaded to the knowledge base.

7. The method for railway alarm processing information text recognition based on a large language model according to claim 6 is characterized in that: The obtaining of the updated knowledge base and the use of the search enhancement generation technology combined with the input text to generate input prompt information include: After obtaining the updated knowledge base, the search enhancement generation technology is used to retrieve the k text blocks with the highest similarity to the input text from the updated knowledge base; where k is a set value; The k retrieved text blocks are combined with the input text to generate input prompt information.

8. The method for railway alarm processing information text recognition based on a large language model according to claim 1 is characterized in that: The output end of the second largest language model is configured with a third adapter, and the second largest language model is fine-tuned by fine-tuning the third adapter; The third adapter includes: a second fully connected layer, a third nonlinear activation layer and a second Softmax layer arranged in sequence; wherein Softmax is a normalized exponential function.

9. The method for railway alarm processing information text recognition based on a large language model according to claim 1, characterized in that: It also includes the steps of pre-building the knowledge base, including: Acquire the railway alarm processing information text in the database, and divide each railway alarm processing information text into multiple text blocks according to the preset text block segmentation size and overlap size; Each text block is vectorized by the text embedding model to obtain the corresponding text block vector; The knowledge base is constructed using all the obtained text block vectors.

Citation Information

Patent Citations

  • Railway signal monitoring data generation method and system

    CN119128802A

  • Attention mechanism-based domain large model knowledge base classification and construction method

    CN117573829A

  • User risk behavior perception method based on large language model and related equipment

    CN118504586A

  • Marine early warning intelligent question and answer method based on large language model and related device

    CN119179756A

  • Credit business system evolution method and device based on large model RAG technology

    CN119204205A