Rail transit large model security check detection method and system based on multi-source information fusion

By adopting a large-scale model detection method of multi-source information fusion in rail transit security inspection, the problem of poor accuracy and interpretability in existing security inspection technologies is solved, and the comprehensive analysis and judgment of multi-dimensional security inspection information is realized, operational efficiency and security are improved, and the results are output in natural language.

CN120126082APending Publication Date: 2025-06-10JINAN RAILWAY TRANSPORT GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510274264.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing rail transit security inspection technology has poor accuracy and interpretability, and the security inspection information is single and cannot be analyzed in multiple dimensions, resulting in the inability to effectively and comprehensively utilize the information, and there are many manual interventions and are greatly affected by human experience.

Method used

The rail transit large-scale security detection method based on multi-source information fusion is adopted. By obtaining optical video images, X-ray transmission images and terahertz human body images, the object detection algorithm is used for detection, and the fused target information is obtained after embedding operations, and the security inspection evaluation model is input for processing. The natural language description is combined with the large language model to realize the comprehensive analysis and judgment of multi-dimensional security inspection information.

Benefits of technology

It improves the accuracy and interpretability of security inspection results, reduces manual intervention, improves the "pass inspection efficiency" and "safety coefficient" of rail transit operations, and outputs the security inspection results in natural language that is easy for humans to understand.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126082A_ABST
    Figure CN120126082A_ABST
Patent Text Reader

Abstract

The invention discloses a rail transit large model security check detection method and system based on multi-source information fusion, and the method comprises the steps: carrying out the detection of an optical video image, an X-ray transmission image and a terahertz human body image through an optical target detection algorithm, an X-ray target detection algorithm and a terahertz target detection algorithm, and obtaining the corresponding target information; performing embedding operation on the optical detection target information, the X-ray detection target information and the terahertz detection target information, and splicing to obtain fused target information; inputting the fused target information into a security check evaluation model for processing to obtain a passenger security check evaluation result; the security check evaluation model comprises an encoder and a decoder, the decoder generates output data and inputs the output data into a large language model of rail transit security check for processing, and a natural language description corresponding to a passenger security check evaluation result is obtained. Various security check information is comprehensively utilized, and natural language summarized security check information is output in combination with a language big model, so that the security check accuracy and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of security inspection, and particularly relates to a security inspection and detection method and system for rail transit large models based on multi-source information fusion. Background Technique

[0002] The statements in this part only provide background technical information related to the present disclosure and do not necessarily constitute prior art.

[0003] With the rapid development of science and technology, different security inspection technologies and means are increasingly applied in the field of urban rail transit. At present, common urban rail transit security inspections include passenger body security inspections and item security inspections. The main body security inspection methods are hand-held metal detectors, optical camera monitoring, and metal security gates, and items are mainly detected by X-ray security inspection machines. However, with the continuous advancement of the construction of smart rail transit, a series of new security inspection technologies have been widely applied in the field of urban rail transit. The most representative one is the passive terahertz body security inspection equipment, which is continuously applied to the body security inspection field due to its characteristics such as detecting multi-material dangerous goods such as metals, liquids, and powders, non-contact and non-stop security inspection, and no ionizing radiation.

[0004] The existing security inspection means only stay at detecting by using target detection algorithms and do not further process the security inspection result information, resulting in poor accuracy and interpretability of the security inspection results. At the same time, the existing security inspection means have deficiencies such as single acquisition of security inspection information, independence of various security inspection means, and inability to analyze security inspection information multi-dimensionally, resulting in ineffective comprehensive utilization of security inspection information, a large amount of manual intervention, and being greatly affected by human experience. In addition, under the traditional security inspection method, security inspection personnel need to manually identify dangerous items, which also has relatively high requirements for the professionalism of security inspection practitioners. Summary of the Invention

[0005] To overcome the above-mentioned deficiencies of the prior art, the present invention provides a security inspection and detection method and system for rail transit large models based on multi-source information fusion, which comprehensively utilizes the security inspection information of various security inspection devices such as optical video monitoring, X-ray luggage inspection, and terahertz body security inspection equipment, analyzes the multi-dimensional security inspection information of rail transit passengers through comprehensive analysis and processing and large language models, and gives a comprehensive judgment result, without excessive intervention of security inspection personnel, and can effectively promote the double improvement of the "security inspection efficiency" and "safety factor" of rail transit operations.

[0006] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions: In the first aspect, the present invention provides a security inspection and detection method for rail transit large models based on multi-source information fusion, including: Obtain optical video images, X-ray transmission images, and terahertz body images; The optical video image is detected using an optical target detection algorithm to obtain optical detection target information, the X-ray transmission image is detected using an X-ray target detection algorithm to obtain X-ray detection target information, the terahertz human body image is detected using a terahertz target detection algorithm to obtain terahertz detection target information, and after performing an embedding operation on the optical detection target information, X-ray detection target information, and terahertz detection target information, they are spliced to obtain fused target information; The fused target information is input into a security inspection evaluation model for processing to obtain a passenger security inspection evaluation result; the security inspection evaluation model includes an encoder and a decoder, the decoder receives the context vector generated by the encoder, and generates output data based on the context vector; The output data is input into a large language model for rail transit security inspection for processing to obtain a natural language description corresponding to the passenger security inspection evaluation result.

[0007] In a further technical solution, Embedding embedding is used to perform an embedding operation on the optical detection target information, X-ray detection target information, and terahertz detection target information respectively, mapping a one-dimensional vector into a two-dimensional matrix to obtain corresponding matrices.

[0008] In a further technical solution, the encoder consists of two sub-layer connection structures, successively including a multi-head self-attention sub-layer and a feed-forward fully connected sub-layer, and each sub-layer is followed by a normalization layer and a residual connection.

[0009] In a further technical solution, after receiving the fused target information, the encoder outputs a context vector, and the context vector is input into a classification network to obtain a passenger security inspection evaluation result.

[0010] In a further technical solution, the classification network is a fully connected network with an activation function, successively including an input layer, a hidden layer, and an output layer.

[0011] In a further technical solution, the decoder consists of three sub-layer connection structures, successively including a masked multi-head self-attention sub-layer, a multi-head attention sub-layer, and a feed-forward fully connected sub-layer, and each sub-layer is followed by a normalization layer and a residual connection.

[0012] In a further technical solution, the large language model for rail transit security inspection is obtained by fine-tuning the language large model based on the output data of the decoder.

[0013] In a second aspect, the present invention provides a rail transit large model security inspection detection system based on multi-source information fusion, including: A data acquisition module, which is configured to: acquire an optical video image, an X-ray transmission image, and a terahertz human body image; A multi-source information fusion module, which is configured to: use an optical target detection algorithm to detect the optical video image to obtain optical detection target information, use an X-ray target detection algorithm to detect the X-ray transmission image to obtain X-ray detection target information, use a terahertz target detection algorithm to detect the terahertz human body image to obtain terahertz detection target information, and perform an embedding operation on the optical detection target information, X-ray detection target information, and terahertz detection target information and then splice them to obtain fusion target information; A security inspection evaluation module, which is configured to: input the fusion target information into a security inspection evaluation model for processing to obtain a passenger security inspection evaluation result; the security inspection evaluation model includes an encoder and a decoder, and the decoder receives the context vector generated by the encoder and generates output data based on the context vector; A security inspection description module, which is configured to: input the output data into a large language model for rail transit security inspection for processing to obtain a natural language description corresponding to the passenger security inspection evaluation result.

[0014] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the rail transit large model security inspection detection method based on multi-source information fusion as described in the first aspect.

[0015] In a fourth aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps in the rail transit large model security inspection detection method based on multi-source information fusion as described in the first aspect.

[0016] The above one or more technical solutions have the following beneficial effects: The present invention comprehensively utilizes the security inspection information of various security inspection devices such as optical video monitoring, X-ray luggage inspection, and terahertz human body security inspection equipment, designs a security inspection evaluation model for comprehensive analysis and processing and a large language model for fine-tuning, analyzes the multi-dimensional security inspection information of rail transit passengers, and gives a comprehensive judgment result, which is more accurate, requires little intervention from security inspection personnel, and can effectively promote the double improvement of the "security inspection efficiency" and "safety factor" of rail transit operations.

[0017] Based on the fine-tuned large language model, the present invention obtains the security inspection summary information in natural language, and finally outputs it as natural language that is easy for humans to understand, which is concise and friendly to security inspection personnel using the equipment.

[0018] In view of the requirements of intelligent security inspection for rail transit, the present invention effectively combines the current mainstream security inspection technology information and large language models. By utilizing the designed network structure, the security inspection information of passengers is finally summarized and output in natural language that is easy for security inspection personnel to understand. On the one hand, it integrates various security inspection information. After training and learning various security inspection information such as optical video surveillance, X-ray luggage detection, and terahertz body security inspection through the neural network structure designed by the present invention, the security inspection information is effectively comprehensively utilized, solving the problem of mutual isolation of different current security inspection information and enabling a more multi-dimensional and comprehensive description of the security inspection information of passengers. On the other hand, it combines with the language large model. By fine-tuning the large language model and inputting the security inspection information features learned by the neural network into the large model, the security inspection results are finally output in the form of natural language, and the security inspection evaluation results of passengers output by the model are explained, effectively alleviating the work pressure of security inspection personnel and improving the security inspection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings forming a part of this invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0020] Figure 1 is a flowchart of the security inspection detection method of the rail transit large model according to the embodiment of the present invention; Figure 2 is a processing flowchart of the terahertz target detection algorithm according to the embodiment of the present invention; Figure 3 is a schematic diagram of the target information embedding operation according to the embodiment of the present invention; Figure 4 is a splicing schematic diagram of multiple target information embeddings according to the embodiment of the present invention; Figure 5 is a network structure diagram of the encoder according to the embodiment of the present invention; Figure 6 is a network structure diagram of the classifier according to the embodiment of the present invention; Figure 7 is a data flow schematic diagram of the encoder-decoder according to the embodiment of the present invention; Figure 8 is a network structure diagram of the decoder according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0022] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0023] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0024] Embodiment 1 As Figure 1 shown, this embodiment discloses a security inspection detection method for a large-scale rail transit model based on multi-source information fusion. The method includes the following steps: S1: Obtain optical video images, X-ray transmission images, and terahertz human body images; In this embodiment, taking three security inspection technologies, namely optical video monitoring, X-ray luggage detection, and terahertz human body security inspection detection technology, as examples to illustrate the security inspection detection method for a large-scale rail transit model based on multi-source information fusion, it is not limited to these three types of security inspection data. The security inspection data of various security inspection technologies can be processed by the method of the present invention to obtain more accurate security inspection results. This embodiment does not specifically limit the quantity and type of the fusion of various security inspection data, and can be flexibly selected according to the actual situation.

[0025] The optical video images are collected through optical video monitoring (video security gates); the X-ray transmission images are collected through X-ray luggage detectors to form images of the internal structure of the luggage; the terahertz human body images are collected through terahertz human body security inspection equipment.

[0026] S2: Use an optical target detection algorithm to detect the optical video images to obtain optical detection target information, use an X-ray target detection algorithm to detect the X-ray transmission images to obtain X-ray detection target information, use a terahertz target detection algorithm to detect the terahertz human body images to obtain terahertz detection target information, and perform an embedding operation on the optical detection target information, X-ray detection target information, and terahertz detection target information and then splice them to obtain fusion target information; In this embodiment, after the optical video images, X-ray transmission images, and terahertz human body images are respectively detected by the corresponding target detection algorithms, the corresponding detection results are obtained. The optical target detection algorithm, X-ray target detection algorithm, and terahertz target detection algorithm all adopt existing target detection algorithms, generally target detection algorithms based on deep learning, and the algorithm content will not be elaborated here.

[0027] The following takes the terahertz target detection algorithm as an example for illustration. AsFigure 2 As shown in Figure 2 , this algorithm uses a target detection convolutional neural network to obtain the position coordinates of the target in the terahertz human body image (i.e., terahertz detection target information). Among them, the result output by the terahertz target detection algorithm is the coordinate information of the target , and the coordinate information respectively represents the category of the target ( ), the upper left corner coordinates in the image ( ), the width of the target ( ), and the height of the target ( ); if there are multiple targets, it is expressed as: . Similarly, the optical target detection algorithm and the X-ray target detection algorithm will obtain target information in the same form, that is, optical detection target information and X-ray detection target information

[0028] Perform an "Embedding" operation on the obtained optical detection target information, X-ray detection target information, and terahertz detection target information respectively. This part "upgrades the dimension" of the position information of the target and maps the one-dimensional vector to a two-dimensional matrix. As Figure 3 shown, for example, is a one-dimensional vector with a vector length of 5. In order to learn more useful information features, each value in the vector is encoded using a one-dimensional vector; assuming that a one-dimensional vector with a length of 6 is used, then the Embedding can transform the vector into a 5*6 two-dimensional matrix

[0029] Perform a concatenation (Concat) operation on the embedded optical detection target information, X-ray detection target information, and terahertz detection target information to obtain fused target information. As Figure 4 shown, for example, the data size of the embedding layer of the optical detection target information is , the data size of the embedding layer of the X-ray detection target information is , the data size of the embedding layer of the terahertz detection target information is , then after the Concat operation, the data size is

[0030] The present invention fuses the detection results of three security inspection methods, and fuses different data sources and different data formats in the same processing framework, encodes and embeds the categories and coordinate information of the detection results of different source data to improve the accuracy of security inspection detection

[0031] ​S3: Input the fused target information into the security inspection evaluation model for processing to obtain the passenger security inspection evaluation result; the security inspection evaluation model includes an encoder and a decoder, the decoder receives the context vector generated by the encoder, and generates output data based on the context vector; In this embodiment, as Figure 5 shown, the encoder Encoder consists of two sub-layer connection structures, successively including a multi-head self-attention sub-layer (Multi-Head Attention) and a feed-forward fully-connected sub-layer (Feed-Forward), and each sub-layer is followed by a normalization layer (LayerNorm) and a residual connection ( Figure 5 the dotted line in). The normalization layer and the residual connection together are called the Add&Norm operation.

[0032] Input the fused target information obtained by splicing into the encoder Encoder for feature extraction. The encoder receives the input matrix and converts it into a context vector of a fixed length (context vector). The context vector is an internal representation of the input sequence, capturing the key features of the input information.

[0033] Furthermore, the encoder Encoder outputs a context vector, and the output of the encoder is classified through a classification network to obtain the passenger security inspection evaluation result, and this result is used as a score for the danger level of the passenger.

[0034] As Figure 6 shown, the classification network is a fully-connected network with an activation function, successively including an input layer, a hidden layer, and an output layer.

[0035] Furthermore, the encoder Encoder outputs a context vector, and the context vector is input into the decoder Decoder to generate output data. As Figure 8 shown, the decoder Decoder consists of three sub-layer connection structures, successively including a masked multi-head self-attention sub-layer (Masked Self-Attention), a multi-head attention sub-layer (Encoder-DecoderAttention), and a feed-forward fully-connected sub-layer (Feed-Forward), and each sub-layer is followed by a normalization layer and a residual connection, which together are called the Add&Norm operation. The decoder receives the context vector generated by the encoder, and then generates output data based on this vector. The output data is a multi-dimensional matrix.

[0036] S4: Input the output data into the large language model for rail transit security inspection for processing to obtain the natural language description corresponding to the passenger security inspection evaluation result.

[0037] In this embodiment, the output data of the decoder Decoder is used to fine-tune and train an existing large language model, so that the large language model is adapted to the application scenario of intelligent rail transit security inspection, and a large language model for rail transit security inspection is obtained.

[0038] In the training stage of fine-tuning the large language model, the detection results of optical video images, X-ray transmission images, and terahertz human body images are known. The detection results are formatted and stored by the training personnel in the form of labels (this is known data that needs to be manually labeled). For example, if a mobile phone is detected in the optical image, the coordinate data of the mobile phone is used as training data. After operations such as Embedding, it is input into the security inspection evaluation model. Then the label of this piece of data is "holding a mobile phone in the hand"; the label and the output of the Decoder decoding network are the inputs for fine-tuning the large language model. This is a supervised learning method for neural networks (not limited to large model training).

[0039] The fine-tuned large language model outputs a natural language description based on the output data. The natural language description corresponds to the passenger security inspection evaluation result output by the classification network, and comprehensively describes the information of the current passenger when passing through multiple security inspection devices in multiple dimensions. For example, the overall output is: The risk coefficient of this passenger is 30%; This passenger is carrying a mobile phone, a bag...; There is a rectangular object at the abdominal position...; There is liquid, a laptop... in the luggage.

[0040] The present invention embeds and learns various security inspection results, designs Embedding embedding to fuse various security inspection information in the neural network, and then inputs the obtained embedding vector into the security inspection evaluation model to evaluate the current passenger, obtaining the passenger security inspection evaluation result, which is used as the score for the passenger's risk level; at the same time, the output of the encoder in the security inspection evaluation model is also input into the fine-tuned large language model (the large language model for rail transit security inspection), obtaining a comprehensive description in multiple dimensions of the information of the current passenger when passing through multiple security inspection devices (i.e., the comprehensive security inspection result). Therefore, the present invention innovatively proposes a network that fuses various security inspection results, comprehensively utilizes various security inspection result information. On the one hand, the fused target information is classified by using a traditional classification network to obtain the classification probability, and on the other hand, the fused target information is used for fine-tuning the natural language large model, and natural language output is performed through the large language model for rail transit security inspection to interpret the network output result.

[0041] The present invention is based on security inspection methods in the current rail transit scenario, including but not limited to technologies such as optical video surveillance, X-ray baggage inspection, terahertz body security inspection equipment, and metal detection; a new detection network is designed; specifically: through an Embedding network, information embedding is performed on multiple independent security inspection results, the information is trained and learned through an encoder-decoder network, and a large language model is used to output security inspection information; an Encoder network is designed to encode and learn all security inspection information and extract feature vectors; a convolutional-based classification network is proposed and innovatively used to classify various encoded security inspection information, and finally the network outputs the evaluation result of this security inspection information; based on the large language model, fine-tuning is performed using the output of the Decoder network, and then a large language model for rail transit security inspection is trained.

[0042] Embodiment 2 This embodiment discloses a large model security inspection detection system for rail transit based on multi-source information fusion, including: A data acquisition module, which is configured to: acquire optical video images, X-ray transmission images, and terahertz body images; A multi-source information fusion module, which is configured to: use an optical target detection algorithm to detect the optical video images to obtain optical detection target information, use an X-ray target detection algorithm to detect the X-ray transmission images to obtain X-ray detection target information, use a terahertz target detection algorithm to detect the terahertz body images to obtain terahertz detection target information, perform an embedding operation on the optical detection target information, X-ray detection target information, and terahertz detection target information, and then splice them to obtain fusion target information; A security inspection evaluation module, which is configured to: input the fusion target information into a security inspection evaluation model for processing to obtain a passenger security inspection evaluation result; the security inspection evaluation model includes an encoder and a decoder, and the decoder receives the context vector generated by the encoder and generates output data based on the context vector; A security inspection description module, which is configured to: input the output data into a large language model for rail transit security inspection for processing to obtain a natural language description corresponding to the passenger security inspection evaluation result.

[0043] Embodiment 3 The purpose of this embodiment is to provide a computing device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method in Embodiment 1 are implemented.

[0044] Embodiment 4 The purpose of this embodiment is to provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it performs the steps of the method in Embodiment 1.

[0045] The steps involved in the devices in the above Embodiments 3 and 4 correspond to those in Method Embodiment 1. For specific implementation manners, reference may be made to the relevant description part of Embodiment 1. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0046] Those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0047] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0048] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that based on the technical solutions of the present invention, various modifications or deformations that can be made without creative efforts by those skilled in the art are still within the protection scope of the present invention.

Claims

1. A large-scale rail transit model security inspection method based on multi-source information fusion, characterized in that: include: Acquire optical video images, X-ray transmission images, and terahertz human body images; The optical video image is detected by an optical target detection algorithm to obtain optical detection target information, the X-ray transmission image is detected by an X-ray target detection algorithm to obtain X-ray detection target information, the terahertz human body image is detected by a terahertz target detection algorithm to obtain terahertz detection target information, and the optical detection target information, the X-ray detection target information and the terahertz detection target information are embedded and spliced ​​to obtain fused target information; Inputting the fusion target information into the security inspection assessment model for processing to obtain a passenger security inspection assessment result; The security inspection assessment model includes an encoder and a decoder, wherein the decoder receives a context vector generated by the encoder and generates output data based on the context vector; The output data is input into a large language model for rail transit security inspection for processing to obtain a natural language description corresponding to the passenger security inspection evaluation result.

2. The rail transit large model security inspection detection method based on multi-source information fusion as claimed in claim 1 is characterized in that: Embedding is used to embed the optical detection target information, the X-ray detection target information and the terahertz detection target information respectively, and the one-dimensional vector is mapped into a two-dimensional matrix to obtain a corresponding matrix.

3. The rail transit large model security inspection detection method based on multi-source information fusion as claimed in claim 1 is characterized in that: The encoder consists of two sub-layer connection structures, including a multi-head self-attention sub-layer and a feed-forward fully connected sub-layer in sequence, and each sub-layer is followed by a normalization layer and a residual connection.

4. The rail transit large model security inspection detection method based on multi-source information fusion as claimed in claim 3 is characterized in that: The encoder receives the fusion target information and outputs a context vector, and inputs the context vector into a classification network to obtain a passenger security inspection evaluation result.

5. The rail transit large model security inspection detection method based on multi-source information fusion as claimed in claim 4 is characterized in that: The classification network is a fully connected network with an activation function, which includes an input layer, a hidden layer and an output layer in sequence.

6. The rail transit large model security inspection detection method based on multi-source information fusion as claimed in claim 1 is characterized in that: The decoder consists of a three-sublayer connection structure, including a masked multi-head self-attention sublayer, a multi-head attention sublayer, and a feedforward fully connected sublayer, each of which is followed by a normalization layer and a residual connection.

7. The rail transit large model security inspection detection method based on multi-source information fusion as claimed in claim 1 is characterized in that: The large language model is fine-tuned based on the output data of the decoder to obtain a large language model for rail transit security inspection.

8. A large-scale rail transit model security inspection system based on multi-source information fusion is characterized by: include: A data acquisition module, which is configured to: acquire optical video images, X-ray transmission images, and terahertz human body images; A multi-source information fusion module is configured to: detect the optical video image using an optical target detection algorithm to obtain optical detection target information, detect the X-ray transmission image using an X-ray target detection algorithm to obtain X-ray detection target information, detect the terahertz human body image using a terahertz target detection algorithm to obtain terahertz detection target information, and embed and splice the optical detection target information, the X-ray detection target information, and the terahertz detection target information to obtain fused target information; A security inspection and evaluation module is configured to: input the fusion target information into a security inspection and evaluation model for processing to obtain a passenger security inspection and evaluation result; The security inspection assessment model includes an encoder and a decoder, wherein the decoder receives a context vector generated by the encoder and generates output data based on the context vector; The security inspection description module is configured to: input the output data into the fine-tuned language model for processing to obtain a natural language description corresponding to the passenger security inspection evaluation result.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the rail transit large model security inspection and detection method based on multi-source information fusion as described in any one of claims 1 to 7 are implemented.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the rail transit large model security inspection and detection method based on multi-source information fusion as described in any one of claims 1-7 are implemented.