Multi-modal question and answer model training method and system in urban rail transit field
By collecting data in the field of urban rail transit and extracting images and text features, building an image-text problem pair, using XLNet and ResNet models for feature fusion, training multimodal question-and-answer model, solving the problem of insufficient understanding of professional terms in traditional models, and improving the accuracy and robustness of the question-and-answer system.
Patent Information
- Application Number
- CN202510446700.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-25
AI Technical Summary
The information query method of traditional large-language models in the field of urban rail transit is difficult to meet passengers' requirements for real-time information, service quality and personalized experience, especially in the field of transportation, which lacks professional terms for related content, resulting in inaccurate or unrelated output.
By collecting data from the urban rail transit field, extracting images and text features, building image-text question pairs, and using XLNet and ResNet models for feature extraction and fusion, the multimodal question-and-answer model is trained to answer user questions.
It improves the accuracy and robustness of the Q&A system, provides richer and multidimensional information representation, especially in complex tasks, making up for the lack of information in a single mode.
Smart Images

Figure CN120372033A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of urban rail transit, and particularly to a method and system for training a multi-modal question-answering model in the field of urban rail transit. Background Art
[0002] In recent years, with the acceleration of the urbanization process, urban rail transit, as an important mode of public transportation, has become an important means to alleviate urban traffic congestion, reduce energy consumption and environmental pollution. With the continuous expansion and complexity of the rail transit system, the diversity and immediacy requirements of passengers' information needs are increasing day by day. Traditional information query methods are difficult to meet the requirements of passengers for real-time information, service quality and personalized experience. Therefore, it has become increasingly important to develop an efficient and intelligent question-answering system.
[0003] Traditional large language models (LLMs) rely on advanced machine learning techniques based on deep neural networks. They learn the semantic and syntactic structures of languages by analyzing a large amount of text data. These models perform well in natural language processing (NLP) tasks and can be used in various applications such as text generation, translation, and question-answering systems. Most large language models are pre-trained on general corpora of multiple topics. Although these corpora are rich, they often lack in-depth understanding of content related to the transportation field. For example, the models may have insufficient knowledge of professional terms in aspects such as traffic signs, road condition assessment, and changes in traffic regulations, resulting in inaccurate or irrelevant outputs. Summary of the Invention
[0004] To solve the above problems, the purpose of the embodiments of the present invention is to provide a method and system for training a multi-modal question-answering model in the field of urban rail transit.
[0005] A method for training a multi-modal question-answering model in the field of urban rail transit includes:
[0006] Step 1: Collect data in the field of urban rail transit;
[0007] Step 2: Extract images and texts from the data in the field of urban rail transit;
[0008] Step 3: Construct image-text question pairs based on the extracted images and texts;
[0009] Step 4: Extract image features and text features of the image-text question pairs;
[0010] Step 5: Fuse the image features and text features to form fused features;
[0011] Step 6: Use the fused features to train a neural network model to obtain a multi-modal question-answering model;
[0012] Step 7: Use the multi-modal Q&A model to answer the user's questions in the field of urban rail transit.
[0013] Preferably, the step 4: Extract the image features and text features of the image-text question pair, including:
[0014] Step 4.1: Input the text in the image-text question pair into the XLNet model to extract the initial text features;
[0015] Step 4.2: Input the initial text features into the text feature network to extract the global text features;
[0016] Step 4.3: Enhance the features of the image in the image-text question pair to form an image with enhanced features;
[0017] Step 4.4: Input the image with enhanced features into the ResNet model to extract the image features.
[0018] Preferably, the step 4.2: Input the initial text features into the text feature network to extract the global text features, including:
[0019] The initial text features are input into different convolutional kernels to extract feature vectors, and then the extracted feature vectors are sequentially input into the max-pooling layer and the Dropout layer to obtain the global text features; among them, the global text feature extraction formula is:
[0020] G = dropout(H)
[0021]
[0022] In the formula, D1, D2, and D3 represent the values of the feature vectors extracted by 3 different convolutional kernels after passing through the max-pooling layer, and G represents the extracted global text features.
[0023] Preferably, the step 4.3: Enhance the features of the image in the image-text question pair to form an image with enhanced features, including:
[0024] Step 4.3.1: Take an n*n window centered on any point on the image;
[0025] Step 4.3.2: Calculate the distances between the central pixel point and other pixel points within the n*n window;
[0026] Step 4.3.2: Calculate the weight of each pixel point based on the distances between the central pixel point and other pixel points;
[0027] Step 4.3.4: Perform weighted summation on the pixel points within the n*n window according to the weight of each pixel point to obtain the output pixel value.
[0028] Preferably, step 4.3.2: calculating the weight of each pixel point based on the distance between the central pixel point and other pixel points includes:
[0029] Calculating the entropy of the image, and calculating the weight of each pixel point based on the entropy of the image and the distance between the central pixel point and other pixel points;
[0030]
[0031] Among them, g(i,j) represents the weight of the pixel point (i,j), d(i,j) represents the distance between the central pixel point and the pixel point (i,j), α represents the entropy of the image, N represents the gray level of the image, and p(i) represents the probability that a pixel point with a pixel value of m appears in the image.
[0032] Preferably, step 4.3.4: performing weighted summation on the pixel points within the n*n window according to the weight of each pixel point to obtain the output pixel value, includes:
[0033] Using the formula:
[0034]
[0035] Performing weighted summation on the pixel points within the n*n window to obtain the output pixel value; where Z(i,j) represents the pixel points within the n*n window, Represents the output pixel point.
[0036] Preferably, it further includes:
[0037] Using the formula:
[0038]
[0039] Evaluating the feature-enhanced image, and when the evaluation value is not within the preset range, readjusting the size of the n*n window until the evaluation value is within the preset range; where MSE represents the first evaluation value, AG represents the second evaluation value, SD represents the third evaluation value, W represents the length of the feature-enhanced image, H represents the width of the feature-enhanced image, I(x,y) represents the original image, R(x,y) represents the feature-enhanced image, Represents the gradient of the feature-enhanced image in the x-axis direction, Represents the gradient of the feature-enhanced image in the y-axis direction, avg represents the pixel mean of the feature-enhanced image, and gray(x,y) represents the pixel value on the feature-enhanced image.
[0040] The present invention also provides a multi-modal question-answering model training system in the field of urban rail transit, including:
[0041] A data acquisition module, configured to collect data in the field of urban rail transit;
[0042] A data extraction module, configured to extract images and texts from the data in the field of urban rail transit;
[0043] A sample construction module, configured to construct image-text question pairs according to the extracted images and texts;
[0044] A feature extraction module, configured to extract image features and text features of the image-text question pairs;
[0045] A fusion module, configured to fuse the image features and text features to form fused features;
[0046] A training module, configured to train a neural network model using the fused features to obtain a multi-modal question answering model;
[0047] A question answering module, configured to answer questions of users in the field of urban rail transit using the multi-modal question answering model.
[0048] The present invention also provides an electronic device, including a bus, a transceiver, a memory, a processor, and a computer program stored on the memory and executable on the processor. The transceiver, the memory, and the processor are connected through the bus. It is characterized in that when the computer program is executed by the processor, the steps in the above-mentioned multi-modal question answering model training method in the field of urban rail transit are implemented.
[0049] The present invention also provides a computer-readable storage medium, on which a computer program is stored. It is characterized in that when the computer program is executed by a processor, the steps in the above-mentioned multi-modal question answering model training method in the field of urban rail transit are implemented.
[0050] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:
[0051] The present invention relates to a multi-modal question answering model training method in the field of urban rail transit. Compared with the prior art, by fusing image and text features, the present invention can provide a richer and more multi-dimensional information representation. Through multi-modal fusion, the performance of the model can be significantly improved, especially in complex tasks. Cross-modal knowledge can often make up for the information deficiency of a single modality, thereby improving the accuracy and robustness of the question answering system.
[0052] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following preferred embodiments are specifically given below, and detailed descriptions are made in conjunction with the accompanying drawings as follows. Description of the Drawings
[0053] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0054] Figure 1 Flowchart of the multi-modal question-answering model training method in the field of urban rail transit provided by the present invention;
[0055] Figure 2 System diagram of the multi-modal question-answering model training in the field of urban rail transit provided by the present invention. Detailed implementation manners
[0056] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.
[0057] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined.
[0058] In the present invention, unless otherwise clearly defined and limited, the terms "mounted", "connected", "connected to", "fixed", etc. should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0059] Please refer to Figure 1 , the multi-modal question-answering model training method in the field of urban rail transit includes:
[0060] Step 1: Collect data in the field of urban rail transit;
[0061] The data collected in the field of urban rail transit in the present invention includes papers, research reports, patents, etc. related to urban rail transit.
[0062] Step 2: Extract images and texts from the data in the field of urban rail transit;
[0063] Step 3: Construct image-text question pairs based on the extracted images and texts;
[0064] In practical applications, the present invention first needs to extract images and texts from the collected data in the field of urban rail transit. Next, the present invention needs to construct image-text question pairs in the required order according to the user's question habits.
[0065] Step 4: Extract the image features and text features of the image-text question pairs;
[0066] Further, the said Step 4 includes:
[0067] Step 4.1: Input the text in the image-text question pair into the XLNet model to extract initial text features;
[0068] The XLNet model is a pre-trained language model based on the combination of autoregression and autoencoding. Its core part is the Transformer-XL architecture, which solves the defect of the input length constraint of the context of the BERT model (Bidirectional Encoder Representations from Transformers), so that the pre-trained model can learn more distant context information.
[0069] Step 4.2: Input the initial text features into the text feature network to extract global text features;
[0070] In the said Step 4.2, it includes:
[0071] Input the initial text features into different convolutional kernels to extract feature vectors, and then input the extracted feature vectors into the max pooling layer and the Dropout layer in sequence to obtain global text features; where the global text feature extraction formula is:
[0072] G = dropout(H)
[0073]
[0074] In the formula, represents the values of the feature vectors extracted by 3 different convolutional kernels after passing through the max pooling layer, and G represents the extracted global text features.
[0075] Step 4.3: Enhance the features of the image in the image-text question pair to form an image with enhanced features;
[0076] Among them, Step 4.3 includes:
[0077] Step 4.3.1: Take an n*n window centered at any point on the image;
[0078] Step 4.3.2: Calculate the distances between the central pixel point and other pixel points within the n*n window;
[0079] Step 4.3.2: Calculate the weight of each pixel point based on the distances between the central pixel point and other pixel points;
[0080] In Step 4.3.2, calculate the entropy of the image, and calculate the weight of each pixel point based on the entropy of the image and the distances between the central pixel point and other pixel points;
[0081]
[0082] Among them, g(i,j) represents the weight of the pixel point (i,j), d(i,j) represents the distance between the central pixel point and the pixel point (i,j), α represents the entropy of the image, N represents the gray level of the image, and p(i) represents the probability that the pixel point with pixel value m appears in the image.
[0083] Step 4.3.4: Perform weighted summation on the pixel points within the n*n window according to the weights of each pixel point to obtain the output pixel value.
[0084] In the above steps, the present invention can adopt the formula:
[0085]
[0086] Perform weighted summation on the pixel points within the n*n window to obtain the output pixel value; among them, Z(i,j) represents the pixel points within the n*n window, represents the output pixel point.
[0087] By calculating the distances between the central pixel and the surrounding pixels to adjust the weights, the algorithm can have stronger robustness to noise. The influence of noise points farther from the center point on the output pixel value will be reduced, thereby enhancing the quality of the image.
[0088] After Step 4.3.4, the present invention also needs to adopt the formula:
[0089]
[0090] Evaluate the image after feature enhancement. When the evaluation value is not within the preset range, readjust the window size of n*n until the evaluation value is within the preset range; where MSE represents the first evaluation value, AG represents the second evaluation value, SD represents the third evaluation value, W represents the length of the image after feature enhancement, H represents the width of the image after feature enhancement, I(x,y) represents the original image, and R(x,y) represents the image after feature enhancement. represents the gradient of the image in the x-axis direction. represents the gradient of the image in the y-axis direction. avg represents the pixel mean of the image after feature enhancement, and gray(x,y) represents the pixel value on the image after feature enhancement.
[0091] Step 4.4: Input the image after feature enhancement into the ResNet model to extract image features.
[0092] Step 5: Fuse the image features and text features to form the fused features.
[0093] Concatenate the image features and text features to form the fused features; where the fused features are: C is the image feature, and G represents the extracted global text feature.
[0094] Step 6: Use the fused features to train a neural network model to obtain a multi-modal question answering model.
[0095] Step 7: Use the multi-modal question answering model to answer questions from users in the field of urban rail transit.
[0096] The present invention can provide a richer and more multi-dimensional information representation by fusing image and text features. Through multi-modal fusion, the performance of the model can be significantly improved, especially in complex tasks. Cross-modal knowledge can often make up for the information deficiency of a single modality, thereby improving the accuracy and robustness of the question answering system.
[0097] Please refer to Figure 2 , the present invention also provides a multi-modal question answering model training system in the field of urban rail transit, including:
[0098] A data acquisition module for collecting data in the field of urban rail transit;
[0099] A data extraction module for extracting images and texts from the data in the field of urban rail transit;
[0100] A sample construction module for constructing image-text question pairs according to the extracted images and texts;
[0101] A feature extraction module for extracting image features and text features of the image-text question pairs;
[0102] A fusion module, configured to fuse image features and text features to form fused features;
[0103] A training module, configured to train a neural network model using the fused features to obtain a multi-modal question answering model;
[0104] A question answering module, configured to answer questions of users in the field of urban rail transit using the multi-modal question answering model.
[0105] Compared with the prior art, the beneficial effects of a training system for a multi-modal model in the field of urban rail transit provided by the present invention are the same as those of the multi-modal question answering model training method in the field of urban rail transit described in the above technical solution, and will not be elaborated here.
[0106] The present invention further provides an electronic device, including a bus, a transceiver, a memory, a processor, and a computer program stored on the memory and executable on the processor. The transceiver, the memory, and the processor are connected through the bus. It is characterized in that when the computer program is executed by the processor, the steps in the above multi-modal question answering model training method in the field of urban rail transit are implemented.
[0107] The present invention further provides a computer-readable storage medium, on which a computer program is stored. It is characterized in that when the computer program is executed by a processor, the steps in the above multi-modal question answering model training method in the field of urban rail transit are implemented. Compared with the prior art, the beneficial effects of a computer-readable storage medium provided by the present invention are the same as those of the multi-modal question answering model training method in the field of urban rail transit described in the above technical solution, and will not be elaborated here.
[0108] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technical solution that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for training a multimodal question-answering model in the field of urban rail transit, characterized in that Including: Step 1: Collect data in the field of urban rail transit; Step 2: Extract images and texts from the data in the field of urban rail transit; Step 3: Construct image-text question pairs based on the extracted images and texts; Step 4: Extract the image features and text features of the image-text question pairs; Step 5: Fuse the image features and text features to form the fused features; Step 6: Use the fused features to train a neural network model to obtain a multi-modal question answering model; Step 7: Use the multi-modal question answering model to answer the questions of users in the field of urban rail transit.
2. The training method of the multi-modal question answering model in the field of urban rail transit according to claim 1, characterized in that The said Step 4: Extract the image features and text features of the image-text question pairs, including: Step 4.1: Input the text in the image-text question pair into the XLNet model to extract the initial text features; Step 4.2: Input the initial text features into the text feature network to extract the global text features; Step 4.3: Enhance the features of the image in the image-text question pair to form the image with enhanced features; Step 4.4: Input the image with enhanced features into the ResNet model to extract the image features.
3. The training method of the multi-modal question answering model in the field of urban rail transit according to claim 2, wherein The said Step 4.2: Input the initial text features into the text feature network to extract the global text features, including: Input the initial text features into different convolutional kernels to extract feature vectors, and then input the extracted feature vectors into the max pooling layer and the Dropout layer in sequence to obtain the global text features; where the formula for extracting the global text features is: G = dropout(H) In the formula, D1, D2, and D3 represent the values of the feature vectors extracted by 3 different convolutional kernels after passing through the max pooling layer, H represents the concatenated feature vectors, and G represents the extracted global text features.
4. The training method of the multi-modal question answering model in the field of urban rail transit according to claim 3, wherein The said Step 4.3: Enhance the features of the image in the image-text question pair to form the image with enhanced features, including: Step 4.3.1: Take an n*n window with an arbitrary point on the image as the center; Step 4.3.2: Calculate the distances between the central pixel point and other pixel points within the n*n window; Step 4.3.2: Calculate the weights of each pixel point based on the distances between the central pixel point and other pixel points; Step 4.3.4: Perform weighted summation on the pixel points within the n*n window according to the weights of each pixel point to obtain the output pixel value.
5. The training method of the multi-modal question answering model in the field of urban rail transit according to claim 4, wherein The said Step 4.3.2: Calculate the weights of each pixel point based on the distances between the central pixel point and other pixel points, including: Calculate the entropy of the image, and calculate the weights of each pixel point based on the entropy of the image and the distances between the central pixel point and other pixel points; Among them, g(i,j) represents the weight of the pixel point (i,j), d(i,j) represents the distance between the central pixel point and the pixel point (i,j), α represents the entropy of the image, N represents the gray level of the image, and p(i) represents the probability that the pixel point with the pixel value of m appears in the image.
6. The training method of the multi-modal question-answering model in the field of urban rail transit according to claim 5, wherein, The said Step 4.3.4: Perform weighted summation on the pixel points within the n*n window according to the weights of each pixel point to obtain the output pixel value, including: Adopt the formula: The pixel values within an n*n window are weighted and summed to obtain the output pixel value; where Z(i,j) represents the pixel points within the n*n window, represents the output pixel point.
7. The training method of the multi-modal question answering model in the field of urban rail transit according to any one of claims 6, characterized in that Also including: Adopt the formula: Evaluate the feature-enhanced image. When the evaluation value is not within the preset range, readjust the window size of n*n until the evaluation value is within the preset range. Among them, MSE represents the first evaluation value, AG represents the second evaluation value, SD represents the third evaluation value, W represents the length of the feature-enhanced image, H represents the width of the feature-enhanced image, I(x,y) represents the original image, and R(x,y) represents the feature-enhanced image. represents the gradient of the image in the x-axis direction represents the gradient of the image in the y-axis direction, avg represents the pixel mean of the feature-enhanced image, and gray(x,y) represents the pixel value on the feature-enhanced image.
8. A multimodal question-answering model training system in the field of urban rail transit, characterized in that, Including: A data acquisition module for collecting data in the field of urban rail transit; A data extraction module, configured to extract images and texts from the data in the urban rail transit field; A sample construction module, configured to construct image-text question pairs according to the extracted images and texts; A feature extraction module, configured to extract image features and text features of the image-text question pairs; A fusion module, configured to fuse the image features and text features to form fused features; A training module, configured to train a neural network model using the fused features to obtain a multi-modal question answering model; A question answering module, configured to answer questions of users in the urban rail transit field using the multi-modal question answering model.
9. An electronic device, comprising a bus, a transceiver, a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the transceiver, the memory, and the processor are connected via the bus, characterized in that, When the computer program is executed by the processor, it implements the steps in the multi-modal question answering model training method in the urban rail transit field as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps in the multi-modal question answering model training method in the urban rail transit field as described in any one of claims 1-7.
Citation Information
Patent Citations
Method for detecting texts in any shape in natural scene based on multiple polar coordinates
CN112446356A
Robust visual question and answer model training method based on comparative learning
CN116662591A
Remote sensing image cross-modal retrieval method based on language and visual detail feature fusion
CN116775922A
Image joint text sentiment analysis method based on modal fusion graph convolutional network
CN117540023A
Training method and system for urban rail transit multi-modal model
CN119226493A