A multi-modal based maglev operation control intelligent system

By designing a multimodal artificial intelligence model, real-time intelligent interaction and information autonomy of the maglev transportation system were realized, solving the problem that existing systems could not respond to user needs in a timely manner and improving the system's intelligence and safety.

CN119227988BActive Publication Date: 2026-05-29TONGJI UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TONGJI UNIV
Filing Date
2024-01-20
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing maglev transportation systems cannot achieve end-to-end real-time intelligent interaction, cannot respond to system users' needs in a timely manner, and are unable to meet increasingly complex transportation demands.

Method used

Design a multimodal maglev intelligent operation and control system, including a user subsystem, user interface, artificial intelligence subsystem, ubiquitous sensing subsystem and communication subsystem. Through a multimodal artificial intelligence model, achieve the collaborative completion of system-level and vehicle-level safety-critical and non-safety-critical tasks, and achieve deep integration of real-time sensing information with traffic physical entities.

Benefits of technology

It achieves information autonomy, control autonomy, and traffic process autonomy in the end-to-end service process, provides an end-to-end service solution in the field of demanding safety, perceives the correlation between information and traffic physical entities in real time, and improves the intelligence level of the maglev transportation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119227988B_ABST
    Figure CN119227988B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of artificial intelligence and magnetic suspension transportation, and discloses a magnetic suspension transportation control intelligent system based on multi-modalities, which comprises a user subsystem, a user interface, an artificial intelligence subsystem, a ubiquitous perception subsystem, a magnetic suspension transportation physical subsystem and a communication subsystem; the user subsystem is used for submitting transportation demand, user demand and transportation objects; the artificial intelligence subsystem is used for completing system-level safety demanding tasks, vehicle-level safety demanding tasks and vehicle-level non-safety demanding tasks; the ubiquitous perception subsystem provides multi-modal perception information of the user subsystem and the artificial intelligence subsystem; the magnetic suspension transportation physical subsystem provides physical reality for completing the transportation demand; and the communication subsystem provides ubiquitous interconnection basis for other subsystems. The application effectively promotes the potential implementation of multi-modal artificial intelligence in the safety demanding field, accelerates the intelligent transformation of the magnetic suspension transportation system, and promotes the deep integration of computation, information and physics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence technology and magnetic levitation transportation technology, and in particular to a multimodal magnetic levitation operation and control intelligent system. Background Technology

[0002] Maglev transportation relies on electromagnetic control to achieve train levitation, traction, and braking. Compared to traditional wheeled and wheel-rail transportation, it offers advantages such as contactless wear, safety and reliability, low vibration and noise, and a wider operating speed range. As China's transportation network becomes increasingly complex and sophisticated, traditional transportation methods are facing bottlenecks in meeting future transport demands, making the development of maglev transportation systems urgently needed. With the gradual maturation of supporting technologies for maglev transportation, such as superconducting technology, automatic control technology, and communication technology, maglev transportation is gradually becoming a key focus of next-generation national infrastructure construction. Furthermore, while focusing on increasing the operating speed of maglev trains, improving the intelligence level of the transportation system also requires attention.

[0003] OpenAI's Large Language Model (LLM), ChatGPT, demonstrates superior language understanding, generation, and knowledge reasoning capabilities, making it a next-generation knowledge representation and retrieval method following databases and search engines. ChatGPT largely overcomes the problems of poor robustness and lack of reasoning ability in deep learning models, providing targeted services based on user needs, including copywriting and code writing. Its Infrastructure-as-a-Service (IaaS) operating model signifies that large language models are gradually moving towards Artificial General Intelligence (AGI), which is defined as human-like intelligence capable of autonomous learning and evolution, developing self-awareness, possessing human-like emotions, and solving various complex problems and completing corresponding tasks in different domains.

[0004] The emergence of general artificial intelligence (GA) models is driving AI technology into a new stage of development. The controllability and commercialization of GA to assist various fields, especially those with stringent safety requirements, will be a key focus of future research. For maglev transportation systems, its application presents both a developmental technological empowerment and a challenge in terms of integration. Leveraging GA and coupling it with diverse information technologies to assist in the design of intelligent systems, thereby empowering maglev physical systems, is an urgent need to create a new industry ecosystem, improve intelligence levels, and meet the needs of more potentially complex transportation scenarios.

[0005] Existing maglev transportation systems are based on pre-written programs, which cannot provide end-to-end (demand-service end) real-time intelligent interaction for system users, cannot respond to system user needs in a timely manner, and are difficult to meet the increasingly complex maglev transportation needs. Summary of the Invention

[0006] In response to increasingly complex transportation scenarios, this application provides a multimodal maglev intelligent operation and control system, aiming to promote the integration of artificial intelligence computing processes with maglev transportation physical processes.

[0007] This invention provides a multimodal maglev intelligent operation and control system, which is an intelligent maglev transportation system, comprising: a user subsystem, a user interface, an artificial intelligence subsystem, a ubiquitous sensing subsystem, a maglev transportation physical subsystem, and a communication subsystem; wherein:

[0008] The user subsystem is used to submit transportation requests, user requests, and transportation objects, including system users and transportation objects.

[0009] The user interface is used to enable interaction and information exchange between the user subsystem and the artificial intelligence subsystem;

[0010] The artificial intelligence subsystem is used to complete system-level safety-critical tasks, vehicle-level safety-critical tasks, and vehicle-level non-safety-critical tasks, including system-level multimodal dedicated artificial intelligence MDAI-MS (Multimodal Dedicated Artificial Intelligence-Maglev System), vehicle-level multimodal dedicated artificial intelligence MDAI-MV (Multimodal Dedicated Artificial Intelligence-Maglev Vehicle), and vehicle-level multimodal general artificial intelligence MAGI-MV (Multimodal Artificial General Intelligence-Maglev Vehicle);

[0011] The ubiquitous sensing subsystem collects real-time multimodal information from the user subsystem and the magnetic levitation transportation physical subsystem on large spatial and temporal scales, and transmits it to the artificial intelligence subsystem.

[0012] The magnetic levitation physical subsystem is used to realize the physical reality of the transportation needs of system users;

[0013] The communication subsystem provides a ubiquitous interconnection foundation for the sensing interaction and data transmission between the subsystems of the maglev intelligent operation and control system.

[0014] The advantages of this invention include:

[0015] This invention provides an intelligent magnetic levitation transportation system, a potential realization of artificial intelligence technology in magnetic levitation transportation systems, real-time perception information is associated with traffic physical entities, general artificial intelligence and special artificial intelligence cooperate to complete non-safety-critical and safety-critical tasks of the magnetic levitation transportation system, realize information autonomy, control autonomy and traffic process autonomy in the end-to-end service process, and provide a system solution for the application of large-scale models to end-to-end services in the field of safety-critical requirements.

[0016] This invention provides a scheme for mapping the generalized knowledge inherent in general artificial intelligence to the physical reality of a magnetic levitation transportation system. Specifically, it maps the uncertain expression of user needs and ubiquitous perception information in the semantic space to the deterministic expression of control commands in the semantic space, and then maps them to the deterministic expression of mechanical equipment in the behavioral space. The model computation and reasoning process and the physical process influence each other and interact in real time, realizing a deep integration of computation, information and physics. Attached Figure Description

[0017] Figure 1 This is a schematic diagram illustrating the specific structure and connection relationships of a multimodal maglev motion control intelligent system according to this application;

[0018] Figure 2 This is a schematic diagram of the network structure of the multimodal artificial intelligence model in this application;

[0019] Figure 3 This is a schematic diagram illustrating the specific structure of the multimodal information interaction layer in the multimodal artificial intelligence model of this application;

[0020] Figure 4 This is a schematic diagram illustrating the training steps of the multimodal artificial intelligence model in this application:

[0021] Figure 5 This is a schematic diagram illustrating the structure and connection relationships of the communication subsystem of this application;

[0022] Figure 6 This is a schematic diagram of the operation scenario of the maglev intelligent operation and control system provided in the implementation of this application;

[0023] Figure 7 This is a schematic diagram of the operation process of the intelligent maglev operation and control system provided in the implementation of this application.

[0024] Appendix Figure 1 In the text, the meanings of each symbol are as follows:

[0025] Icons: 1 represents the user subsystem; 2 represents the user interface; 3 represents the artificial intelligence subsystem; 4 represents the ubiquitous sensing subsystem; 5 represents the magnetic levitation physical subsystem; 6 represents the communication subsystem.

[0026] System-level multimodal dedicated artificial intelligence MDAI-MS (Multimodal Dedicated Artificial Intelligence-Maglev System);

[0027] Multimodal Dedicated Artificial Intelligence-Maglev Vehicle (MDAI-MV)

[0028] Multimodal Artificial General Intelligence-Maglev Vehicle (MAGI-MV) is a vehicle-level multimodal general artificial intelligence. Detailed Implementation

[0029] The technical solutions provided in this application will be further described below with reference to specific embodiments and accompanying drawings. The advantages and features of this application will become clearer from the following description.

[0030] like Figure 1 As shown, a multimodal maglev operation control intelligent system is an intelligent maglev transportation system, including: a user subsystem (1), a user interface (2), an artificial intelligence subsystem (3), a ubiquitous sensing subsystem (4), a maglev transportation physical subsystem (5), and a communication subsystem (6).

[0031] The user subsystem (1) includes system users and transport objects. The system users submit transport requests to the maglev transportation system through the user interface. The system users include, but are not limited to, real users (humans) or virtual users (artificial intelligence models). The transport objects are the contents transported by the maglev transportation system and are defined by the system users, including, but not limited to, humans or goods.

[0032] Furthermore, the transportation demand is presented in text or voice modality, and the content of the transportation demand includes at least the object of transportation, the place of origin of transportation, and the destination of transportation, and may also include transportation time limits, etc.

[0033] The user interface (2) enables interaction and information exchange between the user subsystem (1) and the artificial intelligence subsystem (3). The interface forms include, but are not limited to, software interfaces and brain-computer interfaces.

[0034] The artificial intelligence subsystem (3) includes system-level multimodal dedicated artificial intelligence MDAI-MS (Multimodal Dedicated Artificial Intelligence-Maglev System), vehicle-level multimodal dedicated artificial intelligence MDAI-MV (Multimodal Dedicated Artificial Intelligence-Maglev Vehicle), and vehicle-level multimodal general artificial intelligence MAGI-MV (Multimodal Artificial General Intelligence-Maglev Vehicle);

[0035] Furthermore, the MDAI-MS serves the entire maglev transportation system and is deployed in the maglev transportation system control center to complete the stringent system-level safety requirements involved in the operation of the maglev transportation system.

[0036] Specifically, the system-level security requirements include: scheduling maglev vehicles according to transportation needs, assigning maglev vehicle IDs to users, solving for the optimal route from the origin to the destination, calculating the static running curve from the origin to the destination, and transmitting the generated information to the corresponding MDAI-MV in the form of messages.

[0037] Specifically, the MDAI-MS is built on a neural network with a decoder-encoder architecture. It performs self-supervised pre-training with a large amount of unlabeled multimodal data (text, speech, video, etc.) and supervised fine-tuning with labeled multimodal data. It particularly focuses on learning general and professional knowledge in fields such as vehicles, transportation, control, communication, computers, and electrical engineering. It has semantic understanding and logical reasoning capabilities, and its generation strategy is guided by stringent safety requirements based on reinforcement learning with human feedback. Downstream task interfaces are set according to the functions of MDAI-MS, including: transportation demand parsing interface, static operation curve generation interface, etc.

[0038] Furthermore, the MDAI-MV corresponds one-to-one with the maglev vehicles included in the maglev transportation system, and is deployed on the onboard computer of the corresponding maglev vehicle, such as the MDAI-MV. i Serving the i-th maglev vehicle, it is used to complete vehicle-level safety requirements related to transportation needs;

[0039] Specifically, the vehicle-level safety requirement task refers to MDAI-MV i Based on information from MDAI-MS and real-time perception information from the ubiquitous perception subsystem, and based on communication technology and MDAI-MV of the preceding neighboring vehicle. iInteractive information is used to dynamically adjust the static operation curve, generate a dynamic operation curve, and convert the kinematic parameters such as speed and acceleration of the maglev car into motor voltage and current, generate control commands and send them to the corresponding safety-critical equipment, as well as identify the transport object and monitor the behavior of the transport object.

[0040] Specifically, the safety-critical equipment includes, but is not limited to, doors, traction motors, levitation electromagnets, guide electromagnets, brakes, and turnout mechanisms.

[0041] Specifically, the MDAI-MV i Based on neural network construction and with decoder-encoder as the basic architecture, it performs self-supervised pre-training through a large amount of unlabeled multimodal data, with a particular focus on learning general and professional knowledge in fields such as vehicles, transportation, control, communication, computer, and electrical engineering. It has semantic understanding and logical reasoning capabilities, and its generation strategy is guided by stringent safety requirements based on reinforcement learning with human feedback. Downstream task interfaces are set according to the MDAI-MS function settings, including: message parsing interface, dynamic operation curve generation interface, and control command generation interface.

[0042] Furthermore, the MAGI-MV also corresponds one-to-one with the maglev vehicles included in the maglev transportation system, and is deployed in the onboard computer of the corresponding maglev vehicle, such as MAGI-MV. i Serving the i-th maglev car, and MAGI-MV i With MDAI-MV i Together they serve the i-th maglev vehicle to complete vehicle-level non-safety-critical tasks involved in the transportation process;

[0043] Specifically, the vehicle-level non-safety-critical tasks can come from system requirements or system user needs, and are defined as tasks that do not affect the driving safety of the maglev vehicle, including user need reasoning, user non-safety-critical requirements, and non-safety-critical on-board equipment control, etc.

[0044] Specifically, the non-safety-critical in-vehicle equipment includes, but is not limited to, lighting, air conditioning, and audio;

[0045] Specifically, the MAGI-MV i Based on neural network construction and with decoder-encoder as the basic architecture, it performs self-supervised pre-training through a large amount of unlabeled multimodal data to learn general knowledge in fields such as humans, society, environment, vehicles, medicine, and transportation. It has semantic understanding and logical reasoning capabilities, and reinforcement learning based on human feedback ensures that its generation strategy does not violate ethics. Downstream task interfaces are set according to the functions of MDAI-MS, including: user requirement parsing interface, control command generation interface, and multimodal signal generation interface.

[0046] It should be noted that system users can only access MDAI-MS and MAGI-MV allocated by MDAI-MS according to transportation needs through the user interface. i Unable to access MDAI-MV.

[0047] The ubiquitous sensing subsystem (4) collects real-time multimodal information from the user subsystem (1) and the magnetic levitation transportation physical subsystem (5) on a large spatial and temporal scale, and transmits it to the artificial intelligence subsystem (3). It consists of multiple sensing submodules, including at least a first sensing module, a second sensing module, a third sensing module, and a fourth sensing module; wherein:

[0048] The first sensing module is used to estimate the speed and spatiotemporal position of the maglev vehicle, and consists of a speed sensor, a synchronization clock, a positioning sensor, and satellite navigation.

[0049] The second sensing module, used to estimate the guide gap and suspension gap of the maglev vehicle, consists of a guide gap sensor and a suspension gap sensor.

[0050] The third sensing module is used to collect dynamic shape information of objects transported inside the maglev vehicle, track lines, and external environment. It consists of an in-vehicle 3D depth camera, a drone vision camera, an infrared vision camera, and a lidar.

[0051] The fourth sensing module is used to monitor the operation and fault status of equipment such as maglev cars, track lines, turnout mechanisms, motors, and auxiliary sub-equipment. It consists of vibration sensors, acoustic emission sensors, voltage sensors, current sensors, etc.

[0052] Furthermore, the ubiquitous sensing subsystem can be further classified into vehicle-mounted sensing modules, trackside sensing modules, and far-field sensing modules (satellites) based on their deployment location.

[0053] The magnetic levitation physical subsystem (5) is used to fulfill the physical reality of the transportation needs of system users, including but not limited to magnetic levitation vehicles, on-board equipment, track lines, linear motors, turnout mechanisms, traction substations, 5G communication base stations and related auxiliary equipment.

[0054] The communication subsystem (6) provides a ubiquitous interconnection foundation for the perceptual interaction and data transmission between the various subsystems of the maglev intelligent operation and control system to form a ubiquitous interconnection.

[0055] Specifically, the vehicle base station is connected to the vehicle computer, which in turn is connected to the vehicle safety-critical equipment, the vehicle non-safety-critical equipment, and the vehicle sensing module via bus technology. The bus uses fiber optic transmission medium.

[0056] Specifically, the vehicle-mounted base station is connected to the ground base station via 38GHz millimeter-wave communication technology, and the ground base station is wired to the track safety-critical equipment (traction motor, turnout mechanism, etc.) and trackside sensing module via a trackside switch.

[0057] Specifically, the trackside switch is connected to MDAI-MS via IP-based Ethernet communication technology, and information transmission passes through Ethernet switches, routing devices, and firewalls;

[0058] Specifically, the far-field perception module connects to the vehicle computer based on the information transmission specific medium of the far-field perception device. For example, the vehicle computer receives satellite navigation signals through the vehicle antenna and the vehicle receiver.

[0059] Specifically, the user subsystem (1) is connected to MDAI-MS and MAGI-MV via IP-based Ethernet communication technology.

[0060] It should be noted that MAGI-MS, MDAI-MS, and MDAI-MV provided in this application use the same network structure and training method, but differ in the scale of the training dataset, the scale of the model parameters, and the downstream task interface; the network structure is the network structure of a multimodal artificial intelligence model.

[0061] The specific network structure and computational principles of the multimodal artificial intelligence model provided in this application are explained in detail below.

[0062] like Figure 2 As shown, the network structure of a multimodal artificial intelligence model (including dedicated and general-purpose models) consists of a set of m modal signals as network input. The network structure can be divided into an encoder and feature layer for each modal signal, a multimodal information interaction layer, an output layer for each modal signal, and a downstream task interface.

[0063] Taking the m-th mode encoder as an example, the m-th mode encoder discretizes, vectorizes, and positions the m-th mode signal;

[0064] Specifically, there is the m-th mode signal x m As input to a multimodal artificial intelligence model, the m-th modality code is first discretized into N subsamples of equal size, resulting in a sample sequence. N is determined based on the characteristics of the m-th modal signal. For example, when the signal is two-dimensional image data, N = 14, and it is encoded by vector E. m Convert the discrete sample sequence into a computable vector. in Encode the m-th mode signal vector E m The weight parameters are then concatenated with the category label cls. mThe semantic information represented by the sample sequence is used to finally add positional encoding E. p To obtain the output y, we use the positional information of the subsample in the sample sequence. m =[cls m V m ]+E p .

[0065] Taking the m-th modal feature layer as an example, the m-th modal feature layer responds to the output from the m-th modal encoder. Perform feature extraction;

[0066] Specifically, for the m-th mode signal, the output y of the m-th mode encoder is used. m As input, the m-th modality feature layer is decoded by an L-layer Transformer neural network decoder T. en These layers are stacked, with the output of each layer serving as the input to the next layer, i.e., features. in L represents the weight parameters of the j-th layer Transformer neural network decoder in the m-th modal feature layer. L is determined based on the characteristics of the m-th modal signal. For example, when the signal is two-dimensional image data, L = 6. Finally, the m-th modal signal features are output from the m-th modal feature layer.

[0067] It should be noted that the Transformer neural network T is an existing technology, in which one layer of the Transformer neural network consists of a decoder T. en and a layer decoder T de It is stacked together, which will not be described in detail here.

[0068] The multimodal information interaction layer performs modality category encoding, modality fusion, and modality position encoding on the features from each modal signal. Then, the multimodal information feature layer extracts the fused features to achieve a unified expression of the same semantic information in the interaction space for different modal signals.

[0069] Figure 3 The specific structure of the multimodal information interaction layer provided in this application is shown.

[0070] Specifically, the multimodal information interaction layer takes the output of the feature layer of multiple modal signals as input; it is assumed that feature pairs are obtained from n modal signals with the same semantics. The multimodal information interaction layer first defines the characteristics of each modal signal. Add the corresponding modal category tag (mod). j ,get Then, the signal features of each modality mentioned above are sequentially concatenated to obtain the fusion sequence. Then, modal position codes E′ are added to the features of different modal signals.p To characterize the positional information of different modalities in the fused sequence, we obtain The information is then fed into a multimodal information feature layer, which consists of Z layers of stacked Transformer neural networks T. The output of each layer serves as the input to the next layer, i.e., y1′=T(y′;θ1),…,y′ Z =T(y′) Z-1 ,θ Z ), where θ j ,j=1,…,Z represents the weight parameters of the j-th layer of the Transformer neural network in the multimodal interaction layer. In this embodiment, Z=3. Finally, the multimodal feature y′ is output through the multimodal information interaction layer. Z .

[0071] Taking the m-th modal output layer as an example, the m-th modal output layer analyzes the multimodal features from the multimodal interaction layer and extracts the implicit features of the m-th modal signal to achieve cross-modal feature conversion;

[0072] Specifically, assuming the first modal signal is used as the input to the multimodal artificial intelligence model, and multimodal features are obtained by passing through the first modal encoder, the first modal feature layer, and the multimodal information interaction layer, the m-th modal output layer takes the multimodal features as input; and there are multimodal features y′ Z The m-th modal output layer is composed of a G-layer Transformer neural network decoder T. de They are stacked, with the output of each layer serving as the input of the next layer. in G represents the weight parameters of the j-th layer Transformer neural network decoder in the m-th mode output layer. G is determined based on the characteristics of the m-th mode signal. For example, when the signal is two-dimensional image data, G=3. Finally, the m-th mode implicit features are output through the m-th mode output layer. The downstream task interface takes the implicit features output by each modal output layer as input. The downstream task interface can call a single modal output layer or multiple modal output layers to complete specific safety-critical or non-safety-critical tasks, which generally include user requirement analysis tasks, path planning tasks, speed target curve estimation tasks, generative tasks, etc.

[0073] Specifically, assuming the q-th downstream task interface needs to output a signal in the form of the m-th modality, then the implicit features of the m-th modality are used. As input, the q-th downstream task interface can be derived from the B-layer Transformer neural network decoder T. de It is composed of a Q-layer multilayer perceptron M stacked together, with the output of each layer serving as the input of the next layer. in The weight parameters B and Q represent the weight parameters of the j-th layer Transformer neural network decoder or multilayer perceptron in the q-th downstream task interface. B and Q are determined according to the characteristics of the m-th modal signal. For example, when the signal is two-dimensional image data, B=2 and Q=2. Finally, the m-th modal signal is output through the q-th downstream task interface.

[0074] It should be noted that the multilayer perceptron is an existing technology implemented using neural networks, which will not be discussed in detail here.

[0075] Figure 4 The training steps of the multimodal artificial intelligence model provided in this application are illustrated, and a detailed explanation is given below:

[0076] Step S101: Pre-train each modal encoder, each modal feature layer, multimodal information interaction layer, and each modal output layer;

[0077] Specifically, each modal encoder, each modal feature layer, and each modal output is trained only with data consisting of the corresponding modal signals. For example, the first modal encoder and the first modal feature layer are trained only with a large amount of unlabeled data or structured knowledge of the first modality type using a self-supervised method, and the other modal encoders and feature layers are trained in the same way.

[0078] In each modal encoder, only vector encoding E is used. i parameters Training is required, and the training objective is to map the corresponding modal signal into a computable vector suitable for feature extraction by the corresponding modal feature layer; each modal feature layer has L layers of Transformer neural network decoders T. en Weight parameters Training is required, and the training objective is to enable each modal feature layer to fully extract the shallow texture features and deep semantic features of the corresponding modal signal, and to have the ability to generalize to understand the corresponding modal signal.

[0079] Among them, the weight parameters θ of the Z-layer Transformer neural network T in the multimodal information interaction layer j The training process, j = 1, ..., Z, is conducted using a large number of unlabeled, mutually matched sample pairs. For a mutually matched sample pair {x}, ... 1 ,…,x m}, where x i ,i=1,…,m represent samples from the i-th modal signal. Mutual matching means that samples from different modal signals have the same semantic content or can describe each other in pairs. The training objective of the multimodal information interaction layer is to eliminate the inconsistency between different modal features extracted by different modal feature layers based on their respective modal signals for a mutually matched sample pair.

[0080] Among them, the weight parameters of the G-layer Transformer neural network decoder in each modality output layer Training is required, and the training objective is to enable each modal output layer to extract implicit features that can be converted into corresponding modal signals based on the multimodal features output by the multimodal information interaction layer.

[0081] The self-supervised pre-training process for each modal encoder, each modal feature layer, multimodal information interaction layer, and each modal output layer uses predictive loss and contrastive loss functions to enable the multimodal AI model to have cross-modal understanding capabilities;

[0082] Taking the m-th mode signal as an example, there is a sample x m After discretization, the sample sequence is obtained. N is the number of subsamples. The purpose of the prediction loss function is to enable the multimodal AI model to understand single-modal signals; specifically, it refers to maximizing the multimodal AI model's ability to interpret sample sequences. The conditional probability P of the j-th subsample is predicted from the jk-th to the (j-1)-th subsample, as follows:

[0083]

[0084] It should be noted that each modal signal has a corresponding prediction loss function. The overall prediction loss function is as follows:

[0085]

[0086] The purpose of the contrastive loss function is to ensure that the semantic representations of different modal features extracted from mutually matching modal signals have high correlation, while the semantic representations of different modal features extracted from mismatched modal signals have low correlation. This aligns different modal features in the multimodal space, enabling the multimodal AI model to have cross-modal understanding capabilities. Taking the contrastive loss function between the 1st modal signal and the mth modal signal as an example, a batch_size contains b mutually matching sample pairs, as follows:

[0087]

[0088] Here, Sr(·,·) is a function for calculating the semantic relevance of features. The higher the relevance, the higher the Sr value, and vice versa. and These are the implicit features of the 1st and mth mode signals in the i-th sample pair, respectively. and These are the implicit features of the 1st and mth mode signals in the j-th sample pair, respectively;

[0089] Total contrastive loss function It is the sum of the pairwise contrast losses calculated between all modal signals, that is:

[0090]

[0091] It should be noted that step S101 above is the pre-training process of the multimodal artificial intelligence model. The training data can come from web crawling, public databases, public datasets, etc. The scale of the training dataset is on the GB level, the number of model parameters is on the hundreds of millions level, and the order is: MAGI-MS > MDAI-MS > MDAI-MV.

[0092] Step S102: Supervised training of the weight parameters of the downstream task interface, and fine-tuning with instructions to fine-tune the multimodal artificial intelligence model;

[0093] Specifically, based on the type of task the system needs to complete, downstream task interfaces are set up, and the weight parameters of the downstream task interfaces are trained in a supervised manner using labeled training data. The training objective is to ensure that the labels predicted by the multimodal AI model are consistent with the true labels, so as to adapt to downstream tasks in specific domains; for example, classification tasks, given a sample x of the m-th modality. m There are corresponding real labels. m The cross-entropy loss function for this sample is as follows:

[0094]

[0095] Wherein, P(x m For multimodal artificial intelligence models, the sample x m The probability of category prediction;

[0096] Furthermore, instruction tuning is used to fine-tune the weight parameters of the multimodal AI model. Instruction tuning adds extra instructions to the input to guide the multimodal AI model in understanding human intentions. The goal is to help the multimodal AI model better complete downstream tasks in a specific domain through prompts. Given an input x presented in the form of the m-th modality... m After fine-tuning with instructions, the result is Multimodal Artificial Intelligence Model Analysis x m It generates the correct response based on the instructions to complete specific downstream tasks;

[0097] Specifically, for MDAI-MS, let's take user requirements analysis as an example:

[0098] Input: Transportation requirements (in text or voice);

[0099] After fine-tuning the instructions: ① "Transportation request | Please assign maglev car ID", ② "Transportation request, please confirm the transport object, origin, destination, etc.", ③ "Transportation object, origin, destination, etc. | Request to find the optimal route from origin to destination"...

[0100] Predefined answers: ① "Magnetic levitation vehicle ID", ② "Transportation object, origin of transportation, destination of transportation, etc.", ③ "Optimal route from origin of transportation to destination";

[0101] Specifically, for MDAI-MV i Taking vehicle-level safety requirements as an example:

[0102] Input: Section route instructions, static operation curves, and real-time sensing information (in text form);

[0103] After fine-tuning the instructions: ① "Section route instructions, static operating curve, real-time sensing information | Please dynamically adjust the static operating curve based on the sensing information", ② "Section route instructions, optimal speed target curve | Please generate control instructions for safety-critical equipment within the section"...

[0104] Predefined answers: ① "Optimal speed target curve within the section", ② "Control commands for equipment with stringent safety requirements, such as motors";

[0105] Step S103: Constrain the multimodal artificial intelligence model through reinforcement learning with human feedback;

[0106] Reinforcement learning based on human feedback refers to: establishing a reward model that takes actions as input and outputs scalar reward values. During training, the reward model continuously fits the safety-obsessed thinking of human intentions based on human feedback (e.g., EN 50126, EN50128, IEC 61508, etc.). Furthermore, within the reinforcement learning framework, the intelligent maglev transportation system is regarded as the environment, the weight parameters of the multimodal artificial intelligence model are regarded as policies, and the output is regarded as actions. The reward model evaluates the actions and calculates the reward. The loss function aims to maximize the reward, thereby constraining the output of the multimodal artificial intelligence model.

[0107] For example, for MDAI-MS, the boundary conditions are the track line, virtual section settings, rated parameters of the maglev car hardware, and power supply system parameters. The objective function is to optimize traction, levitation, guidance stability, and operating energy consumption. Reinforcement learning with human feedback is used to generate optimal paths and static operating curves, all guided by stringent safety requirements. For MDAI-MV, real-time sensing information is used as dynamic parameters. The model is based on the quantitative relationship between the kinematic parameters of the maglev car and the coil current and voltage. The objective function is to optimize traction, levitation, guidance stability, and operating energy consumption. Reinforcement learning with human feedback is used to generate optimal speed target curves and control commands, all guided by stringent safety requirements.

[0108] The following details the operation flow of the multimodal maglev intelligent control system provided in this application. For example... Figure 6 As shown, taking electromagnetic suspension (EMS) rail transit as an example, the train traction is completed by a linear synchronous motor. The line in the figure crosses 4 power supply zones, each of which consists of 6 traction motors (stator sections). Each traction motor is 2km long, and the maglev train is 100m long. Each traction motor has 2 auxiliary stopping points. In this embodiment, the maglev train needs to run from station A to station B, then from station B to the traction motor before the switch mechanism. After the switch mechanism is turned, it finally runs to station E to complete the transportation requirements.

[0109] like Figure 7 As shown, the operation flow of the multimodal maglev intelligent control system is as follows:

[0110] Step S201: System users submit transportation requests to MDAI-MS through the user interface;

[0111] Specifically, three users submitted transportation requests to the system: x1 = "From station A to station C", x2 = "I boarded at station B, please take me to station C", and x3 = "Deliver the goods from station A to station C". These requests were transmitted to MDAI-MS through the user interface.

[0112] Step S202: MDAI-MS analyzes transportation demand, allocates maglev vehicles, generates the necessary driving information based on physical static parameters, and transmits it to the corresponding MDAI-MV.

[0113] Specifically, MDAI-MS parses the transportation demand to obtain the transportation object, transportation origin, and transportation destination, which are respectively "System User 1, Station A, Station C", "System User 2, Station B, Station C", and "Goods, Station A, Station C". It then dispatches the maglev train 01 to Station A to serve the system users, and the system users obtain access to MAGI-MV. 01With the necessary permissions, MDAI-MS generates the necessary information for train operation based on the physical static parameters from station A to station C, such as track parameters, speed limits, maglev train traction, and braking performance. This information includes the optimal route for maglev train 01 and the static operating curve from the origin to the destination. The MDAI-MS then packages and transmits this information to MDAI-MV. 01 ;

[0114] In addition, the standard 3D shape information of the transported object is uploaded by the system user or retrieved from the 3D face database, and then transmitted from MDAI-MS to MDAI-MV. 01 ;

[0115] Step S203: MDAI-MV uses in-vehicle visual information to determine whether the identity, behavior, location, and other information of the transported object meet the requirements for safe departure of the maglev train;

[0116] Specifically, MDAI-MV 01 The 3D modeling information of objects inside the maglev vehicle is obtained through an in-vehicle 3D depth camera, and it is determined whether the modeling information matches the standard 3D shape information of the transported object. If the match is successful, MDAI-MV... 01 Based on the visual monitoring information provided by the in-vehicle 3D depth camera, the system infers the position information of the transported object, including its behavior and location information, and determines whether it meets the requirements for safe operation of the maglev vehicle.

[0117] Step S204: Close the car door, and MDAI-MV generates the control commands required for the maglev car to drive based on real-time ubiquitous sensing information and sends them to relevant safety-critical equipment to complete the transportation requirements;

[0118] Specifically, MDAI-MV 01 After confirming that the safe departure conditions of the maglev train are met, a control command to close the doors is generated and transmitted to the door controller; MDAI-MV 01 Based on the static operation curves and track resource information provided by MDAI-MS, and real-time ubiquitous sensing information, including track parameters, surrounding weather conditions, maglev train spatiotemporal positioning information, operating speed, and switch mechanism status, dynamic operation curves are generated using a step-by-step approach. Furthermore, based on the mapping relationship between the maglev train's operating speed and the electromagnetic control quantities of the traction motor, kinematic parameters such as maglev train speed and acceleration are converted into electromagnetic control quantities such as traction voltage and current, which are then directly controlled by MDAI-MV. 01 Generate control commands and issue them to the corresponding safety-critical equipment, such as enabling the traction motor to provide the necessary electrical resources or moving the switch to the correct position.

[0119] It should be noted that "step-forward" means that, under normal circumstances, the current movement authorization endpoint of the maglev car is the forward auxiliary stopping point, and the current dynamic operation curve is calculated based on this. If there is a switch mechanism between the maglev car and the forward auxiliary stopping point, the switch mechanism is used as the endpoint of the movement authorization. If there are other maglev cars between the maglev car and the forward auxiliary stopping point, the braking stopping point of the car with the worst braking performance is used as the endpoint of the movement authorization, thus forming a tracking of the car in front. However, the same traction motor is not allowed to provide traction to two maglev cars at the same time, and there must be a traction motor between adjacent cars. In this case, the generation of the dynamic operation curve also needs to consider the information provided by the MDAI-MV of the car in front, including: the running status of the car in front, the running speed of the car in front, the driving acceleration of the car in front, and the destination of the car in front.

[0120] It should be noted that, based on the communication subsystem, the MDAI-MVs of different maglev cars can exchange information with each other through the MDAI-MS as an intermediary, and the control commands of the MDAI-MV can be directly issued to the traction motor in the maglev physical subsystem, as well as the forward-adjacent traction motors, switch mechanisms, etc.

[0121] It should be noted that, based on the vehicle-motor communication strategy, MDAI-MV directly issues control commands to the traction motors where it is located and those in front of it based on the generated dynamic operating curve, so as to reduce the impact of switching power supply zones on the smoothness of operation.

[0122] It should be noted that MDAI-MV only allows communication with the traction motor currently in which the maglev car is located and the traction motor that is adjacent to it in the forward direction. When the maglev car leaves the traction motor, the communication with the traction motor that was left is interrupted, and the electromagnetic control quantity of the traction motor that was left is reset. A traction motor or switch mechanism is only allowed to communicate with the MDAI-MV of one maglev car to avoid control confusion of the traction motor or switch mechanism.

[0123] It should be noted that the format of the control command is: control command identifier + communication protocol identifier + maglev vehicle ID + controlled object IP address + length identifier + content + check code;

[0124] It should be noted that during the operation of the maglev car, the guide gap and suspension gap are stabilized near the set value by the guide electromagnet and suspension electromagnet based on closed-loop control technology.

[0125] Step S205: During the operation of the maglev car, MAGI-MV completes the non-safety-critical requirements submitted by the system users.

[0126] During the operation of the maglev car, the system user submits user requests to the MAGI-MV of the assigned maglev car 01. The MAGI-MV analyzes whether the user requests include safety-critical tasks and only allows the MAGI-MV to execute non-safety-critical tasks in the user requests, such as turning on the on-board air conditioning, audio services, and checking the status of goods.

[0127] Step S206: The maglev vehicle travels to the transportation destination, completes the transportation request submitted by the system user, and closes the system user's access rights to MAGI-MV.

[0128] The above description is merely a description of preferred embodiments of this application and is not intended to limit the scope of this application in any way. Any changes or modifications made by those skilled in the art based on the above-disclosed technical content should be considered as equivalent and valid embodiments and fall within the scope of protection of the technical solution of this application.

Claims

1. A multimodal maglev intelligent operation and control system, which is an intelligent maglev transportation system, characterized in that: include: User subsystem (1), user interface (2), artificial intelligence subsystem (3), ubiquitous sensing subsystem (4), magnetic levitation transportation physical subsystem (5), communication subsystem (6); among which: The user subsystem (1) is used to submit transportation requirements, user requirements and transportation objects, including system users and transportation objects; The user interface (2) is used to realize the interaction and information exchange between the user subsystem (1) and the artificial intelligence subsystem (3); The artificial intelligence subsystem (3) is used to complete system-level safety-critical tasks, vehicle-level safety-critical tasks, and vehicle-level non-safety-critical tasks, including system-level multimodal dedicated artificial intelligence MDAI-MS, vehicle-level multimodal dedicated artificial intelligence MDAI-MV, and vehicle-level multimodal general artificial intelligence MAGI-MV. The ubiquitous sensing subsystem (4) collects real-time multimodal information from the user subsystem (1) and the magnetic levitation transportation physical subsystem (5) on a large spatial and temporal scale, and transmits it to the artificial intelligence subsystem (3). The magnetic levitation transportation physical subsystem (5) is used to realize the physical reality of the transportation needs of system users; The communication subsystem (6) provides a ubiquitous interconnection basis for the perception, interaction and data transmission between the subsystems of the maglev intelligent operation and control system; MAGI-MS, MDAI-MS, and MDAI-MV all use the same network structure, which is the network structure of a multimodal artificial intelligence model. The network input consists of a set of modal signals. The network structure is divided into an encoder and feature layer for each modal signal, a multimodal information interaction layer, an output layer for each modal signal, and a downstream task interface. Specifically: No. Modal signal encoder for the first Modal signals are discretized, vector-encoded, and position-encoded; No. Modal signal feature layer pairs from the first Feature extraction is performed on the output of the modal signal encoder; The multimodal information interaction layer performs modality category encoding, modality fusion, and modality position encoding on the features from each modal signal. Then, the multimodal information feature layer extracts the fused features to achieve a unified expression of the same semantic information in the interaction space for different modal signals. No. The modal signal output layer analyzes the multimodal features from the multimodal interaction layer and extracts the first modal signal. Implicit features of modal signals are used to achieve cross-modal feature conversion; The downstream task interface takes the implicit features output by each modal output layer as input. The downstream task interface calls a single modal output layer or multiple modal output layers to complete specific security-critical or non-security-critical tasks.

2. The multi-modal magnetic levitation intelligent control system as described in claim 1, characterized in that, The artificial intelligence subsystem (3): The MDAI-MS serves the entire maglev transportation system and is deployed in the maglev transportation system control center to complete the stringent system-level safety requirements involved in the operation of the maglev transportation system. The MDAI-MS is built on a neural network with a decoder-encoder architecture. It is self-supervised pre-trained using a large amount of unlabeled multimodal data and supervised fine-tuned using labeled multimodal data. It has semantic understanding and logical reasoning capabilities, and its generation strategy is guided by strict safety requirements based on reinforcement learning with human feedback. Downstream task interfaces are set according to the functions of MDAI-MS. The MDAI-MV corresponds one-to-one with the maglev vehicles included in the maglev transportation system, and is deployed in the onboard computer of the corresponding maglev vehicle. i Serving the i-th maglev vehicle, it is used to complete vehicle-level safety requirements related to transportation needs; The MDAI-MV i Based on neural network construction, with decoder-encoder as the basic architecture, it is self-supervised pre-trained with a large amount of unlabeled multimodal data, and has semantic understanding and logical reasoning capabilities. Based on reinforcement learning with human feedback, its generation strategy is guided by strict safety requirements, and the downstream task interface is set according to the MDAI-MS function. The MAGI-MV corresponds one-to-one with the maglev vehicles included in the maglev transportation system, and is deployed in the onboard computer of the corresponding maglev vehicle. i Serving the i-th maglev car, and MAGI-MV i With MDAI-MV i Together they serve the i-th maglev vehicle to complete vehicle-level non-safety-critical tasks involved in the transportation process; The MAGI-MV i Based on neural network construction and with decoder-encoder as the basic architecture, it is self-supervised pre-trained using a large amount of unlabeled multimodal data, and has semantic understanding and logical reasoning capabilities. Furthermore, it uses reinforcement learning based on human feedback to ensure that its generation strategy does not violate ethical standards, and sets up downstream task interfaces according to the functions of MDAI-MS.

3. The multi-modal magnetic levitation intelligent control system as described in claim 1, characterized in that, The ubiquitous sensing subsystem (4): It includes a first sensing module, a second sensing module, a third sensing module, and a fourth sensing module; The first sensing module is used to estimate the operating speed and spatiotemporal position of the maglev vehicle, and includes a speed sensor, a synchronization clock, a positioning sensor, and satellite navigation. The second sensing module is used to estimate the guide gap and suspension gap of the maglev vehicle, including a guide gap sensor and a suspension gap sensor; The third sensing module is used to collect dynamic shape information of the transported objects inside the maglev vehicle, the track line, and the external environment, including an in-vehicle 3D depth camera, a drone vision camera, an infrared vision camera, and a lidar. The fourth sensing module is used to monitor the operation and fault status of the maglev car, track line, turnout mechanism, motor equipment and auxiliary equipment, including vibration sensors, acoustic emission sensors, voltage sensors and current sensors.

4. The multi-modal magnetic levitation intelligent control system as described in claim 1, characterized in that, The training steps for a multimodal artificial intelligence model include: Step S101: Pre-train each modal encoder, each modal feature layer, multimodal information interaction layer, and each modal output layer; Step S102: Supervised training of the weight parameters of the downstream task interface, and fine-tuning with instructions to fine-tune the multimodal artificial intelligence model; Step S103: Constrain the multimodal artificial intelligence model through reinforcement learning with human feedback.

5. The multimodal magnetic levitation intelligent control system as described in claim 4, characterized in that, Step S101: Each modal encoder, each modal feature layer, and each modal output is trained using data consisting only of the corresponding modal signals; The self-supervised pre-training process for each modal encoder, each modal feature layer, multimodal information interaction layer, and each modal output layer uses predictive loss functions and contrastive loss functions to enable the multimodal AI model to have cross-modal understanding capabilities.

6. The multimodal magnetic levitation intelligent control system as described in claim 1, characterized in that, The system operation process is as follows: Step S201: System users submit transportation requests to MDAI-MS through the user interface; Step S202: MDAI-MS analyzes transportation demand, allocates maglev vehicles, generates the necessary driving information based on physical static parameters, and transmits it to the corresponding MDAI-MV; Step S203: MDAI-MV uses in-vehicle visual information to determine whether the identity, behavior, and location information of the transported object meet the requirements for safe departure of the maglev train; Step S204: Close the doors and MDAI-MV generates control commands required for the maglev car to operate based on real-time ubiquitous sensing information and sends them to relevant safety-critical equipment to complete the transportation requirements; Step S205: During the operation of the maglev car, MAGI-MV completes the non-safety-critical requirements submitted by the system users; Step S206: The maglev vehicle travels to the transportation destination, completes the transportation request submitted by the system user, and closes the system user's access rights to MAGI-MV.