A vehicle load identification method based on a multi-modal large model

By constructing a benchmark bridge case library and transferring cross-bridge knowledge through a multimodal large model, rapid and accurate identification of bridge vehicle loads was achieved, solving the problems of equipment fragility, high cost, and low efficiency in existing technologies, and meeting the high-precision requirements of bridge safety monitoring.

CN121388633BActive Publication Date: 2026-03-20NINGBO LANGDA ENG TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511924793.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-03-20
Estimated Expiration
2045-12-19

AI Technical Summary

Technical Problem

Existing technologies for bridge vehicle load identification suffer from problems such as equipment fragility, poor environmental adaptability, complex calibration, high cost, low efficiency, and poor model transferability, making it difficult to achieve continuous and dynamic monitoring of bridges and large-scale promotion.

Method used

We construct a benchmark bridge case library covering multiple types of bridges and various working conditions. We use a multimodal large model for retrieval enhancement generation and cross-bridge knowledge transfer. We analyze the core features of new bridges through the multimodal large model, match similar benchmark bridge cases, and combine real-time traffic flow conditions to reconstruct and identify the architecture of the load response model.

Benefits of technology

It has enabled the rapid deployment and application of the system on dozens to hundreds of bridges in the region, reducing deployment and time costs, solving the identification confusion problem in multi-vehicle coupled scenarios, meeting the high-precision requirements of bridge safety monitoring, and improving the system's sustainability and economy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121388633B_ABST
    Figure CN121388633B_ABST
Patent Text Reader

Abstract

The application discloses a vehicle load identification method based on a multi-modal large model, which comprises the following steps: selecting a reference bridge in a region for testing to obtain a reference bridge case library; analyzing a new bridge through a multi-modal large model to match a plurality of reference bridges with the closest core characteristics of the new bridge from the reference bridge case library; the multi-modal large model migrates the load response model and model parameters of the reference bridge case to the new bridge through a case reasoning mechanism; and the multi-modal large model reconstructs the migrated load response model according to the structural differences between the new bridge and the matched reference bridge and the actual working conditions of the new bridge. The application has the beneficial effect that, through the constructed reference bridge case library, the generation of the multi-modal large model retrieval enhancement and the cross-bridge knowledge migration mechanism, the new bridge can directly reuse the model parameters and load response law of the reference bridge, thereby greatly reducing the time cost and economic cost of cross-bridge deployment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent traffic monitoring, in particular to a vehicle load identification method based on a multi-modal large model. BACKGROUND

[0002] Currently, the identification of vehicle loads for regional bridge groups mainly relies on two types of means: one is a traditional physical sensor system, and the other is an intelligent identification model based on deep learning. The methods adopted by the existing technology for the above two types of means include:

[0003] Static weighing method: by setting a weight indicator or strain gauge, piezoelectric sensor, etc. weighing device on the bridge deck, the vehicle is weighed at low speed or static state. This method has high precision, but it must interrupt traffic, has high construction and maintenance cost, and is difficult to realize continuous and dynamic monitoring of the bridge.

[0004] Dynamic weighing system: install sensor arrays under the bridge deck or road surface, and collect axle load and total weight data in real time during vehicle travel. Such systems have been widely used in highways, but they rely on fixedly deployed sensor networks, have problems such as easy damage of equipment, poor environmental adaptability, and complex calibration; and are difficult to implement in old bridges or scenes lacking pre-buried sensors.

[0005] Vehicle load identification method based on deep learning: identify the vehicle type on the bridge deck through a high-definition camera combined with a target detection algorithm such as YOLO, and inversely deduce the vehicle mass combined with the deflection response data of the bridge. Although it has a certain level of intelligence, it requires high data quality and model generalization ability. The current method needs to collect data and train models for each bridge, which is difficult to quickly promote in large-scale bridge groups, has high training cost and low efficiency, and poor model migration. SUMMARY

[0006] One of the purposes of the present application is to provide a vehicle load identification method based on a multi-modal large model, which can solve at least one of the defects in the background art.

[0007] To achieve the above at least one purpose, the technical solution adopted by the present application is: a vehicle load identification method based on a multi-modal large model, comprising the following steps:

[0008] S100: Select reference bridges of different construction ages, structure types, and stiffness levels in the region, and perform tests under multiple load conditions to obtain a reference bridge case library;

[0009] S200: Analyze the core characteristic parameters of newly deployed bridges in the region through a multi-modal large model, and according to the analysis results, use an enhanced generation mechanism to match multiple reference bridges with the closest core characteristics to the new bridge from the reference bridge case library;

[0010] S300: Based on the matched benchmark bridge case, the multi-modal large model migrates the load response model and model parameters of the benchmark bridge case to the new bridge through the case reasoning mechanism;

[0011] S400: According to the structural differences between the new bridge and the matched benchmark bridge and the actual working conditions of the new bridge, the multi-modal large model restructures the migrated load response model, obtains the load response model architecture with the highest matching degree for the current working condition, and performs load identification.

[0012] Preferably, step S200 includes the following process: based on the multi-modal large model for analyzing the core characteristic parameters of the new bridge, the obtained structured data and the unstructured data based on real-time traffic are vectorized to obtain a new bridge feature vector; the multi-modal large model is used to extract multi-modal features from the benchmark bridge case library, and the obtained structured data and unstructured data are vectorized to obtain benchmark bridge feature vectors corresponding to different benchmark bridges; the cosine similarity of the new bridge feature vector and each benchmark bridge feature vector is calculated, and the multiple benchmark bridges that meet the cosine similarity requirement are matched with the new bridge.

[0013] Preferably, for the structured data of the new bridge and the benchmark bridge, a difference enhancement strategy is adopted when constructing the feature vector, which includes the following process: based on the structural feature types of the bridge, the structural features of the new bridge and the benchmark bridge are clustered to obtain structural clusters corresponding to different structural feature types; principal component analysis is performed on each structural cluster, and the first two principal components are taken as structural common vectors to obtain structural deviation vectors; the obtained structural deviation vectors are amplified based on the channel attention mechanism to obtain structural deviation enhancement vectors; the structural common vectors and the structural deviation enhancement vectors are re-spliced to obtain the structural feature vectors of the bridge after difference enhancement.

[0014] Preferably, the expression of the structural feature vector of the bridge after difference enhancement is as follows:

[0015] ;

[0016] ;

[0017] ;

[0018] Among them, denotes the structural common vector, denotes the structural difference enhancement vector, and μ denotes the difference amplification multiple, denotes the gain vector, denotes the structural deviation vector, ​denotes an element-wise multiplication, σ(·) denotes a sigmoid activation function, W1 denotes a downweighting matrix, W2 denotes an upweighting matrix, and ReLU(·) denotes a rectified linear unit (ReLU) activation function.

[0019] Preferably, the step S300 comprises the following process: taking the load response model with the highest occurrence frequency in the matched multiple reference bridges as the initial template for new bridge load identification; the multi-modal large model takes the corresponding case data of the matched multiple reference bridges as the anchor point, and generates structured reasoning basis combining the real-time working condition characteristics of the new bridge; the initial template is modified according to the obtained structured reasoning basis, and the modified new bridge data is marked and supplemented to the reference bridge case library.

[0020] Preferably, when performing the step S400, the multi-modal large model reconstructs the neural network architecture of the structural response model of the new bridge based on the real-time traffic condition of the new bridge and the super parameters of the matched reference bridges, and combines the performance feedback data of different architecture models in the same scene in the reference bridge case library, including the following process: dividing an adaptive neural network architecture candidate pool based on the data of the reference bridge case library; designing a grid structure based on grid depth, neuron number and feature fusion mode according to the structural characteristics and traffic condition characteristics of the new bridge; and selecting the optimal neural network architecture with the highest matching to the current scene through simulation training and verification.

[0021] Preferably, the division of the neural network architecture candidate pool comprises the following process: according to the obtained structured reasoning basis, the core constraints of the neural network architecture design are determined; the neural network architecture candidate pool is divided directionally combining the traffic condition characteristics of the new bridge and the model experience data of the matched reference bridges, and the neural network architectures with poor performance feedback in the same scene in the reference bridge case library are excluded.

[0022] Preferably, when designing the neural network architecture, the feature dimension of the new bridge needs to be adaptively mapped to the input layer, which specifically comprises the following process: preprocessing the input features of the new bridge, if the stiffness of the new bridge and the reference bridge is less than a set value, linearly correcting and adjusting the deflection response data of the new bridge; in view of the interference of environmental factors, the "environment-deflection correction curve" of the matched reference bridge is called from the reference bridge case library to correct the response data of the new bridge; combining the corrected data features and the input layer design experience of the same data of the reference bridge, the optimal sequence length of the time series data of the reference bridge is taken as the input layer parameter.

[0023] Preferably, in the design of the neural network architecture, the generation of the grid structure includes the following process: according to the complexity of the mechanical properties of the new bridge, the complexity of the working condition and the structural difference between the new bridge and the reference bridge, a plurality of groups of grid structure candidates are dynamically generated; if the inference indicates that the complexity of the new bridge is higher than that of the reference bridge, the optimal depth of the reference bridge in the same scene is taken as the reference, and the grid depth is increased by 1-2 layers; for scenes with a grid depth greater than 5 layers, residual connections are automatically added; the number of input layer neurons is strongly bound with the modified feature dimension, and the number of hidden layer neurons is adaptively attenuated layer by layer.

[0024] Preferably, the process of fine-tuning the load response model of the new bridge is as follows: the space-time correlation of vehicles driving on the reference bridge and the new bridge is established through license plate recognition technology; if there are vehicles continuously driving through the reference bridge and the new bridge, the load response data of the vehicles on the reference bridge are automatically associated with the new bridge data set for training of the load response model.

[0025] Compared with the prior art, the beneficial effects of the present application are as follows:

[0026] (1) The technical scheme of the present application builds a reference bridge case library covering multiple types of bridges and multiple working conditions, combines the retrieval enhancement generation and cross-bridge knowledge transfer mechanism of a multi-modal large model, and directly reuses the model parameters and load response rules of the reference bridge for the new bridge, without the need to repeat experiments, thereby greatly reducing the time cost and economic cost of cross-bridge deployment, and realizing the rapid popularization and application of dozens to hundreds of bridges in a region.

[0027] (2) The technical scheme of the present application relies on the accurate analysis capability of the multi-modal large model for single vehicle, multiple vehicles driving together, following vehicle and other working conditions, and combines the adaptation mechanism of the special model for different working conditions, to solve the problem of identification confusion in the multi-vehicle coupling scene of the traditional method, and meet the high-precision requirement of load data for bridge safety monitoring.

[0028] (3) The technical scheme of the present application utilizes the automatic structure search capability driven by the multi-modal large model, dynamically adjusts the model architecture according to the structural characteristics of different bridges and the characteristics of vehicle flow, and replaces the traditional manual parameter tuning mode which relies on expert experience.

[0029] (4) The technical scheme of the present application uses the cross-bridge data supplement mechanism and the case library dynamic updating capability through license plate recognition, so that the system can autonomously accumulate real operation data without the need for periodic production calibration or full-scale model retraining. In long-term operation and maintenance, the labor cost of data collection and model updating is significantly reduced, and the sustainability and economy of the system in the whole life cycle monitoring of the bridge are improved. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 It is a schematic diagram of the overall working steps of the present application. DETAILED DESCRIPTION

[0031] Hereinafter, the present application will be further described with reference to the specific embodiments, it should be noted that in the description of the present specification, the description referring to the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms should not be understood as necessarily referring to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in the present specification.

[0032] In the description of the present application, it should be noted that for orientation words such as the terms "center", "transverse", "longitudinal", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise" and the like indicate the orientation and positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the purpose of facilitating the description of the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and cannot be understood as limiting the specific protection scope of the present application.

[0033] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence.

[0034] In the present application, unless otherwise explicitly specified and limited, the terms "mounting", "connection", "connection", "fixing" and the like should be understood in a broad sense, for example, it can be connected, or detachable connection, or integrated; it can be mechanical connection, or electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, it can be the internal communication of two elements or the interaction relationship between two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0035] In the present application, unless specifically stated and limited otherwise, the first feature is "on" or "under" the second feature can include that the first and second features are in direct contact, or that the first and second features are not in direct contact but are in contact through another feature between them. Moreover, the first feature is "on", "above" and "over" the second feature includes that the first feature is directly above and obliquely above the second feature, or only indicates that the first feature is higher in horizontal height than the second feature. The first feature is "under", "below" and "underneath" the second feature includes that the first feature is directly below and obliquely below the second feature, or only indicates that the first feature is lower in horizontal height than the second feature.

[0036] The terms "comprising" and "having" and any variations thereof in the specification and claims of the present application are intended to cover the non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units includes not necessarily limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0037] One of the preferred embodiments of the present application is shown as Figure 1 A vehicle load identification method based on a multi-modal large model, comprising the following steps:

[0038] S100: Select reference bridges of different construction ages, structure types and stiffness levels in the region to conduct tests under multiple load conditions to obtain a reference bridge case library.

[0039] It can be understood that before the deployment of the load response model for the new bridge in the region, there are multiple historical bridges in the region that have completed the deployment; for some historical bridges in the same region, there are similarities in structure design and working conditions with the new bridge, so the data obtained by testing the historical bridges in the region which have deployed the load response model can be used as the data basis for the new bridge.

[0040] S200: Analyze the core characteristic parameters of the new bridge deployed in the region by a multi-modal large model, and match multiple reference bridges with the closest core characteristics to the new bridge from the reference bridge case library according to the analysis results using an enhanced generation mechanism.

[0041] It can be understood that the selected reference bridges in the region cover all bridge structure types, that is, there are bridges in the reference bridges that are inconsistent with the structure type of the new bridge; due to the inconsistency of the bridge structure, the generality of the corresponding load response model is poor in theory, so by comparing the characteristics of the new bridge and the reference bridge, the reference bridge with similar characteristics to the new bridge is selected as the basis for subsequent knowledge transfer.

[0042] S300: Based on the matched benchmark bridge case, the multi-modal large model migrates the load response model and model parameters of the benchmark bridge case to the new bridge through the case reasoning mechanism.

[0043] S400: According to the structural differences between the new bridge and the matched benchmark bridge and the actual working conditions of the new bridge, the multi-modal large model reconstructs the migrated load response model, obtains the load response model architecture with the highest matching degree for the current working condition, and performs load identification.

[0044] It can be understood that since there are still some differences in structure between the new bridge and the benchmark bridge, the actual working conditions caused by the structural differences may also have some differences, so the load response model of the migrated benchmark bridge needs to be adaptively adjusted and reconstructed based on the actual situation of the new bridge, to ensure that the obtained load response model can well match the actual working condition of the new bridge, to ensure the accuracy of vehicle load identification.

[0045] Compared with the traditional way of collecting data and training models for each new bridge in the regional bridge group. The technical solution of the present application constructs a benchmark bridge case library covering multiple types of bridges and multiple working conditions, combines the retrieval enhancement generation and cross-bridge knowledge migration mechanism of the multi-modal large model, and the new bridge can directly reuse the model parameters and load response law of the benchmark bridge. There is no need to repeat the experiment, so as to greatly reduce the time cost and economic cost of cross-bridge deployment, and realize the rapid popularization and application of dozens to hundreds of bridges in the region.

[0046] In this embodiment, for the establishment of the benchmark bridge case library in step S100, representative bridges can be selected from the regional bridge group as benchmark bridges. First, representative bridges covering different construction years (new bridges, bridges in service for 10 years, bridges in service for more than 20 years), structural types (simple beam bridges, continuous beam bridges, arch bridges, etc.) and stiffness grades are selected as benchmark bridges. Then, a multi-working condition load test scheme is designed for each benchmark bridge: a systematic test is carried out through a calibrated weight car (equipped with a standard load module with known weight), covering different load levels (10t, 20t, 30t, etc.), driving speeds (20km / h, 40km / h, 60km / h, etc.) and traffic working conditions (single vehicle driving, double vehicle driving, multiple vehicle following (vehicle distance 5m / 10m)), to ensure coverage of typical scenes in actual traffic.

[0047] For each test process of the benchmark bridge, multi-dimensional data, vehicle load and working condition data, load response data, and visual auxiliary data are synchronously collected. The multi-dimensional data includes bridge name, geographic location, construction year, structure type, span combination, main beam material, design load level, measured stiffness parameters (such as mid-span cross-section bending stiffness), bridge deck pavement type, etc., which are obtained and quantified through drawing analysis and field detection. The vehicle load and working condition data include the actual total weight of the weight car, the driving lane position, and the number of vehicles under multi-vehicle working conditions, which are synchronously recorded by the vehicle positioning equipment and the bridge deck camera. The load response data are collected by the target arranged at the key section of the benchmark bridge under the action of the vehicle load, using digital image correlation (DIC) algorithm, optical flow method or millimeter wave radar, etc. The dynamic deflection time response data under the vehicle load are synchronously recorded, and the environmental interference parameters (such as temperature) are recorded for later data correction. The visual auxiliary data are collected by the bridge deck high-definition camera to obtain the vehicle driving images / videos, which are used to extract the vehicle type features, driving state and traffic working condition visual labels.

[0048] It can be understood that the core data of the benchmark bridge case library obtained through the above process includes three key elements: first, bridge meta information, covering the construction age, structure type, bridge deck common traffic working condition and other basic attributes; second, load-response correlation data, that is, the structured mapping relationship collected under different test working conditions (such as the correlation between the total weight of the weight car and the corresponding mid-span dynamic deflection peak value, the load action position and the deflection distribution characteristics); and third, model experience data, including the load response model structure (the most suitable network architecture such as BiLSTM) and model parameters (such as the optimal hyperparameter combination, the weight matrix of training convergence) of the benchmark bridge verified and adapted, and other historical optimization results. For the above multi-source heterogeneous data, the numerical structure parameters can be preprocessed, and the features (such as structure type and working condition type) can be embedded and vectorized. The cases are vectorized and coded and stored in a vector database supporting semantic retrieval, to build a benchmark bridge case library covering multiple bridge types and multiple load working conditions, and to provide core data support for the retrieval, enhanced reasoning and cross-bridge knowledge transfer of subsequent multi-modal large models.

[0049] In this embodiment, step S200 includes the following process:

[0050] S210: Based on the multi-modal large model, the core characteristic parameters of the new bridge are analyzed, the obtained structured data and unstructured data based on real-time traffic are vectorized, and the new bridge feature vector is obtained.

[0051] It can be understood that the structured data of the new bridge includes the construction age, the structure type, the span component, the design load level and the measured stiffness, and the unstructured data includes real-time traffic visual data, etc.

[0052] S220: Multi-modal feature extraction is performed on the benchmark bridge case library by a multi-modal large model, the obtained structured data and unstructured data are vectorized, and benchmark bridge feature vectors corresponding to different benchmark bridges are obtained.

[0053] It can be understood that the structured data of the benchmark bridge case library includes construction age, structure type, span component, design load level, and measured stiffness, etc., which are converted into numerical feature vectors through standardized coding; the unstructured data includes dynamic deflection time series response curve and vehicle flow video frame features, etc., which can be converted into dimension semantic vectors by a multi-modal large model; after obtaining the feature vectors of the structural data and the feature vectors of the unstructured data, the multi-modal large model can uniformly map all the feature vectors to the same semantic space to obtain the benchmark bridge feature vectors corresponding to different benchmark bridges, and all the benchmark bridge feature vectors can form a benchmark bridge case vector library and be stored in a vector database.

[0054] S230: Calculate the cosine similarity of the new bridge feature vector and each benchmark bridge feature vector, and match the new bridge with multiple benchmark bridges that meet the cosine similarity requirement.

[0055] It can be understood that a similarity threshold can be set to determine the similarity between the benchmark bridge and the new bridge, i.e., if the calculated cosine similarity of the new bridge feature vector and the benchmark bridge feature vector is greater than or equal to the set similarity threshold, the cosine similarity requirement is met, and the benchmark bridge and the new bridge can be matched. The specific value of the similarity threshold can be set by the person skilled in the art according to actual needs, for example, the value of the similarity threshold can be 0.8-0.9. Generally speaking, there can be multiple benchmark bridges that can meet the similarity threshold, the benchmark bridges that meet the similarity threshold can be arranged in descending order of the cosine similarity, and then the first few benchmark bridges can be selected to match the new bridge, for example, the first 3-5 benchmark bridges can be selected to match the new bridge. The specific calculation process of the cosine similarity is known to those skilled in the art, and therefore will not be described in detail here.

[0056] It needs to be known that based on the above content, it can be known that the reference bridge feature vector and the new bridge feature vector both include structural feature items and non-structural feature items; for the non-structural feature items, they are mainly affected by structural features and external environment features, and the external environment features can be regarded as consistent in the region, so the similarity of the reference bridge and the new bridge mainly reflects on the structural features. However, the differences of the structural parameters such as the stiffness and span component of the bridge are often 5% to 15%, but in the case that the high-dimensional vector is submerged by a large amount of redundant information, the cosine similarity may relatively reduce the structural difference, so that the matching result of the reference bridge and the new bridge has an error, and the accuracy of the subsequent migrated load response model is low. Therefore, before the calculation of the cosine similarity, the structural data of the new bridge and the reference bridge need to be subjected to difference enhancement, so that the mechanical deviation determining the load response relationship of the new bridge and the reference bridge is amplified.

[0057] Specifically, for the structured data of the new bridge and the reference bridge, a difference enhancement strategy is adopted when constructing the feature vector, and the specific process includes the following steps:

[0058] S201: Based on the structural feature types of the bridge, the structural features of the new bridge and the reference bridge are clustered to obtain structural clusters corresponding to different structural feature types.

[0059] It can be understood that, as known from the foregoing, the structural feature types of the bridge include the construction age, the structural type, the span component, the design load level, and the measured stiffness. There are significant differences between these structural feature data, so the same type of data can be classified by clustering to obtain corresponding structural clusters; the structural clusters are maintained in the form of vectors.

[0060] S202: Principal component analysis is performed on each structural cluster, and the first two principal components are taken as the structural common vectors to obtain the structural deviation vectors.

[0061] It can be understood that the specific process of principal component analysis is known to those skilled in the art, and therefore will not be described in detail here; based on principal component analysis, the commonness of the features of the earlier dimensions is stronger, so in this embodiment, the first two principal components are preferred as the structural common vectors, and the remaining components are taken as the structural deviation vectors. It needs to be noted that the dimension of the structural deviation vector is consistent with the dimension of the structural cluster, and the dimension corresponding to the common vector can be replaced by a unit vector.

[0062] S203: The obtained structural deviation vector is amplified based on the channel attention mechanism to obtain a structural deviation enhancement vector; after the structural common vector and the structural deviation enhancement vector are re-spliced, a bridge structural feature vector based on difference enhancement is obtained. The expression of the bridge structural feature vector after difference enhancement is as follows:

[0063] .

[0064] .

[0065] .

[0066] wherein, denotes a structure common vector, denotes a structure difference enhancement vector; μ denotes a difference amplification multiple, and a specific value can be selected by a person skilled in the art according to actual needs, for example, can be taken as 2-4; denotes a gain vector, denotes a structure deviation vector, denotes an element-by-element multiplication, σ(·) denotes a Sigmoid activation function, W1 denotes a weight reduction matrix, W2 denotes a weight increase matrix, and ReLU(·) denotes a linear rectification activation function.

[0067] In the embodiment, when step S300 is performed, the multi-modal large model realizes cross-bridge knowledge transfer through a case reasoning mechanism to obtain model experience data of the reference bridge, including a load response model structure and model parameters of the reference bridge which have been verified to be adapted; wherein the case reasoning mechanism is based on four parts including “retrieval-reuse-correction-preservation”, which will be described in detail below.

[0068] Specifically, for case retrieval, the case data corresponding to the multiple reference bridges matched by the retrieval enhancement generation mechanism of step S200 is taken as a source case. For case reuse, the load response model with the highest frequency of occurrence in the source case is extracted as an initial template for new bridge load identification. For case correction, the multi-modal large model takes the retrieved source case as an anchor point, combines the real-time working condition characteristics of the new bridge, generates a structured reasoning basis, and modifies the initial template according to the obtained structured reasoning basis; for example, if the stiffness difference ΔK≤10%, the multi-modal large model will infer that “linear correction (correction coefficient = new bridge stiffness / source case bridge stiffness) is recommended”. For case preservation, the modified new bridge data is also supplemented to the reference bridge case library, but the bridge meta-information label is a newly deployed bridge whose model parameters are intelligently generated by the multi-modal large model, rather than a bridge obtained by actually running a weight car. Although only reference bridges are retrieved at present, they will also be considered in the future.

[0069] In this embodiment, after the new bridge completes the knowledge transfer of the load response model, to meet the demand for accurate load identification in the mixed traffic scenario in actual traffic, through the high-precision visual understanding and scene reasoning ability of the multi-modal large model, relying on the real-time visual data collected by the bridge deck traffic monitoring camera, an intelligent response mechanism of "working condition identification-model adaptation" is constructed: first, the multi-modal large model accurately analyzes the real-time visual data of the traffic scene, automatically identifies different traffic working condition types such as single vehicle driving, multi-vehicle driving, multi-vehicle following, and extracts key features such as vehicle number, relative position, and driving distance; Then, based on the identification result, the corresponding adaptive deflection-traffic load deep learning model is intelligently matched. If it is a single vehicle working condition, the load response model trained by single vehicle data is directly called to accurately output the load data of a single vehicle; if it is a complex working condition such as multi-vehicle driving and following, the load response model trained by multi-vehicle coupled data is automatically called, and through the accurate modeling of the model on the coupling relationship between multi-vehicle load and deflection response, the synchronous and accurate output of the load data of each vehicle is realized. Compared with the traditional method, the technical scheme of the present application relies on the accurate analysis ability of the multi-modal large model for single vehicle, multi-vehicle driving, and following working conditions, and combines the adaptation mechanism of the special model for different working conditions, solves the problem of identification confusion in the multi-vehicle coupling scene of the traditional method, and meets the high-precision requirements of bridge safety monitoring for load data.

[0070] It can be understood that the multi-modal large model deeply analyzes the real-time visual data, identifies key information such as vehicle type, license plate, and lane, and judges the traffic working condition type combined with time sequence analysis technology. If there is only one vehicle in the image and no other vehicles interfere, it is determined as "single vehicle driving working condition"; if there are two or more vehicles in adjacent lanes and the lateral distance is ≤2m, it is determined as "multi-vehicle driving working condition"; if multiple vehicles are driving in the same lane and the longitudinal distance is ≤5m, it is determined as "multi-vehicle following working condition". At the same time, the multi-modal large model extracts key features of the working condition: records the vehicle speed and lane position in single vehicle working condition; records the number of vehicles, relative distance (lateral / longitudinal), vehicle type combination (such as mixed driving of trucks and passenger cars) and other parameters in multi-vehicle working condition, forming a structured working condition feature vector, which provides accurate scene labels for subsequent model adaptation.

[0071] It should be noted that based on the structural characteristics between the reference bridges in the region, multiple neural network architectures of load response models need to be designed for different working conditions. Commonly used neural network architectures include CNN, BiLSTM, Transformer, and CNN-BiLSTM hybrid architecture. At the same time, the performance optimization of the model needs to rely on manual adjustment of structural parameters such as network layer number and neuron number, and hyperparameters such as learning rate and batch size. Since the initial template of the load response model of the new bridge is migrated from the reference bridge, although the reference bridge and the new bridge are similar in structural characteristics after matching by cosine similarity, there are still some differences. Therefore, based on the structural difference scenario, the neural network architecture of the load response model of the new bridge with the reference bridge as the initial template is not necessarily optimal. Therefore, in this embodiment, after the new bridge obtains the initial template through the knowledge transfer of the multi-modal large model, intelligent search and optimization design of the neural network architecture are automatically carried out. At the same time, the multi-modal large model can automatically test various deep learning model architectures and hyperparameter combinations according to the measured data characteristics of different bridges. By comparing the performance indicators such as recognition accuracy and convergence speed of different schemes, the optimal model configuration is adaptively selected. Compared with the traditional way, the technical solution of the present application can dynamically adjust the model architecture according to the structural characteristics and traffic characteristics of different bridges through the automatic structure search and code generation capability of the multi-modal large model, replace the traditional manual parameter adjustment mode which relies on expert experience, realize the rapid iteration and dynamic adaptation of the model structure, significantly improve the development efficiency, and reduce the maintenance cost.

[0072] Specifically, when step S400 is executed, the multi-modal large model reconstructs the neural network architecture of the structural response model of the new bridge based on the real-time traffic conditions of the new bridge and the matched reference bridge hyperparameters, and the performance feedback data of different architecture models in the reference bridge case library in the same scene, including the following processes: dividing the adaptive neural network architecture candidate pool based on the data of the reference bridge case library; designing a grid structure based on grid depth, neuron number, and feature fusion method according to the structural characteristics and traffic condition characteristics of the new bridge; and selecting the optimal neural network architecture that matches the current scene the highest through simulation training and verification.

[0073] The process of defining the candidate pool for neural network architectures includes the following steps: First, based on the structured reasoning evidence obtained in the preceding steps (such as "structural difference analysis between the new bridge and the matching benchmark bridge"), intelligent modeling prompts are used to clarify the core constraints of the neural network architecture design (for example, if the reasoning indicates that "the stiffness of the new bridge is slightly lower than that of the benchmark bridge but the consistency of operating conditions is high", then the architecture type that has been verified as valid by the benchmark bridge is retained first). Combining the traffic flow operating conditions characteristics of the new bridge (such as the spatial locality of single-vehicle operating conditions and the temporal coupling of multi-vehicle following) and the model experience data of the matching benchmark bridge, the candidate architecture pool is defined in a targeted manner. Among them, the single-vehicle operating conditions focus on BiLSTM and CNN-BiLSTM architectures, while the multi-vehicle operating conditions focus on Transformer and hybrid architectures. Architectures with poor performance feedback in similar scenarios in the benchmark bridge case library are excluded (such as pure CNN architectures with accuracy below 85% in multi-vehicle operating conditions), thus reducing the invalid search space.

[0074] In designing the neural network architecture, an adaptive mapping between the feature dimensions of the new bridge and the input layer is required. This process includes the following steps: Preprocessing the input features of the new bridge; if the stiffness difference ΔK between the new bridge and the benchmark bridge is ≤10%, linear correction is used to adjust the deflection response data. To address environmental interference, the "environment-deflection correction curve" of the matching benchmark bridge is retrieved from the benchmark bridge case library to correct the new bridge response data. For example, using temperature interference, the "temperature-deflection correction curve" of the benchmark bridge is retrieved to correct the new bridge response data, eliminating environmental bias. Combining the corrected data features (such as the dynamic deflection time series length) and the input layer design experience of similar data from the benchmark bridge, the optimal sequence length of the benchmark bridge's time series data is used as the input layer parameter. In other words, the multimodal large model automatically uses the optimal time step obtained by the benchmark bridge during the experiment as the input layer parameter of the reconstructed neural network architecture.

[0075] In the design of the neural network architecture, the generation of the grid structure includes the following process: according to the complexity of the mechanical properties of the new bridge (such as the nonlinear response of the long-span arch bridge > the linear response of the simply supported beam bridge), the complexity of the working condition (multi-vehicle following > double-vehicle driving > single-vehicle), and the structural difference between the new bridge and the reference bridge, a plurality of candidate grid structures are dynamically generated. In the design of the grid depth, i.e., the number of grid layers, if the inference indicates that the complexity of the new bridge is higher than that of the reference bridge, the optimal depth of the reference bridge in the same scenario is taken as the reference, and the grid depth is increased by 1 to 2 layers; for example, the grid depth of the simply supported beam bridge under single-vehicle working condition is 3-5 layers, and the grid depth of the complex arch bridge under multi-vehicle working condition is 6-8 layers. For scenarios with a grid depth greater than 5 layers, residual connections are automatically added to ensure training stability; for the automatic addition of residual connections, the deep network anti-gradient disappearance scheme of the reference bridge can be referred to, and therefore it will not be described here. The number of input layer neurons is strongly bound to the modified feature dimension, and the number of hidden layer neurons is adaptively attenuated layer by layer; that is, the number of input layer neurons is equal to the modified feature dimension, and the attenuation ratio of the number of neurons in the hidden layer layer by layer can be set by the actual needs of those skilled in the art, for example, attenuated by half layer by layer.

[0076] It should be noted that the network structure of the load response model used by the new bridge is obtained based on the retrieval and inference of the multi-modal large model, and the model weight of the reference bridge is currently used as the initialization weight of the model. However, if a better model result is desired, fine-tuning with real data on the new bridge is required. Since the new bridge does not perform a running test with a weight car, in this embodiment, the dynamic supplement of real running data of the new bridge is realized by spatiotemporal correlation between the reference bridge and the new bridge, without the need for separate weight car tests for the new bridge. For the convenience of understanding, the following will be described in detail.

[0077] In this embodiment, the process of fine-tuning the load response model of the new bridge is as follows: the spatiotemporal correlation between the reference bridge and the new bridge is established through license plate recognition technology; if there are vehicles continuously driving through the reference bridge and the new bridge, the load response data of the vehicles on the reference bridge are automatically associated with the new bridge data set for training of the load response model.

[0078] It can be understood that the new bridge and the reference bridge are located in the same regional bridge group, that is, they belong to the same region in geographical position. Then the vehicle driving space-time correlation of the new bridge and the reference bridge can be constructed through the recognition of license plates by the cameras on the new bridge and the reference bridge. For example, the reference bridge passes a large truck at time t1, and the multi-modal large model identifies the license plate of the large truck through the camera on the reference bridge. Since the reference bridge is subjected to the weight vehicle test process, the data of the load response model of the reference bridge is relatively accurate. Therefore, the multi-modal large model can record the load response data of the large truck passing through the reference bridge and identify it through the license plate. When the large truck drives to the new bridge at time t2, the multi-modal large model can identify the license plate of the large truck through the camera on the new bridge, and call the load response data of the large truck on the reference bridge to the data set of the new bridge for training the load response model of the new bridge according to the identified license plate. It should be noted that in order to ensure the accuracy of the load response model under different working conditions, the data set supplemented by the space-time correlation from the reference bridge to the new bridge passing through needs to include different working condition data, and the load response data under different working conditions are used for training the load response model of the working condition respectively. Compared with the traditional way, the technical scheme of the present application can autonomously accumulate real running data through the cross-bridge data supplement mechanism and the case library dynamic updating capability of license plate recognition correlation, without the need for regular production calibration or full-quantity retraining of the model. In long-term operation and maintenance, the labor cost of data acquisition and model updating is significantly reduced, and the sustainability and economy of the whole life cycle monitoring of the bridge are improved.

[0079] Specifically, in the training of the load response model of the new bridge based on different working conditions, for single vehicle working condition, the multi-modal large model automatically loads the single vehicle load-response data set sample library, and uses the MSE loss function to optimize the analytical ability of the model to single load; for multi-vehicle working condition, the multi-vehicle coupling data set is called, and the cross-entropy loss function is introduced to strengthen the learning of the load distribution ratio between vehicles, and the model is quickly converged and the precision is optimized through dynamic adjustment of the learning rate. In order to facilitate understanding, the training of the load response model under different working conditions will be described in detail below.

[0080] For the training of the load response model of single vehicle working condition: based on the code generation capability of the multi-modal large model, the supplemented single vehicle load-response dataset is automatically called, and the training set and the validation set are divided in the ratio of 8:2. In data preprocessing, based on the inference results of the multi-modal large model, such as the need for stiffness correction, the code is automatically generated to linearly correct the deflection data, and the weight label of the vehicle is normalized (mapped to the [0, 1] interval) and the like, to improve the model convergence efficiency. The loss function uses the MSE (mean square error) loss function, the optimizer uses Adam, and the initial learning rate is taken from the optimal hyperparameters of the benchmark bridge single vehicle model, which is set to 0.001 by default; the weight decay coefficient is set to 1e-5 to suppress overfitting. Stop when the total weight error of the validation set is ≤1% or the training is full 50 rounds, save the optimal model parameters.

[0081] For the training of the load response model of multiple vehicle working conditions: based on the code generation capability of the multi-modal large model, the supplemented multi-vehicle load-response dataset is automatically called, and the training set and the validation set are divided in the ratio of 8:2; at the same time, the data is grouped according to the subdivided working conditions (two vehicles in parallel, two vehicles following (vehicle distance 5m / 10m), three vehicles following, etc.), to ensure that the samples of each subdivided scene are balanced. In data preprocessing, based on the inference results of the multi-modal large model (such as stiffness difference, temperature interference, etc.), the code is automatically generated to linearly correct the deflection data (stiffness correction coefficient = new bridge stiffness / benchmark bridge stiffness) and temperature compensation (call the temperature-deflection correction curve of the benchmark bridge), etc.; for different subdivided working conditions, and the weight label of each vehicle is normalized (mapped to the [0, 1] interval), to strengthen the recognition ability of the load response model to the working condition difference. The loss function uses a hybrid loss of weighted cross-entropy (60%) + MSE (40%). Cross-entropy loss focuses on learning the load distribution ratio between vehicles, and MSE loss optimizes the absolute accuracy of single vehicle weight; the optimizer uses AdamW, the initial learning rate is taken from the optimal hyperparameters of the benchmark bridge corresponding to the subdivided working condition model, which is set to 0.002 by default, and the weight decay coefficient is set to 5e-5 to suppress overfitting. Stop when the total weight error of the validation set is ≤2% or the training is full 80 rounds, save the optimal model parameters.

[0082] The above describes the basic principles, main features and advantages of the present application. Those skilled in the art should understand that the present application is not limited to the above embodiments, and the above embodiments and descriptions in the specification are only the principles of the present application. Without departing from the spirit and scope of the present application, various changes and improvements can be made to the present application, and these changes and improvements all fall within the scope of the claimed present application. The scope of protection claimed by the present application is defined by the appended claims and their equivalents.

Claims

1. A vehicle load identification method based on a multimodal large model, characterized in that, Includes the following steps: S100: Select benchmark bridges with different years of construction, structural types and stiffness levels in the region, conduct tests under multiple working conditions, and obtain a benchmark bridge case library. S200: Analyze the core feature parameters of new bridges deployed in the region using a multimodal large model, and based on the analysis results, use an enhanced generation mechanism to match multiple benchmark bridges from the benchmark bridge case library that are most similar to the core features of the new bridge. S300: Based on the matching benchmark bridge case, the multimodal large model transfers the load response model and model parameters of the benchmark bridge case to the new bridge through the case reasoning mechanism; S400: Based on the structural differences between the new bridge and the matching benchmark bridge, as well as the actual working conditions of the new bridge, the multimodal large model reconstructs the architecture of the migrated load response model to obtain the load response model architecture with the highest matching degree to the current working conditions and performs load identification. Step S200 includes the following process: Based on the analysis of the core characteristic parameters of the new bridge using a multimodal large model, the obtained structured data and unstructured data based on real-time traffic flow are vectorized to obtain the feature vector of the new bridge. Multimodal feature extraction is performed on the benchmark bridge case library using a multimodal large model. The obtained structured and unstructured data are then vectorized to obtain benchmark bridge feature vectors corresponding to different benchmark bridges. Calculate the similarity between the feature vector of the new bridge and the feature vectors of each reference bridge, and match the new bridge with the reference bridges that meet the similarity requirements; For the structured data of the new bridge and the benchmark bridge, a difference enhancement strategy is adopted when constructing the feature vectors, which specifically includes the following process: Based on the structural feature types of bridges, the structural features of new bridges and benchmark bridges are clustered to obtain structural clusters corresponding to different structural feature types; Principal component analysis is performed on each structural cluster, and the structural deviation vector is obtained by taking the first two principal components as the structural commonality vector. The obtained structural deviation vector is amplified based on the channel attention mechanism to obtain the structural deviation enhancement vector. After re-concatenating the structural commonality vector and the structural deviation enhancement vector, the structural feature vector of the bridge based on the difference enhancement is obtained.

2. The vehicle load identification method based on a multimodal large model as described in claim 1, characterized in that, The structural feature vector of the bridge after differential enhancement The expression is as follows: ; ; ; in, Represents structural commonality vectors. This represents the structural difference enhancement vector, where μ represents the difference amplification factor. Represents the gain vector. Represents the structural deviation vector. σ(·) represents element-wise multiplication, W1 represents the weighting matrix, W2 represents the weighting matrix, and ReLU(·) represents the linear rectified activation function.

3. The vehicle load identification method based on a multimodal large model as described in claim 1 or 2, characterized in that, Step S300 includes the following process: The load response model that appears most frequently among the matched benchmark bridges is used as the initial template for load identification of the new bridge. The multimodal large model uses case data corresponding to multiple benchmark bridges as anchor points, and combines the real-time operating characteristics of the new bridge to generate structured reasoning basis. The initial template was modified based on the obtained structured reasoning, and the modified new bridge data was added to the benchmark bridge case library after being marked.

4. The vehicle load identification method based on a multimodal large model as described in claim 3, characterized in that, During step S400, the multimodal large model, based on the real-time traffic flow conditions of the new bridge and the matched hyperparameters of the benchmark bridge, and combined with the performance feedback data of different architecture models in the benchmark bridge case library under similar scenarios, restructures the neural network architecture of the new bridge's structural response model, including the following process: Based on data from the benchmark bridge case library, a candidate pool of suitable neural network architectures is defined; for the structural features and traffic flow characteristics of the new bridge, a grid structure based on grid depth, number of neurons, and feature fusion method is designed; and the optimal neural network architecture with the highest matching pair with the current scenario is selected through simulation training and verification.

5. The vehicle load identification method based on a multimodal large model as described in claim 4, characterized in that, The process of defining the candidate pool for neural network architectures includes the following steps: Based on the obtained structured reasoning, the core constraints of neural network architecture design are clarified; By combining the traffic flow characteristics of the new bridge with the model experience data of the matched benchmark bridge, a candidate pool of neural network architectures is defined in a targeted manner, and neural network architectures with poor performance feedback in similar scenarios in the benchmark bridge case library are excluded.

6. The vehicle load identification method based on a multimodal large model as described in claim 4, characterized in that, When designing a neural network architecture, it is necessary to adaptively map the feature dimensions of the new bridge to the input layer, which includes the following process: The input characteristics of the new bridge are preprocessed. If the stiffness of the new bridge and the reference bridge is less than the set value, the deflection response data of the new bridge is linearly corrected and adjusted. To address the interference from environmental factors, the "environment-deflection correction curve" of a matching benchmark bridge is retrieved from the benchmark bridge case library to correct the response data of the new bridge. Based on the corrected data characteristics and the input layer design experience of similar data from the benchmark bridge, the optimal sequence length of the benchmark bridge time series data is used as the input layer parameter.

7. The vehicle load identification method based on a multimodal large model as described in claim 4, characterized in that, When designing a neural network architecture, the generation of the mesh structure includes the following process: Based on the complexity of the mechanical properties of the new bridge, the complexity of the working conditions, and the structural differences between the new bridge and the reference bridge, multiple sets of candidate mesh structures are dynamically generated. If the inference indicates that the complexity of the new bridge is higher than that of the benchmark bridge, the mesh depth is increased by 1 to 2 layers based on the optimal depth of the benchmark bridge in similar scenarios; for scenarios with a mesh depth greater than 5 layers, residual connections are automatically added. The number of neurons in the input layer is strongly bound to the corrected feature dimension, and the number of neurons in the hidden layer is adaptively reduced layer by layer.

8. The vehicle load identification method based on a multimodal large model as described in claim 4, characterized in that, The process of fine-tuning the load response model of the new bridge is as follows: Establish a spatiotemporal correlation of vehicle travel between the benchmark bridge and the new bridge using license plate recognition technology; If there are vehicles that continuously cross the reference bridge and the new bridge, the load response data of the vehicles on the reference bridge will be automatically associated with the new bridge dataset for training the load response model.

Citation Information

Patent Citations

  • Data-driven modeling method for dynamic weighing of regional bridge

    CN120235056A