Road damage detection method and device based on multimodal large model

Through the multimodal large model fusion of images and text information, the problem of models relying on predefined categories in the prior art is solved, and high accuracy and adaptability of road disease detection is achieved, which is suitable for open vocabulary scenarios.

CN119495027BActive Publication Date: 2025-08-29STREAMAP TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510074499.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-08-29
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

The prior art relies on predefined categories in road disease detection, which makes it difficult to deploy models and not have high detection accuracy, making it difficult to adapt to open vocabulary scenarios.

Method used

The road disease detection method based on multimodal large model is adopted, and the semantic recognition model and multimodal large model are used to perform fusion detection of disease information, including a combination of backbone network, neck network and detection head, and feature extraction and fusion are used using ALCSR, TITFA and LITFF modules.

Benefits of technology

It improves the accuracy and adaptability of road disease detection, can accurately identify disease types and locations in open vocabulary scenarios, and reduces the complexity of model deployment and resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119495027B_ABST
    Figure CN119495027B_ABST
Patent Text Reader

Abstract

This application provides a road defect detection method and device based on a multimodal large model. This method obtains road image information and textual information about the defect type contained in the image information. The textual information is then input into a semantic recognition model to determine semantic information. The image information and semantic information are then input into the multimodal large model to determine the road defect information. This method enables road defect detection based on multimodal information, improving detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of road damage detection, and in particular relates to a road damage detection method and device based on a multimodal large model. Background Art

[0002] Roads are an essential component of urban planning and infrastructure development, and their condition has a crucial impact on urban operational efficiency and residents' quality of life. However, due to multiple factors, including the passage of time, environmental factors, traffic load, and material aging, road structures gradually develop defects such as cracks, potholes, and subsidence. These defects not only threaten driving safety and reduce driving comfort, but also shorten road lifespans and increase maintenance costs. Therefore, accurate detection of road defects is crucial. Relevant technologies have proposed methods based on object detection and image segmentation for road defect detection, but these methods often rely too heavily on predefined categories, significantly limiting their applicability in open-vocabulary scenarios. For example, if a model is trained specifically to identify road defects such as "potholes," "cracks," and "grid cracks," the introduction of new defect categories often requires retraining the entire model. This clearly imposes a heavy burden and significant challenges for large-scale deployment. Furthermore, these methods still suffer from low detection accuracy. Summary of the Invention

[0003] In response to the above-mentioned problems, the embodiments of the present application provide a road deterioration detection method and device based on a multimodal large model, which can improve the detection accuracy of road deterioration.

[0004] In a first aspect, an embodiment of the present application provides a road damage detection method based on a multimodal large model, the method comprising:

[0005] Image information of the road and text information about the type of damage in the image information are obtained.

[0006] Outputting the text information to a semantic recognition model to determine semantic information;

[0007] The image information and the semantic information are input into a multimodal large model to determine the road damage information.

[0008] In some embodiments, the multimodal large model includes: a backbone network, a neck network and a detection head. The input of the backbone network is the image information. The backbone network is used to extract multi-dimensional feature information based on the image information and output the multi-dimensional feature information to the neck network. The input of the neck network is the multi-dimensional feature information and the semantic information. The neck network is used to determine the fused feature information based on the multi-dimensional feature information and the semantic information, and output the fused feature information to the detection head. The detection head is used to detect and output the road disease information based on the fused feature information.

[0009] In some embodiments, the backbone network includes: multiple deformable adaptive lightweight channel segmentation and rearrangement ALCSR modules, each ALCSR module is used to perform adaptive channel segmentation on the input feature information, obtain the feature information of each channel, and merge the feature information of each channel for output, each channel exchanges information between channels through a channel rearrangement mechanism, each channel is provided with a deformable convolution module, the deformable convolution module is used to change the shape of the feature information in the corresponding channel and then output it, multiple ALCSR modules include: a first ALCSR module, a second ALCSR module, a third ALCSR module, a fourth ALCSR module and a fifth ALCSR module, the backbone network also includes: a first CBS module, a second CBS module Block, the first CBS module and the fourth CBS module, the first CBS module, the second CBS module, the first CBS module and the fourth CBS module are connected in sequence, the input of the first CBS module is the image information, the output of the fourth CBS module is the input of the first ALCSR module, the input of the second ALCSR module is the output of the first ALCSR module, the output of the second ALCSR module is the input of the third ALCSR module and the neck network, the output of the third ALCSR module is the input of the fourth ALCSR module, the output of the fourth ALCSR module is the input of the fifth ALCSR and the neck network, and the output of the fifth ALCSR is the input of the neck network.

[0010] In some embodiments, the neck network includes: multiple image-text feature alignment TITFA modules based on the attention mechanism, each TITFA module is used to fuse the input semantic information and feature information, and the multiple TITFA modules include: a first TITFA module, a second TITFA module, and a third TITFA module. The neck network also includes: a semantic information interaction module based on the attention mechanism, a first upsampling module, a first feature splicing module, a sixth ALCSR module, a second upsampling module, a second feature splicing module, a seventh ALCSR module, a fifth CBS module, a third feature splicing module, an eighth ALCSR module, a sixth CBS module, a fourth feature splicing module, and a ninth ALCSR module. The input of the semantic information interaction module is the output of the fifth ALCSR module, the output of the semantic information interaction module is the input of the first upsampling module, the input of the first feature splicing module is the output of the first upsampling module and the fourth ALCSR module, the output of the first feature splicing module is the input of the sixth ALCSR module, and the output of the sixth ALCSR module is the output of the second upsampling module and the first The input of the third feature splicing module is the output of the second ALCSR module and the second upsampling module, the input of the first TITFA module is the input of the second feature splicing module and the semantic information, the output of the first TITFA module is the input of the seventh ALCSR module, the output of the seventh ALCSR module is the input of the fifth CBS module and the detection head, the output of the fifth CBS module is the input of the third feature splicing module, the input of the second TITFA module is the semantic information and the output of the third feature splicing module, the output of the second TITFA module is the input of the eighth ALCSR module, the output of the eighth ALCSR module is the input of the sixth CBS module and the detection head, the input of the fourth feature splicing module is the output of the fifth ALCSR module and the sixth CBS module, the input of the third TITFA module is the semantic information and the output of the fourth feature splicing module, the output of the third TITFA module is the input of the ninth ALCSR module, and the output of the ninth ALCSR module is the input of the detection head.

[0011] In some embodiments, the semantic information interaction module includes: multiple shape change units and multiple residual connection units, the shape change unit is used to change the shape of the feature map, the residual connection unit is used to retain the original information, the multiple shape change units include: a first shape change unit and a second shape change unit, the multiple residual connection units include: a first residual connection unit and a second residual connection unit, the semantic information interaction module also includes: a first multi-head attention unit, a first multi-layer perceptron unit, the input of the first shape change unit is the output of the fifth ALCSR module, the output of the first shape change unit is the input of the first multi-head attention unit, the first multi-head attention unit is used to process the input information based on the multi-head attention mechanism, and output feature information to the first residual connection unit, the first residual connection unit is used to residually connect the information input into the semantic information interaction module with the features output by the first multi-head attention unit, the output of the first residual connection unit is the input of the first multi-layer perceptron unit, the first multi-layer perceptron unit performs nonlinear mapping and outputs feature information to the second residual connection unit, the output of the second residual connection unit is the input of the second shape change unit, and the output of the second shape change unit is the input of the first upsampling module.

[0012] In some embodiments, the TITFA module includes: a convolution unit, a third shape change unit, a second multi-head attention unit, a third residual connection unit, a second multi-layer perceptron unit, a fourth residual connection unit and a fourth shape change unit, wherein the convolution unit is used to convolve the semantic information and output the convolved semantic information to the second multi-head attention unit, the third shape change unit is used to change the shape of the feature information of the image and output the changed feature information to the second multi-head attention unit, the second multi-head attention unit performs feature fusion based on the multi-head attention mechanism, and outputs the fused feature information to the third residual connection unit, the output of the third residual connection unit is the second multi-layer perceptron unit, the output of the second multi-layer perceptron unit is the fourth residual connection unit, the output of the fourth residual connection unit is the input of the fourth shape change unit, and the fourth shape change unit is used to output the fused feature information.

[0013] In some embodiments, the detection head includes: a lightweight image-text feature fusion LITFF module, the LITFF module includes: multiple convolution units and a normalization unit, the convolution unit is used to perform dimensionality reduction and scale optimization on the image features of the fused features, the normalization unit is used to normalize the features in the numerical space, the multiple convolution units include: a first convolution unit and a second convolution unit, the multiple normalization units include: a first normalization unit and a second normalization unit, the LITFF module also includes: an Einstein summation convention unit, a parameter adjustment unit and a splicing unit, the input of the first convolution unit and the second convolution unit is the fused image information, the output of the first convolution unit is the input of the first normalization unit, the input of the second normalization unit is the semantic information, the input of the Einstein summation convention unit is the output of the first normalization unit and the second normalization unit, the output of the Einstein summation convention unit is the parameter adjustment unit, the input of the splicing unit is the output of the second convolution unit and the parameter adjustment unit, the splicing unit is used to output joint feature information and perform detection based on the joint feature information.

[0014] In a second aspect, an embodiment of the present application provides a road defect detection device based on a multimodal large model, comprising:

[0015] The acquisition module is used to acquire image information of the road and text information about the disease type in the image information.

[0016] A determination module, configured to output the text information to a semantic recognition model to determine semantic information;

[0017] A detection module is used to input the image information and the semantic information into a multimodal large model to determine the disease information of the road.

[0018] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method provided in the first aspect when executing the computer program.

[0019] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the method provided in the first aspect is implemented.

[0020] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it is used to implement at least one of the methods in the first aspect or the third aspect.

[0021] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0022] The multimodal large model-based road defect detection method provided in the embodiments of the present application obtains road image information and textual information about defect types contained in the image information. The textual information is output to a semantic recognition model to determine semantic information; the image information and semantic information are then input to the multimodal large model to determine the road defect information. This method can detect road defects based on multimodal information and improve detection accuracy.

[0023] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0025] Figure 1 A schematic diagram of a road damage detection method based on a multimodal large model provided by the present application;

[0026] Figure 2 A schematic diagram of the overall structure of a multimodal model provided in an embodiment of the present application;

[0027] Figure 3 A schematic diagram of the structure of an ALCSR module provided in an embodiment of the present application;

[0028] Figure 4 A schematic diagram of the structure of an ALCSR module provided in an embodiment of the present application;

[0029] Figure 5 A schematic diagram of the structure of a semantic information interaction module provided in an embodiment of the present application;

[0030] Figure 6 A schematic structural diagram of a TITFA module provided in an embodiment of the present application;

[0031] Figure 7 A schematic diagram of the structure of a LITFF module provided in an embodiment of the present application;

[0032] Figure 8 A schematic diagram of the structure of a road damage detection device based on a multimodal large model provided in an embodiment of the present application;

[0033] Figure 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0034] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0035] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0036] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0037] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if it is detected" can be interpreted as meaning "upon determining" or "in response to determining" or "upon detecting" or "in response to detecting," depending on the context.

[0038] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0039] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with the embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized.

[0040] Based on the technical issues of related technologies, the present invention provides a road defect detection method based on a multimodal large model that can be applied to electronic devices. These electronic devices may include: mobile phones, tablet computers, wearable devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). The present invention does not impose any restrictions on the specific type of electronic device.

[0041] The present invention provides a road damage detection method based on a multimodal large model. Figure 1 The present application provides a flow chart of a road damage detection method based on a multimodal large model, as shown in FIG. Figure 1 As shown, the method includes:

[0042] Step S101: Obtain image information of the road and text information about the damage type in the image information.

[0043] In the embodiments of this application, the road image information is photographs or video frames of the road surface captured by a camera, drone, satellite, or other image acquisition device. These images contain a visual representation of the road's physical characteristics and potential road damage. The text information describes the types of road damage that may be present in the image, such as cracks, potholes, ruts, and waterlogging.

[0044] In the embodiment of the present application, the text information may be manually input information.

[0045] Step S102: Output the text information to a semantic recognition model to determine semantic information.

[0046] In an embodiment of the present application, the semantic recognition model may be CLIP (Contrastive Language-Image Pre-training), which is a multimodal pre-training model proposed by OpenAI.

[0047] In this embodiment, the CLIP model includes a text encoder for processing textual information and extracting its semantic features. The text encoder converts input textual information (e.g., descriptions of defect types such as "cracks" and "potholes") into an embedding vector that is aligned with the image representation in semantic space.

[0048] Step S103: input the image information and semantic information into the multimodal large model to determine the road damage information.

[0049] In the embodiments of this application, the multimodal large model is a deep learning model that can simultaneously process data from different modalities (such as images and text). It integrates image information and semantic information, leveraging their complementarity to improve the accuracy and efficiency of disease identification.

[0050] In the embodiments of this application, image information and semantic information are passed as input to a large multimodal model to extract visual features from the image and semantic features from the text. The model then fuses these features to generate a comprehensive disease recognition result. This may include accurate classification of the disease type and determination of the disease location. The large multimodal model can then output an image with the disease recognition results annotated.

[0051] The multimodal large-scale model-based road defect detection method provided in the embodiments of the present application obtains road image information and textual information about defect types contained in the image information. The textual information is then fed into a semantic recognition model to determine semantic information. The image information and semantic information are then fed into the multimodal large-scale model to determine road defect information. This method enables road defect detection based on multimodal information, improving detection accuracy.

[0052] In some embodiments, the multimodal large model includes: a backbone network, a neck network and a detection head. The input of the backbone network is image information. The backbone network is used to extract multi-dimensional feature information based on the image information and output the multi-dimensional feature information to the neck network. The input of the neck network is multi-dimensional feature information and semantic information. The neck network is used to determine the fused feature information based on the multi-dimensional feature information and semantic information, and output the fused feature information to the detection head. The detection head is used to detect and output road disease information based on the fused feature information.

[0053] In the embodiments of this application, the backbone network is the foundation of the multimodal large model, responsible for processing the input image information. It is typically a deep convolutional neural network (CNN) that extracts multi-level, multi-dimensional feature information from the image. This feature information includes local image details (such as edges and texture) and global structure (such as shape and layout). The neck network, located between the backbone network and the detection head, is responsible for further processing and fusing information from different sources. The neck network's input includes the multi-dimensional feature information extracted by the backbone network and the semantic information from the semantic recognition module. The neck network's task is to effectively fuse this information to generate more expressive fused feature information. The detection head is the final output component of the multimodal large model and is responsible for defect detection based on the fused feature information. It typically comprises a series of fully connected or convolutional layers that perform classification, regression, or other processing on the fused feature information to output road defect information. This information may include defect type, location, size, etc. The fused feature information refers to the feature representation obtained by the neck network fusing the multi-dimensional feature information with the semantic information. This fusion is usually achieved through feature splicing, feature addition, attention mechanism, etc., aiming to effectively combine information from different sources and improve the detection performance of the model.

[0054] In the embodiment of the present application, the neck network can be an improved network based on the basic structure of the Feature Pyramid Network (FPN). The detection head can include: multiple classifiers and regressors for classifying and regressing the fused feature information.

[0055] In some embodiments, Figure 2 A schematic diagram of the overall structure of a multimodal model provided in an embodiment of the present application is shown in FIG. Figure 2As shown, the backbone network includes: multiple deformable adaptive lightweight channel segmentation and reassembly (ALCSR, Adaptive Lightweight Channel Splitting and Reassembly) modules, each ALCSR module is used to perform adaptive channel segmentation on the input feature information, obtain the feature information of each channel, and merge the feature information of each channel for output, each channel exchanges information between channels through the channel rearrangement mechanism, each channel is provided with a deformable convolution module, the deformable convolution module is used to change the shape of the feature information in the corresponding channel and then output it, multiple ALCSR modules include: a first ALCSR module, a second ALCSR module, a third ALCSR module, a fourth ALCSR module and a fifth ALCSR module, the backbone network also includes: a first CBS module, a second CBS module, a third The CBS module and the fourth CBS module, the first CBS module, the second CBS module, the third CBS module, and the fourth CBS module are connected in sequence. The input of the first CBS module is image information, the output of the fourth CBS module is the input of the first ALCSR module, the input of the second ALCSR module is the output of the first ALCSR module, the output of the second ALCSR module is the input of the third ALCSR module and the neck network, the output of the third ALCSR module is the input of the fourth ALCSR module, the output of the fourth ALCSR module is the input of the fifth ALCSR and the neck network, and the output of the fifth ALCSR is the input of the neck network.

[0056] In the embodiment of the present application, the backbone network adopts an adaptive lightweight channel segmentation and rearrangement (ALCSR) module, and combines it with deformable convolution to achieve efficient feature extraction and information interaction. The ALCSR module optimizes feature extraction through adaptive channel segmentation and strengthens information exchange between channels through the channel rearrangement mechanism. At the same time, the introduction of deformable convolution, combined with learnable offsets, improves the model's perception of object shape, size changes and deformed targets. In order to further enhance the global modeling capability of the model, an efficient hybrid encoder architecture is designed. This architecture combines the advantages of convolutional neural networks (CNNs) in local feature extraction and the powerful capabilities of Transformers in global context modeling, thereby achieving remarkable results in the comprehensive modeling of local details and global semantics, and greatly enhancing the generalization ability of the model.

[0057] In the embodiment of this application, Figure 3 A schematic diagram of the structure of a deformable ALCSR module provided in an embodiment of the present application is shown. Figure 4 A structural diagram of a deformable ALCSR module provided in an embodiment of the present application is shown as follows: Figure 3 and Figure 4 As shown, the deformable ALCSR module can be divided into two structures. If the size of the feature information needs to be changed, you can use Figure 4 If the structure shown does not need to change the size of the feature, you can use Figure 3 The structure shown.

[0058] In the embodiment of the present application, the first ALCSR module, the second ALCSR module, the third ALCSR module, the fourth ALCSR module and the fifth ALCSR module are connected in a specific order to achieve gradual extraction and transformation of feature information.

[0059] In the embodiment of the present application, the ALCSR module includes: a segmentation module, a basic network, a multi-layer network, a splicing unit and a channel mixing / channel rearrangement unit. The segmentation module is used to segment the features and input them into the basic network and the multi-layer network respectively. The output of the basic network and the multi-layer network is the input of the splicing unit, and the output of the splicing unit is the input of the channel mixing / channel rearrangement unit, and finally the processed features are output. In the basic network, the convolution operation may also include a pooling operation. In the multi-layer network, the DW Bottleneck, that is, the depthwise separable convolution bottleneck (Depthwise SeparableConvolution Bottleneck) is included.

[0060] In the embodiment of the present application, the ALCSR module can also be called a Deformable-ALCSR module. The ALCSR module adopts an adaptive feature processing strategy according to different feature levels, and achieves efficient feature extraction by adjusting the channel segmentation ratio and operation mechanism. At the lower feature level, the module focuses on the extraction of fine-grained information, adopts a smaller channel segmentation ratio, and allows most channels to be processed by a multi-layer network to enhance feature capture capabilities. At the higher feature level, the module focuses on the extraction of abstract semantic features, uses a larger channel segmentation ratio and introduces jump connections to retain high-level semantic information and optimize computational efficiency. In addition, the module combines bottleneck structure and deep convolution to achieve a lightweight design. After introducing deformable convolution into the ALCSR module, the module can learn spatial offset, enabling it to better perceive the shape, scale changes and complex deformations of objects. This further enhances the module's perception of irregular shapes, different scales and complex deformation diseased areas, providing efficient and reliable feature extraction support for open vocabulary instance segmentation tasks.

[0061] In the embodiments of the present application, the CBS module is a fundamental component of convolutional neural networks (CNNs) in the field of deep learning, particularly in object detection algorithms. It comprises the following key components: Conv layer (convolution layer): Responsible for performing convolution operations and extracting local features of the input data. Batch Normalization layer (BN layer): Batch normalization of the output of the convolution layer helps accelerate training and improve model stability. Activation function: Typically located after the BN layer, it introduces nonlinear factors and enhances the model's expressiveness. Different CBS module implementations may use different activation functions, such as Sigmoid and ReLU. CBS modules play a crucial role in object detection algorithms. By stacking multiple CBS modules, a convolutional neural network with powerful feature extraction capabilities can be constructed. These features are crucial for subsequent tasks such as object detection and classification. These modules are connected in sequence to form a feature extraction pipeline. The input of the first CBS module is image information, and the output of the fourth CBS module serves as the input to the first ALCSR module.

[0062] In an embodiment of the present application, the convolution kernel K, step size S, padding P and grouping G of the CBS module can be configured. For example, K can be configured to 3, S can be 2, P can be 1, and G can be 1.

[0063] In the embodiment of the present application, after being processed by a series of ALCSR modules and CBS modules, the backbone network outputs the transformed feature information to the neck network for further processing.

[0064] In some embodiments, see Figure 2, the neck network includes: multiple image-text feature alignment TITFA (Transformer-based image-text feature alignment) modules based on the attention mechanism, each TITFA module is used to fuse the input semantic information and feature information, and the multiple TITFA modules include: a first TITFA module, a second TITFA module, and a third TITFA module. The neck network also includes: a semantic information interaction module based on the attention mechanism, a first upsampling module, a first feature splicing module, a sixth ALCSR module, a second upsampling module, a second feature splicing module, a seventh ALCSR module, a fifth CBS module, a third feature splicing module, an eighth ALCSR module, a sixth CBS module, a fourth feature splicing module and a ninth ALCSR module. The input of the semantic information interaction module is the output of the fifth ALCSR module, the output of the semantic information interaction module is the input of the first upsampling module, the input of the first feature splicing module is the output of the first upsampling module and the fourth ALCSR module, the output of the first feature splicing module is the input of the sixth ALCSR module, and the output of the sixth ALCSR module is the output of the second upsampling module. The input of the first TITFA module is the input of the second splicing module and the third feature splicing module, the input of the second splicing module is the output of the second ALCSR module and the second upsampling module, the input of the first TITFA module is the input of the second splicing module and semantic information, the output of the first TITFA module is the input of the seventh ALCSR module, the output of the seventh ALCSR module is the input of the fifth CBS module and the detection head, the output of the fifth CBS module is the input of the third feature splicing module, the input of the second TITFA module is the semantic information and the output of the third feature splicing module, the output of the second TITFA module is the input of the eighth ALCSR module, the output of the eighth ALCSR module is the input of the sixth CBS module and the detection head, the input of the fourth splicing module is the output of the fifth ALCSR module and the sixth CBS module, the input of the third TITFA module is the semantic information and the output of the fourth splicing module, the output of the third TITFA module is the input of the ninth ALCSR module, and the output of the ninth ALCSR module is the input of the detection head.

[0065] In the embodiment of the present application, each feature splicing module in the figure is represented by Concat.

[0066] In the embodiment of the present application, the first, second, and third TITFA modules all receive input from different sources and output processed feature information. The upsampling module upsamples the feature map to increase its resolution. The feature concatenation module concatenates feature maps from different modules to fuse different feature information. The ALCSR module can function as a downsampling module.

[0067] In an embodiment of the present application, the attention-based semantic information interaction module may be referred to as a TSII (Transformer-based semantic information interaction) module.

[0068] In this embodiment of the present application, the sixth, seventh, and eighth ALCSR modules in the neck network are used as downsampling modules, supporting a variety of downsampling mechanisms, such as convolution with a stride of 2, pooling, and their combination, to flexibly adapt to the requirements of different resolutions. The channel rearrangement mechanism enhances information interaction between channels, thereby improving the diversity and expressiveness of feature representation.

[0069] In some embodiments, Figure 5 A structural diagram of a semantic information interaction module provided in an embodiment of the present application is shown as follows: Figure 5 As shown, the semantic information interaction module includes: multiple shape change units and multiple residual connection units, the shape change unit (reshape) is used to change the shape of the feature map, the residual connection unit (Add&Norm) is used to retain the original information, the multiple shape change units include: a first shape change unit and a second shape change unit, the multiple residual connection units include: a first residual connection unit and a second residual connection unit, the semantic information interaction module also includes: a first multi-head attention unit, a first multi-layer perceptron unit, the input of the first shape change unit is the output of the fifth ALCSR module, the output of the first shape change unit is the input of the first multi-head attention unit, the first multi-head attention unit is used to process the input information based on the multi-head attention mechanism, and output feature information to the first residual connection unit, the first residual connection unit is used to residually connect the information input to the semantic information interaction module with the features output by the first multi-head attention unit, the output of the first residual connection unit is the input of the first multi-layer perceptron unit, the first multi-layer perceptron unit performs nonlinear mapping and outputs feature information to the second residual connection unit, the output of the second residual connection unit is the input of the second shape change unit, and the output of the second shape change unit is the input of the first upsampling module.

[0070] In the embodiment of the present application, each shape change unit is used to change the shape of the feature map to meet the needs of the subsequent processing module, and each residual connection unit is used to retain the original information to avoid losing important features in the deep layer of the network. The first multi-head attention unit processes the input information based on the multi-head attention mechanism to capture the correlation between features.

[0071] In an embodiment of the present application, the output of the fifth ALCSR module is first input into the first shape changing unit to change the shape of the feature map. The output of the first shape changing unit is input into the first multi-head attention unit, and the features are processed based on the multi-head attention mechanism. The output of the first multi-head attention unit is input into the first residual connection unit, and residual connection is performed with the information input into the semantic information interaction module. The output of the first residual connection unit is input into the first multi-layer perceptron unit for nonlinear mapping. The output of the first multi-layer perceptron unit is input into the second shape changing unit to change the shape of the feature map again. The output of the second shape changing unit is ultimately used as the input of the first upsampling module. The first multi-head attention unit includes multiple shape changing units, multiple Einstein summation convention units, regularization units, etc.

[0072] In some embodiments, Figure 6 A schematic diagram of the structure of a TITFA module provided in an embodiment of the present application is shown in FIG. Figure 6 As shown, the TITFA module includes: a convolution unit, a third shape change unit, a second multi-head attention unit, a third residual connection unit, a second multi-layer perceptron unit, a fourth residual connection unit and a fourth shape change unit. The convolution unit is used to convolve the semantic information and output the convolved semantic information to the second multi-head attention unit. The third shape change unit is used to change the shape of the feature information of the image and output the changed feature information to the second multi-head attention unit. The second multi-head attention unit performs feature fusion based on the multi-head attention mechanism and outputs the fused feature information to the third residual connection unit. The output of the third residual connection unit is the second multi-layer perceptron unit. The output of the second multi-layer perceptron unit is the fourth residual connection unit. The output of the fourth residual connection unit is the input of the fourth shape change unit. The fourth shape change unit is used to output the fused feature information.

[0073] In an embodiment of the present application, the TITFA module receives semantic information and feature information of the image as input, the semantic information is processed by the convolution unit, and then output to the second multi-head attention unit, and the feature information of the image is changed in shape by the third shape change unit, and then output to the second multi-head attention unit. The second multi-head attention unit fuses the input semantic information and image feature information based on the multi-head attention mechanism, and outputs it to the third residual connection unit. The third residual connection unit performs residual connection and passes the output to the second multi-layer perceptron unit. The second multi-layer perceptron unit performs nonlinear mapping on the features and outputs them to the fourth residual connection unit. The fourth shape change unit performs final shape change or processing on the features and outputs the final fused feature information of the TITFA module.

[0074] In an embodiment of the present application, during the feature fusion stage, the TITFA module generates query (Q), key (K), and value (V) matrices for the semantic information of the image and text, respectively. These matrices are used in the self-attention mechanism to model the correlation within each modality. The self-attention mechanism of text features helps to grasp the contextual information of the input semantics. In the cross-attention mechanism, the query vector of the image feature interacts with the key and value vectors of the text feature, and through weighted operations, effective fusion of cross-modal semantic information is achieved. This design allows text semantic information to guide the optimization of image features at the spatial and semantic levels, especially in open vocabulary scenarios, significantly enhancing the model's ability to understand diverse target instances. Next, the fused features are nonlinearly mapped through a multi-layer perceptron (MLP), and residual connections (Add&Norm) are used to retain the original information, while further strengthening the semantic expression after interaction. Finally, these optimized features are mapped back to the image space.

[0075] The method provided in the embodiments of this application fully leverages the global modeling capabilities and attention mechanisms of the Transformer, successfully integrating the multimodal knowledge of the Chinese CLIP model into the instance segmentation task. This not only improves the model's ability to handle diverse instance descriptions, but also enhances its robustness and generalization in open vocabulary scenarios. Through the deep fusion of image and text, the model demonstrates excellent adaptability to complex instance segmentation tasks, providing an efficient solution for visual tasks in the open vocabulary domain.

[0076] The neck network provided in the embodiment of the present application cleverly combines the advantages of convolutional neural networks (CNNs) and transformers using an efficient hybrid encoder architecture. CNNs excel at capturing local features and spatial relationships, while transformers have excellent global context modeling capabilities and can effectively handle long-distance dependencies and semantic associations. By integrating these two technologies, the model retains the ability to extract accurate local features while improving its understanding of global information, significantly enhancing generalization performance and application scope.

[0077] In the backbone network, we optimized the overall architecture based on the Feature Pyramid Network (FPN). For smaller feature maps, we leveraged the Transformer's global context modeling advantages, specifically applying self-attention to semantically rich high-level feature maps. This design effectively captures long-range dependencies between different conceptual entities, helping subsequent modules more accurately localize and recognize objects. Furthermore, the Transformer aggregates global information at the high-level feature level, preventing object ambiguity and misclassification caused by the loss of local features. However, in larger feature maps, due to the sparse semantic information and the potential for redundancy and confusion when interacting with high-level features, we intentionally avoided excessive interactions within low-level features in our design. This strategy not only reduces computational cost but also prevents the accumulation of meaningless features, thereby improving model efficiency and performance. Furthermore, the FPN's multi-scale feature aggregation capabilities are fully utilized in the architecture. By integrating feature maps of different scales layer by layer, we ensure model consistency across different spatial resolutions, providing stronger support and higher robustness for instance segmentation tasks. This hybrid encoder architecture overcomes the shortcomings of traditional networks in global information modeling while avoiding resource waste in the global modeling process.

[0078] In some embodiments, the detection head includes: a lightweight image text feature fusion LITFF module, Figure 7 A schematic diagram of the structure of a LITFF module provided in an embodiment of the present application is shown in FIG. Figure 7 As shown, the LITFF module includes: multiple convolution units and normalization units, the convolution unit is used to perform dimensionality reduction and scale optimization on the image features of the fused features, the normalization unit is used to normalize the features in the numerical space, the multiple convolution units include: a first convolution unit and a second convolution unit, the multiple normalization units include: a first normalization unit and a second normalization unit, the LITFF module also includes: an Einstein summation convention unit, a parameter adjustment unit and a splicing unit, the input of the first convolution unit and the second convolution unit is the fused image feature, the output of the first convolution unit is the input of the first normalization unit, the input of the second normalization unit is semantic information, the input of the Einstein summation convention unit is the output of the first normalization unit and the second normalization unit, the output of the Einstein summation convention unit is the parameter adjustment unit, the input of the splicing unit is the output of the second convolution unit and the parameter adjustment unit, the splicing unit is used to output joint feature information and perform detection based on the joint feature information.

[0079] In the embodiments of the present application, a convolution unit is used to perform dimensionality reduction and scale optimization on the image features of the fused features. The convolution operation can extract local features from the image and achieve feature transformation and dimensionality reduction using different convolution kernels. The normalization unit is used to normalize the features in the numerical space to accelerate the training process and improve the generalization ability of the model. The normalization operation can distribute the feature values ​​within a reasonable range, thereby avoiding problems such as vanishing or exploding gradients. The Einstein summation convention is a concise way to represent element-wise operations between multidimensional arrays (such as tensors). In the LITFF module, the Einstein summation convention unit may be used to perform some form of element-wise operation, such as summation or weighted summation, on the outputs of the first and second normalization units to achieve feature fusion. The parameter adjustment unit further adjusts the parameters of the Einstein summation convention unit, such as scaling and translation, to meet the needs of subsequent processing. The splicing unit is used to splice the output of the second convolution unit with the output of the parameter adjustment unit to form joint feature information. The concatenation operation can combine features from different sources or different processing stages to enrich the feature representation.

[0080] In an embodiment of the present application, the LITFF module receives information after the fusion of image information and semantic information as input, and the fused image information is subjected to dimensionality reduction processing and scale optimization through the first convolution unit, and then output to the first normalization unit for normalization. The semantic information is normalized through the second normalization unit, and the output of the first normalization unit and the output of the second normalization unit are fused through the Einstein summation convention unit. The output of the Einstein summation convention unit is further parameter-adjusted through the parameter adjustment unit. On the other hand, the fused image information is also processed by another path through the second convolution unit and output to the splicing unit. The splicing unit splices the output of the second convolution unit and the output of the parameter adjustment unit to form joint feature information. The joint feature information is used as the final output of the LITFF module for subsequent detection tasks.

[0081] In the embodiment of the present application, the LITFF module is an efficient decoder head, the core process of which includes extracting features from images and text respectively and efficiently fusing them to generate accurate open vocabulary target area extraction.

[0082] First, the input image features (feature_map) are processed through independent convolutional modules (Conv ModuleList) and normalization operations (normalize). The convolutional modules perform dimensionality reduction and multi-scale optimization on the image features, while the normalization of the text features ensures cross-modal feature consistency in numerical space, providing a foundation for subsequent fusion operations. The pre-processed image and text data are then subjected to the einsum operation for feature interaction and fusion. This operation directly aligns the semantic information of the image and text modalities across the feature dimensions, enabling the semantic guidance of the text to be accurately mapped into the image space. The fused features are then further adjusted using the logit_scale and bias parameters to enhance inter-modal correlation and reduce noise interference. Subsequently, the semantic information of the open vocabulary is concatenated with the boundary regression information of the image features to form a joint feature representation. These joint features not only contain object category information but also clarify their spatial location through boundary features. The concatenated features are then input into the subsequent network, which, incorporating an anchor-free design principle, directly predicts the ground-truth bounding box of the object.

[0083] In some embodiments, the inspection head also includes a Proto network, which efficiently generates segmentation masks for each category, thereby completing the instance segmentation task. The Proto network processes the fused features and restores spatial resolution through upsampling to generate a set of potential segmentation masks. Guided by category prediction and object localization, these masks are assigned corresponding semantic information and instance attribution, thereby achieving more precise segmentation results based on object detection.

[0084] In an embodiment of the present application, a multi-task loss function is used in the network training process, specifically including: Boxes loss for optimizing bounding box regression, DFL (Distribution FocalLoss) loss for improving bounding box prediction accuracy, classification loss for ensuring target classification accuracy, and segmentation loss for generating segmentation masks. The joint design of these loss functions promotes the collaborative optimization between the detection and segmentation modules, thereby improving the overall performance of the model. Thanks to the decoupling of detection and segmentation tasks in the design of the network architecture, the system has a high degree of flexibility and scalability. For detection tasks, the bounding box regression and classification modules can be run separately to achieve target positioning; in segmentation tasks, the segmentation module can be combined to output a high-precision segmentation mask; for instance segmentation tasks, through the collaborative operation of the detection and segmentation modules, an accurate segmentation mask can be output while generating the target box.

[0085] In some embodiments, before step S103, the method further includes:

[0086] Step S1: Obtain a sample data set.

[0087] In the embodiment of the present application, the image resolution of the sample data set is as high as 1920×1080. The data set covers the shooting period from 7 am to 8 pm, and includes images under various weather and lighting conditions. A total of N original pictures are included, which record a variety of road damage conditions in detail. In addition, these images are refined instance segmentation and annotation with the help of the SAM2 model, and text label information is added. The annotated categories are mainly divided into three categories: sidewalk diseases, asphalt pavement diseases, and cement pavement diseases, and are further subdivided into 36 detailed road disease types. In order to protect personal privacy, all facial and license plate information involved in the data set have been mosaiced.

[0088] Step S2: training a large multimodal model based on the sample data set.

[0089] In the embodiment of the present application, a large multimodal model can be obtained by training using a sample data set.

[0090] After training is completed, the multimodal large model can be deployed to detect road defects.

[0091] The method provided in the embodiments of this application features a large multimodal model with efficient feature extraction, good generalization performance, and continuous learning capabilities. It utilizes an adaptive lightweight channel segmentation and reordering (ALCSR) module combined with deformable convolution. The ALCSR module optimizes feature extraction through adaptive channel segmentation and enhances information exchange between channels through a channel reordering mechanism. Combined with deformable convolution, it introduces a learnable offset, improving the model's ability to perceive object shape, size changes, and deformation. The neck network utilizes an efficient hybrid encoder architecture, retaining the advantages of convolutional neural networks (CNNs) in local feature extraction while incorporating the powerful global context modeling capabilities of the Transformer, thereby enhancing the model's generalization performance. The image-text fusion module, based on the Transformer architecture, combines self-attention and cross-attention mechanisms, allowing text information to effectively guide image feature optimization. A lightweight open-vocabulary instance segmentation decoder head is also provided, optimizing computational efficiency and enhancing the adaptability and accuracy of instance segmentation in diverse environments.

[0092] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0093] According to the aforementioned embodiments, the embodiments of the present application provide a road defect detection device based on a multimodal large model. The modules included in the device, and the units included in each module, can be implemented by a processor in a computer device; of course, they can also be implemented by a specific logic circuit. During implementation, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0094] The present invention provides a road damage detection device based on a multimodal large model. Figure 8 A schematic diagram of the structure of a road disease detection device based on a multimodal large model provided in an embodiment of the present application is shown in FIG. Figure 8 As shown, the road disease detection 800 based on the multimodal large model includes:

[0095] The acquisition module 801 is used to acquire image information of the road and text information about the disease type in the image information.

[0096] A determination module 802 is configured to output the text information to a semantic recognition model to determine semantic information;

[0097] The detection module 803 is used to input the image information and the semantic information into a multimodal large model to determine the road damage information.

[0098] In some embodiments, the multimodal large model includes: a backbone network, a neck network and a detection head. The input of the backbone network is the image information. The backbone network is used to extract multi-dimensional feature information based on the image information and output the multi-dimensional feature information to the neck network. The input of the neck network is the multi-dimensional feature information and the semantic information. The neck network is used to determine the fused feature information based on the multi-dimensional feature information and the semantic information, and output the fused feature information to the detection head. The detection head is used to detect and output the road disease information based on the fused feature information.

[0099] In some embodiments, the backbone network includes: multiple adaptive lightweight channel segmentation and rearrangement ALCSR modules, each ALCSR module is used to perform adaptive channel segmentation on the input feature information, obtain feature information of each channel, and merge the feature information of each channel for output, and each channel exchanges information between channels through a channel rearrangement mechanism. The multiple ALCSR modules include: a first ALCSR module, a second ALCSR module, a third ALCSR module, a fourth ALCSR module and a fifth ALCSR module. The backbone network also includes: a first CBS module, a second CBS module, a first CBS module and a fourth CBS module. The first CBS The first module, the second CBS module, the first CBS module, and the fourth CBS module are connected in sequence, the input of the first CBS module is the image information, the output of the fourth CBS module is the input of the first ALCSR module, the input of the second ALCSR module is the output of the first ALCSR module, the output of the second ALCSR module is the input of the third ALCSR module and the neck network, the output of the third ALCSR module is the input of the fourth ALCSR module, the output of the fourth ALCSR module is the input of the fifth ALCSR and the neck network, and the output of the fifth ALCSR is the input of the neck network.

[0100] In some embodiments, the neck network includes: multiple image-text feature alignment TITFA modules based on the attention mechanism, each TITFA module is used to fuse the input semantic information and feature information, and the multiple TITFA modules include: a first TITFA module, a second TITFA module, and a third TITFA module. The neck network also includes: a semantic information interaction module based on the attention mechanism, a first upsampling module, a first feature splicing module, a sixth ALCSR module, a second upsampling module, a second feature splicing module, a seventh ALCSR module, a fifth CBS module, a third feature splicing module, an eighth ALCSR module, a sixth CBS module, a fourth feature splicing module and a ninth ALCSR module. The input of the semantic information interaction module is the output of the fifth ALCSR module, the output of the semantic information interaction module is the input of the first upsampling module, the input of the first feature splicing module is the output of the first upsampling module and the fourth ALCSR module, the output of the first feature splicing module is the input of the sixth ALCSR module, and the output of the sixth ALCSR module is the input of the second upsampling module. and the input of the third feature splicing module, the input of the second splicing module is the output of the second ALCSR module and the second upsampling module, the input of the first TITFA module is the input of the second splicing module and the semantic information, the output of the first TITFA module is the input of the seventh ALCSR module, the output of the seventh ALCSR module is the input of the fifth CBS module and the detection head, the output of the fifth CBS module is the input of the third feature splicing module, the input of the second TITFA module is the semantic information and the output of the third feature splicing module, the output of the second TITFA module is the input of the eighth ALCSR module, the output of the eighth ALCSR module is the input of the sixth CBS module and the detection head, the input of the fourth splicing module is the output of the fifth ALCSR module and the sixth CBS module, the input of the third TITFA module is the semantic information and the output of the fourth splicing module, the output of the third TITFA module is the input of the ninth ALCSR module, and the output of the ninth ALCSR module is the input of the detection head.

[0101] In some embodiments, the semantic information interaction module includes: multiple shape change units and multiple residual connection units, the shape change unit is used to change the shape of the feature map, the residual connection unit is used to retain the original information, the multiple shape change units include: a first shape change unit and a second shape change unit, the multiple residual connection units include: a first residual connection unit and a second residual connection unit, the semantic information interaction module also includes: a first multi-head attention unit, a first multi-layer perceptron unit, the input of the first shape change unit is the output of the fifth ALCSR module, the output of the first shape change unit is the input of the first multi-head attention unit, the first multi-head attention unit is used to process the input information based on the multi-head attention mechanism, and output feature information to the first residual connection unit, the first residual connection unit is used to residually connect the information input into the semantic information interaction module with the features output by the first multi-head attention unit, the output of the first residual connection unit is the input of the first multi-layer perceptron unit, the first multi-layer perceptron unit performs nonlinear mapping and outputs feature information to the second residual connection unit, the output of the second residual connection unit is the input of the second shape change unit, and the output of the second shape change unit is the input of the first upsampling module.

[0102] In some embodiments, the TITFA module includes: a convolution unit, a third shape change unit, a second multi-head attention unit, a third residual connection unit, a second multi-layer perceptron unit, a fourth residual connection unit and a fourth shape change unit, wherein the convolution unit is used to convolve the semantic information and output the convolved semantic information to the second multi-head attention unit, the third shape change unit is used to change the shape of the feature information of the image and output the changed feature information to the second multi-head attention unit, the second multi-head attention unit performs feature fusion based on the multi-head attention mechanism, and outputs the fused feature information to the third residual connection unit, the output of the third residual connection unit is the second multi-layer perceptron unit, the output of the second multi-layer perceptron unit is the fourth residual connection unit, the output of the fourth residual connection unit is the input of the fourth shape change unit, and the fourth shape change unit is used to output the fused feature information.

[0103] In some embodiments, the detection head includes: a lightweight image-text feature fusion LITFF module, the LITFF module includes: multiple convolution units and a normalization unit, the convolution unit is used to perform dimensionality reduction and scale optimization on the image features of the fused features, the normalization unit is used to normalize the features in the numerical space, the multiple convolution units include: a first convolution unit and a second convolution unit, the multiple normalization units include: a first normalization unit and a second normalization unit, the LITFF module also includes: an Einstein summation convention unit, a parameter adjustment unit and a splicing unit, the input of the first convolution unit and the second convolution unit is the fused image information, the output of the first convolution unit is the input of the first normalization unit, the input of the second normalization unit is the semantic information, the input of the Einstein summation convention unit is the output of the first normalization unit and the second normalization unit, the output of the Einstein summation convention unit is the parameter adjustment unit, the input of the splicing unit is the output of the second convolution unit and the parameter adjustment unit, the splicing unit is used to output joint feature information and perform detection based on the joint feature information.

[0104] in addition, Figure 8 The multimodal large model-based road defect detection shown can be a software unit, hardware unit, or a combination of software and hardware units built into existing electronic devices, or can be integrated into electronic devices as an independent pendant, or can exist as an independent terminal device.

[0105] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0106] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0107] Figure 9This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Figure 9 As shown, the electronic device 3 of this embodiment may include: at least one processor 30 ( Figure 9 Only one processor 30 is shown in the figure), a memory 31, and a computer program 32 stored in the memory 31 and executable on at least one processor 30. When the processor 30 executes the computer program 32, the steps of any of the above-mentioned method embodiments are implemented, or when the processor 30 executes the computer program 32, the functions of the modules / units in the above-mentioned device embodiments are implemented.

[0108] Exemplarily, the computer program 32 may be divided into one or more modules / units, one or more of which are stored in the memory 31 and executed by the processor 30 to implement the present application. The one or more modules / units may be a series of computer program 32 instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 32 in the electronic device 3.

[0109] The embodiment of the present application further provides a computer-readable storage medium, which stores a computer program 32. When the computer program 32 is executed by the processor 30, the steps in the above-mentioned method embodiments can be implemented.

[0110] An embodiment of the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0111] If the integrated unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the process steps in the above-mentioned method embodiments can be implemented by computer program 32 instructing the relevant hardware. Computer program 32 can be stored in a computer-readable storage medium. When executed by processor 30, computer program 32 can implement the steps of each of the above-mentioned method embodiments. Computer program 32 includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. Computer-readable media can include at least: any entity or device capable of carrying computer program code to a terminal, recording media, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, removable hard drives, magnetic disks, or optical disks. In some jurisdictions, based on legislation and patent practice, computer-readable media cannot be electric carrier signals or telecommunication signals.

[0112] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0113] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0114] In the embodiments provided in this application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0115] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0116] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

[0117] The relevant user personal information that may be involved in the various embodiments of this application is strictly in accordance with the requirements of laws and regulations, following the principles of legality, legitimacy and necessity, and based on the reasonable purposes of business scenarios, to process the personal information that users actively provide during the use of products / services or generated due to the use of products / services, as well as the personal information obtained with the user's authorization.

[0118] The personal information processed by the Applicant will vary depending on the specific product / service scenario and will be based on the specific scenario in which the user uses the product / service. This may involve the user's account information, device information, driving information, vehicle information, or other related information. The Applicant will treat the user's personal information and its processing with a high degree of diligence.

[0119] The Applicant attaches great importance to the security of user personal information and has taken reasonable and feasible security measures that comply with industry standards to protect user information and prevent personal information from being accessed, disclosed, used, modified, damaged or lost without authorization.

Claims

1. A road damage detection method based on a multimodal large model, characterized in that: include: Acquiring image information of the road and text information about the type of damage in the image information; Outputting the text information to a semantic recognition model to determine semantic information; Inputting the image information and the semantic information into a multimodal large model to determine the road damage information; The multimodal large model includes: a backbone network, a neck network and a detection head. The input of the backbone network is the image information. The backbone network is used to extract multi-dimensional feature information based on the image information and output the multi-dimensional feature information to the neck network. The input of the neck network is the multi-dimensional feature information and the semantic information. The neck network is used to determine fused feature information based on the multi-dimensional feature information and the semantic information, and output the fused feature information to the detection head. The detection head is used to detect and output the road disease information based on the fused feature information; the backbone network includes: a plurality of deformable convolution adaptive lightweight channel segmentation and rearrangement ALCSR modules, each ALCSR module is used to perform adaptive channel segmentation on the input feature information, obtain feature information of each channel, and merge and output the feature information of each channel. Each channel exchanges information between channels through a channel rearrangement mechanism. A deformable convolution module is provided in each channel, and the deformable convolution module is used to perform feature segmentation on the corresponding channel. The feature information is output after the shape is changed, and the multiple ALCSR modules include: a first ALCSR module, a second ALCSR module, a third ALCSR module, a fourth ALCSR module and a fifth ALCSR module. The backbone network also includes: a first CBS module, a second CBS module, a third CBS module and a fourth CBS module. The first CBS module, the second CBS module, the third CBS module and the fourth CBS module are connected in sequence. The input of the first CBS module is the image information, the output of the fourth CBS module is the input of the first ALCSR module, the input of the second ALCSR module is the output of the first ALCSR module, the output of the second ALCSR module is the input of the third ALCSR module and the neck network, the output of the third ALCSR module is the input of the fourth ALCSR module, the output of the fourth ALCSR module is the input of the fifth ALCSR module and the neck network, and the output of the fifth ALCSR module is the input of the neck network;The neck network includes: multiple image-text feature alignment TITFA modules based on the attention mechanism, each TITFA module is used to fuse the input semantic information and feature information, and multiple TITFA modules include: a first TITFA module, a second TITFA module, and a third TITFA module. The neck network also includes: a semantic information interaction module based on the attention mechanism, a first upsampling module, a first feature splicing module, a sixth ALCSR module, a second upsampling module, a second feature splicing module, a seventh ALCSR module, a fifth CBS module, a third feature splicing module, an eighth ALCSR module, a sixth CBS module, a fourth feature splicing module and a ninth ALCSR module. The input of the semantic information interaction module is the output of the fifth ALCSR module, the output of the semantic information interaction module is the input of the first upsampling module, the input of the first feature splicing module is the output of the first upsampling module and the fourth ALCSR module, the output of the first feature splicing module is the input of the sixth ALCSR module, and the output of the sixth ALCSR module is the output of the second upsampling module and the third feature splicing module. The input of the first TITFA module is the input of the second feature splicing module, the input of the second feature splicing module is the output of the second ALCSR module and the second upsampling module, the input of the first TITFA module is the input of the second feature splicing module and the semantic information, the output of the first TITFA module is the input of the seventh ALCSR module, the output of the seventh ALCSR module is the input of the fifth CBS module and the detection head, the output of the fifth CBS module is the input of the third feature splicing module, the input of the second TITFA module is the semantic information and the output of the third feature splicing module, the output of the second TITFA module is the input of the eighth ALCSR module, the output of the eighth ALCSR module is the input of the sixth CBS module and the detection head, the input of the fourth feature splicing module is the output of the fifth ALCSR module and the sixth CBS module, the input of the third TITFA module is the semantic information and the output of the fourth feature splicing module, the output of the third TITFA module is the input of the ninth ALCSR module, and the output of the ninth ALCSR module is the input of the detection head.

2. The method according to claim 1, characterized in that The semantic information interaction module includes: multiple shape change units and multiple residual connection units, the shape change unit is used to change the shape of the feature map, and the residual connection unit is used to retain the original information. The multiple shape change units include: a first shape change unit and a second shape change unit. The multiple residual connection units include: a first residual connection unit and a second residual connection unit. The semantic information interaction module also includes: a first multi-head attention unit and a first multi-layer perceptron unit. The input of the first shape change unit is the output of the fifth ALCSR module, and the output of the first shape change unit is the input of the first multi-head attention unit. The first multi-head attention unit is used to process the input information based on the multi-head attention mechanism and output feature information to the first residual connection unit. The first residual connection unit is used to perform a residual connection between the information input into the semantic information interaction module and the features output by the first multi-head attention unit. The output of the first residual connection unit is the input of the first multi-layer perceptron unit. The first multi-layer perceptron unit performs nonlinear mapping and outputs feature information to the second residual connection unit. The output of the second residual connection unit is the input of the second shape change unit. The output of the second shape change unit is the input of the first upsampling module.

3. The method according to claim 2, characterized in that The TITFA module includes: a convolution unit, a third shape change unit, a second multi-head attention unit, a third residual connection unit, a second multi-layer perceptron unit, a fourth residual connection unit and a fourth shape change unit. The convolution unit is used to convolve the semantic information and output the convolved semantic information to the second multi-head attention unit. The third shape change unit is used to change the shape of the feature information of the image and output the changed feature information to the second multi-head attention unit. The second multi-head attention unit performs feature fusion based on the multi-head attention mechanism and outputs the fused feature information to the third residual connection unit. The output of the third residual connection unit is the second multi-layer perceptron unit. The output of the second multi-layer perceptron unit is the fourth residual connection unit. The output of the fourth residual connection unit is the input of the fourth shape change unit. The fourth shape change unit is used to output the fused feature information.

4. The method according to claim 1, wherein The detection head includes: a lightweight image-text feature fusion LITFF module, the LITFF module includes: multiple convolution units and a normalization unit, the convolution unit is used to perform dimensionality reduction and scale optimization on the image features of the fused features, the normalization unit is used to normalize the features in the numerical space, the multiple convolution units include: a first convolution unit and a second convolution unit, the multiple normalization units include: a first normalization unit and a second normalization unit, the LITFF module also includes: an Einstein summation convention unit, a parameter adjustment unit and a splicing unit, the input of the first convolution unit and the second convolution unit is the fused image information, the output of the first convolution unit is the input of the first normalization unit, the input of the second normalization unit is the semantic information, the input of the Einstein summation convention unit is the output of the first normalization unit and the second normalization unit, the output of the Einstein summation convention unit is the parameter adjustment unit, the input of the splicing unit is the output of the second convolution unit and the parameter adjustment unit, the splicing unit is used to output joint feature information and perform detection based on the joint feature information.

5. A road disease detection device based on a multimodal large model, characterized in that: include: An acquisition module, configured to acquire image information of the road and text information about the type of damage in the image information; A determination module, configured to output the text information to a semantic recognition model to determine semantic information; A detection module is used to input the image information and the semantic information into a multimodal large model to determine the road disease information. The multimodal large model includes: a backbone network, a neck network and a detection head. The input of the backbone network is the image information. The backbone network is used to extract multi-dimensional feature information based on the image information and output the multi-dimensional feature information to the neck network. The input of the neck network is the multi-dimensional feature information and the semantic information. The neck network is used to determine fused feature information based on the multi-dimensional feature information and the semantic information, and output the fused feature information to the detection head. The detection head is used to detect and output the road disease information based on the fused feature information; the backbone network includes: multiple deformable convolution adaptive lightweight channel segmentation and rearrangement ALCSR modules, each ALCSR module is used to perform adaptive channel segmentation on the input feature information, obtain feature information of each channel, and merge and output the feature information of each channel. Each channel exchanges information between channels through a channel rearrangement mechanism, and each channel is provided with a deformable convolution. Module, the deformable convolution module is used to change the shape of the feature information in the corresponding channel and then output it, multiple ALCSR modules include: a first ALCSR module, a second ALCSR module, a third ALCSR module, a fourth ALCSR module and a fifth ALCSR module, the backbone network also includes: a first CBS module, a second CBS module, a third CBS module and a fourth CBS module, the first CBS module, the second CBS module, the third CBS module and the fourth CBS module are connected in sequence, the input of the first CBS module is the image information, the output of the fourth CBS module is the input of the first ALCSR module, the input of the second ALCSR module is the output of the first ALCSR module, the output of the second ALCSR module is the input of the third ALCSR module and the neck network, the output of the third ALCSR module is the input of the fourth ALCSR module, the output of the fourth ALCSR module is the input of the fifth ALCSR module and the neck network, and the output of the fifth ALCSR module is the input of the neck network;The neck network includes: multiple image-text feature alignment TITFA modules based on the attention mechanism, each TITFA module is used to fuse the input semantic information and feature information, and multiple TITFA modules include: a first TITFA module, a second TITFA module, and a third TITFA module. The neck network also includes: a semantic information interaction module based on the attention mechanism, a first upsampling module, a first feature splicing module, a sixth ALCSR module, a second upsampling module, a second feature splicing module, a seventh ALCSR module, a fifth CBS module, a third feature splicing module, an eighth ALCSR module, a sixth CBS module, a fourth feature splicing module and a ninth ALCSR module. The input of the semantic information interaction module is the output of the fifth ALCSR module, the output of the semantic information interaction module is the input of the first upsampling module, the input of the first feature splicing module is the output of the first upsampling module and the fourth ALCSR module, the output of the first feature splicing module is the input of the sixth ALCSR module, and the output of the sixth ALCSR module is the output of the second upsampling module and the third feature splicing module. The input of the first TITFA module is the input of the second feature splicing module, the input of the second feature splicing module is the output of the second ALCSR module and the second upsampling module, the input of the first TITFA module is the input of the second feature splicing module and the semantic information, the output of the first TITFA module is the input of the seventh ALCSR module, the output of the seventh ALCSR module is the input of the fifth CBS module and the detection head, the output of the fifth CBS module is the input of the third feature splicing module, the input of the second TITFA module is the semantic information and the output of the third feature splicing module, the output of the second TITFA module is the input of the eighth ALCSR module, the output of the eighth ALCSR module is the input of the sixth CBS module and the detection head, the input of the fourth feature splicing module is the output of the fifth ALCSR module and the sixth CBS module, the input of the third TITFA module is the semantic information and the output of the fourth feature splicing module, the output of the third TITFA module is the input of the ninth ALCSR module, and the output of the ninth ALCSR module is the input of the detection head.

6. An electronic device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 4 when executing the computer program.

7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Multi-source pavement disease identification method based on language and image large model

    CN119107447A