Method, apparatus, electronic device, and program product for detecting anomalies in urban facilities based on a multimodal model
Through the multimodal model, the multimodal model combines image and text features, and combined with local and global feature extraction, the problems of small sample learning and complex scene understanding in urban facility anomaly detection are solved, and higher detection accuracy and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510245196.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The prior art has problems such as small sample learning, data imbalance and difficulty in understanding complex scenarios in urban facilities anomaly detection, resulting in insufficient generalization capabilities and accuracy of the model.
Using a multimodal model, by introducing TgVL modules into the neck network to fuse images and text features, and setting up an AGLPU module in the backbone network and TgVL modules, local and global features are extracted, and multimodal information fusion is combined with TVF detection heads to improve detection accuracy.
The generalization ability and accuracy of the anomaly detection model are improved, and it can efficiently and accurately detect urban facilities abnormalities in complex scenarios and adapt to dynamically changing urban environments.
Smart Images

Figure CN119741614B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of smart cities, and particularly relates to an urban facility anomaly detection method based on a multimodal model, an urban facility anomaly detection device based on a multimodal model, an electronic device, and a computer program product. Background Art
[0002] With the continuous acceleration of the global urbanization process, the scale and complexity of urban infrastructure are increasing day by day, which requires a higher level of intelligence in urban management to ensure the efficient operation and sustainable development of the city. Inevitably, various facility anomalies will occur in the daily operation of the city, such as damage to different public safety facilities. If these urban facility anomalies cannot be quickly detected and effectively processed, it will not only reduce the urban management efficiency, but also directly reduce the quality of life of citizens and even threaten the safety of citizens.
[0003] Due to the variety and complexity of urban facility anomalies, it is difficult to obtain some sample data, resulting in a limited number of samples and uneven data of various types. Training a model with such samples may cause the model to be biased towards common anomalies and difficult to effectively detect rare anomalies, thereby affecting its generalization ability and accuracy. Summary of the Invention
[0004] This application provides an urban facility anomaly detection method, an urban facility anomaly detection device, an electronic device, and a computer program product, which can improve the generalization ability and accuracy of the anomaly detection model.
[0005] In a first aspect, this application provides an urban facility anomaly detection method based on a multimodal model, including:
[0006] Performing feature extraction on the image to be detected based on the backbone network of the pre-trained anomaly detection model to obtain image features; the image to be detected includes urban facilities;
[0007] Fusing the image features based on the neck network of the anomaly detection model to obtain a first fused feature;
[0008] Detecting the first fused feature based on the detection head of the anomaly detection model to obtain the detection result of the anomaly of the urban facilities in the image to be detected;
[0009] Wherein, the neck network is provided with a TgVL module, and the TgVL module is used to fuse the input first image feature and the text feature of the preset guiding text to obtain a second fused feature, and the first fused feature is obtained based on the second fused feature.
[0010] Furthermore, the TgVL module includes an image feature enhancement branch, a text-guided image branch, and a first fusion layer; for the first image feature and the text feature:
[0011] Process the first image feature through the image feature enhancement branch to obtain an enhanced image feature;
[0012] Perform a summation operation and a weighted activation operation on the preprocessed first image feature and the text feature in sequence through the text-guided image branch to obtain a guided learning weight;
[0013] Multiply the enhanced image feature and the guided learning weight through the first fusion layer and perform a convolution operation to obtain a second fusion feature corresponding to the first image feature and the text feature.
[0014] Furthermore, the text-guided image branch includes an image feature preprocessing sub-branch, a text feature preprocessing sub-branch, an Einstein summation layer, and a MaxSigmoid activation function layer; performing an Einstein summation operation and a weighted activation operation on the preprocessed first image feature and the text feature in sequence through the text-guided image branch to obtain a guided learning weight, including:
[0015] Perform a convolution operation and a reshaping operation on the first image feature in sequence through the image feature preprocessing sub-branch to obtain the preprocessed first image feature;
[0016] Perform a linear operation and a reshaping operation on the text feature in sequence through the text feature preprocessing sub-branch to obtain the preprocessed text feature;
[0017] Perform an Einstein summation operation on the preprocessed first image feature and the preprocessed first image feature through the Einstein summation layer to obtain a third fusion feature;
[0018] Perform a MaxSigmoid function activation operation on the third fusion feature through the MaxSigmoid activation function layer to obtain a guided learning weight.
[0019] Furthermore, the detection head includes a regression branch, a classification branch, and a splicing layer. The detection head based on the anomaly detection model detects the first fusion feature to obtain a detection result of urban facility anomalies in the image to be detected, including:
[0020] Perform a depth convolution operation on the first fusion feature through the regression branch to obtain a regression result;
[0021] Perform a summation operation and a scaling operation on the classified first fusion feature and the preprocessed text feature in sequence through the classification branch to obtain a classification result;
[0022] Splice the regression result and the classification result through the splicing layer and activate based on a preset activation function to obtain a visual detection result.
[0023] Further, an AGLPU module is provided in the backbone network and / or the TgVL module. The AGLPU module is configured to extract and fuse the local features and global features of the input second image features to obtain a fourth fused feature.
[0024] Further, the AGLPU module includes a convolutional layer, a local perception unit branch, a global perception unit branch, and a second fusion layer. For each second image feature:
[0025] Perform a convolution operation on the second image feature through the convolutional layer to obtain a convolutional feature;
[0026] Extract features from the convolutional feature through the local perception unit branch to obtain local features;
[0027] Extract features from the convolutional feature through the global perception unit to obtain global features;
[0028] Perform adaptive weighted fusion on the local features and global features through the second fusion layer to obtain a fourth fused feature.
[0029] Further, the local perception unit branch is composed of two bottleneck structures with residual connections in series; extract features from the convolutional feature through the local perception unit branch to obtain local features;
[0030] Perform convolution operations, depthwise convolution operations, and convolution operations on the convolutional feature in sequence to obtain a local first sub-feature;
[0031] Concatenate the convolutional feature with the local first sub-feature to obtain a first concatenated feature;
[0032] Perform convolution operations, depthwise convolution operations, and convolution operations on the first concatenated feature in sequence to obtain a local second sub-feature;
[0033] Concatenate the first concatenated feature with the local second sub-feature to obtain local features.
[0034] Further, the global perception unit branch is composed of an LMHSA structure with residual connections and an FFN structure with residual connections in series; extract features from the convolutional feature through the global perception unit to obtain global features, including:
[0035] Perform layer normalization operations, multi-head self-attention operations, and regularization operations on the convolutional feature in sequence to obtain a global first sub-feature;
[0036] Concatenate the convolutional feature with the global first sub-feature to obtain a second concatenated feature;
[0037] Perform layer normalization operations, feature transformation operations, and regularization operations on the second concatenated feature to obtain a global second sub-feature;
[0038] Concatenate the second splicing feature with the global second sub - feature to obtain the global feature.
[0039] In a second aspect, the present application provides an urban facility anomaly detection device based on a multimodal model, including:
[0040] A feature extraction module, configured to extract features from the image to be detected based on the backbone network of the pre - trained anomaly detection model to obtain image features; the image to be detected includes urban facilities;
[0041] A feature fusion module, configured to fuse the image features based on the neck network of the anomaly detection model to obtain the first fusion feature;
[0042] A detection module, configured to detect the first fusion feature based on the detection head of the anomaly detection model to obtain the detection result of the anomaly of urban facilities in the image to be detected;
[0043] Among them, the neck network is provided with a TgVL module, and the TgVL module is used to fuse the input first image feature and the text feature of the preset guiding text to obtain the second fusion feature, and the first fusion feature is obtained based on the second fusion feature.
[0044] In a third aspect, the present application provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method in the first aspect are implemented.
[0045] In a fourth aspect, the present application provides a computer - readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the method in the first aspect are implemented.
[0046] In a fifth aspect, the present application provides a computer program product, which includes a computer program. When the computer program is executed by one or more processors, the steps of the method in the first aspect are implemented.
[0047] The beneficial effects of the first aspect of the present application compared with the prior art are as follows: The present application sets up a Text-guided Visual Learning module (TgVL) in the neck network. By fully integrating the features of both image and text modalities, it can enhance the expressiveness of effective image features under the guidance of text. That is, the TgVL module can enable the model to extract more expressive and richer effective features, namely the second fusion features, by using text information to guide image feature learning. The first fusion feature determined based on the second fusion feature is used for urban facility anomaly detection, which can enable the anomaly detection model to obtain stronger generalization ability and higher accuracy.
[0048] It can be understood that the beneficial effects of the above-mentioned second aspect to fifth aspect can refer to the relevant descriptions in the above-mentioned first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0050] Figure 1 It is a schematic diagram of some urban facility anomalies provided by the embodiments of the present application;
[0051] Figure 2 It is a schematic diagram of the network structure of the OV-MFS-YOLO network provided by the embodiments of the present application;
[0052] Figure 3 It is a schematic diagram of the angular loss provided by the embodiments of the present application;
[0053] Figure 4 It is a schematic diagram of the distance loss provided by the embodiments of the present application;
[0054] Figure 5 It is a schematic flowchart of the urban facility anomaly detection method based on a multi-modal model provided by the embodiments of the present application;
[0055] Figure 6 It is a schematic diagram of the network structure of the TgVL module provided by the embodiments of the present application;
[0056] Figure 7 It is a schematic diagram of the network structure of the TVF detection head provided by the embodiments of the present application;
[0057] Figure 8 It is a schematic diagram of the network structure of the AGLPU module provided by the embodiments of the present application;
[0058] Figure 9 is a schematic structural diagram of an urban facility anomaly detection device based on a multimodal model provided by an embodiment of the present application;
[0059] Figure 10 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0060] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from hindering the description of the present application.
[0061] There are a wide variety of urban facilities, including roads, manhole covers, guardrails, walls, expansion devices, drainage facilities, isolation fences, anti-collision columns, curb stones, etc. Based on the abnormal conditions that occur to these urban facilities, the types and manifestations of anomalies are also more diverse and complex.
[0062] This application focuses on identifying and classifying 152 types of facility abnormal conditions in the urban environment, aiming to comprehensively and accurately address various problems faced by urban management. These abnormal conditions can be divided into four major categories: First is the clean category, covering 30 situations such as garbage scattering, trash can overflow, chaotic trash can placement, and drainage facility blockage; second is the order category, including 14 situations such as random parking of non-motor vehicles, non-motor vehicle tipping, unregulated street vending (mobile vendors), and road occupation for business; the third category is the tidy category, involving 83 situations such as stains on the exterior facades of street buildings, damaged or missing manhole covers, damaged or missing store signs, and damaged anti-collision barrels; finally is the greening and aesthetics category, including 25 situations such as greening garbage, messy roads caused by construction, fallen trees, and unclean flower pots, flower boxes, and flower beds.
[0063] Figure 1 Shows some representative actual case pictures to more clearly display the actual situations of various abnormal situations.
[0064] To improve the detection accuracy of the abnormal detection model for urban facility anomalies, this application constructs a professional dataset called Abnormal City Facilities (ACF), focusing on collecting and analyzing urban facility anomaly instances. The uniqueness of this dataset lies in that all image data comes from on-site collection in a real vehicle-mounted environment, ensuring the authenticity and diversity of the data.
[0065] The number of images in the dataset is closely related to the performance of the trained anomaly detection model. Suppose the dataset contains 1 million high-quality images. To ensure the effectiveness and generalization ability of the algorithm model, the dataset will follow the classic 8:2 division ratio: 800,000 images are used as the training set to train and optimize the anomaly detection model; the remaining 200,000 images are used as the validation set to verify the detection ability, accuracy, and generalization ability of the model, ensuring the performance of the model on unknown data.
[0066] Optionally, to comprehensively consider urban facility anomalies in different environments, the collected images cover a variety of lighting conditions (such as strong light, weak light, backlight, etc.), different weather conditions (such as sunny, rainy, snowy, foggy, etc.), and road and environmental characteristics of most cities across the country. This diverse data collection method ensures the authenticity, richness, and comprehensiveness of the dataset, providing a more accurate and reliable basis for related research and applications.
[0067] The urban facility anomaly detection focused on in this application is different from existing research. It does not only focus on frequent or common urban facility anomalies, but comprehensively covers various facility anomaly situations in the urban environment. Therefore, when performing urban facility anomaly detection, an important challenge is faced - few-shot learning, which is mainly reflected in the following three aspects:
[0068] (1) Diversity and complexity: There are many types of urban facility anomalies with different manifestations. From minor garbage spills to severe road collapses, each anomaly has its unique characteristics and manifestations. In addition, the same type of anomaly may show significant differences in different environments and at different times, increasing the difficulty of anomaly detection.
[0069] (2) Data acquisition and annotation costs: Anomaly events are usually sudden and uncertain, resulting in the inability to collect large-scale and systematic data through conventional means. At the same time, the accurate annotation of anomaly images requires a large amount of manpower and time, increasing the cost and complexity of data annotation.
[0070] (3) Data imbalance problem: In practical applications, some types of anomaly events may be more common, while other types are relatively rare. The data imbalance phenomenon will cause the model to deviate during the training process, thus affecting its generalization ability and making the model perform poorly in the detection of rare anomaly types.
[0071] These factors make it necessary for the model to cope with a complex and changing data environment and overcome the challenges brought by few-shot learning when performing urban facility anomaly detection.
[0072] To solve this problem, in addition to collecting high-quality training data based on the above training sample optimization method, this application also improves the network structure of the model and proposes an anomaly detection model, which includes a backbone network, a neck network, and a detection head. Specifically, the backbone network is used to extract features from the input image to obtain image features; the neck network is used to fuse the image features to obtain fused features; and the detection head is used to detect the fused features to obtain the detection results of urban facility anomalies in the image to be processed.
[0073] Specifically, to overcome the challenges brought by few-shot learning in the process of urban facility anomaly detection, this application sets a Text-guided Visual Learning module (TgVL) in the neck network. By fully fusing the features of these two modalities of image and text, it can enhance the expressiveness of effective image features under the guidance of text. That is, the TgVL module uses text information to guide image feature learning, enabling the model to extract more expressive and richer effective features, namely the second fused features. Using the first fused features determined based on the second fused features for urban facility anomaly detection can enable the anomaly detection model to obtain stronger generalization ability and higher accuracy.
[0074] In some embodiments, the TgVL module may include a Vision Feature Enhancement (VFE) branch, a Text-guided Vision (TgV) branch, and a first fusion layer. The first fusion layer is a Text-Vision Feature Fusion (TVFF) layer. The VFE branch is used to process the first image features to enhance the feature expression ability of the first image features and obtain enhanced image features; the TgV branch guides the extraction of effective features in the first image features through text features to learn the guided learning weights; the TVFF layer performs an attention operation on the enhanced image features based on the guided learning weights, which can promote the feature fusion of the two modalities and effectively solve the problem of difficult few-shot learning.
[0075] In some embodiments, the preset text features can be extracted based on the CLIP (Contrastive Language-Image Pretraining) model, so as to obtain text representations with rich semantics and contrastive learning ability. The text data may include descriptions of urban facility anomalies, which are used to provide prior knowledge for the anomaly detection model to help it identify urban facility anomalies of unseen categories in the urban facility anomaly detection task and achieve more accurate cross-modal understanding and generalization ability.
[0076] In some embodiments, to improve the performance of the anomaly detection model, the TgV branch can be implemented through an image feature preprocessing sub-branch, a text feature preprocessing sub-branch, an Einstein summation layer, and a MaxSigmoid activation function layer. After the first image feature and text feature are preprocessed respectively, they can be effectively fused through the Einstein summation layer to combine the feature information of the two modalities; subsequently, the fused feature is non-linearly activated by means of the MaxSigmoid activation function layer, which can further emphasize the key information and generate high-precision guiding learning weights. Thus, the TgV branch can enhance the feature expression ability of the anomaly detection model by learning high-precision guiding learning weights, thereby improving its accuracy and generalization ability.
[0077] In some embodiments, to further strengthen the advantages of text-guided image learning, make full use of multi-modal information fusion to improve the anomaly detection model's understanding ability of complex scenes, and then improve the detection accuracy of the model, the detection head can include a regression branch, a classification branch, and a concatenation layer. Among them, the regression branch can enhance and refine the input first fused feature to obtain a regression result with higher accuracy. The classification branch combines the first fused feature with the text feature. During the combination process, a regularization technique is adopted to ensure the stability and generalization ability of the model, and the Einstein summation convention is introduced to efficiently perform operations between features. Then, the log_scale coefficient is applied to scale the result to achieve more refined feature fusion and obtain a classification result. Finally, the outputs of the regression branch and the classification branch are combined through a concatenation operation. At this time, the concatenated result is actually still a feature. To present a visual detection result, it can be activated through a preset activation function to form the final visual detection result. This detection result is still an image, and there are annotation boxes for urban facility anomalies and the categories of urban facility anomalies in this image.
[0078] The improved detection head can be called a Text-Vision Feature (TVF) detection head. It can combine the semantic information of the text and the spatial features of the image to perform multi-modal information fusion to improve the understanding ability of complex scenes, and then improve the detection accuracy of urban facility anomalies. At the same time, by comprehensively analyzing data from different sources, it can more accurately detect the area corresponding to the text in the image and reduce the probability of misjudgment. It can be considered that the anomaly detection model applying this TVF detection head has significant advantages in multiple dimensions such as accuracy, efficiency, flexibility, and application scope.
[0079] When performing urban facility anomaly detection, in addition to facing the problem of small-sample learning, it also faces semantic challenges in scene understanding, as follows:
[0080] (1) Ambiguity of classification boundaries: In urban infrastructure anomaly detection, the boundaries between many categories are not clear. For example, whether a slight bend in a guardrail should be regarded as damage or normal wear requires a high judgment accuracy from the anomaly detection model.
[0081] (2) Context dependence: The occurrence of abnormal events is often closely related to a specific environmental background. When detecting anomalies, the model needs to fully consider the surrounding scene information, such as weather, lighting, traffic conditions, etc., which undoubtedly increases the complexity of detection.
[0082] (3) Ambiguity of semantic understanding: For some abnormal phenomena, there may be different interpretations and understandings. This semantic ambiguity will cause the model to be confused when processing relevant data, thus affecting the accuracy of the detection results.
[0083] In order to overcome the semantic challenges of scene understanding to a certain extent, the present application also improves the backbone network and / or the TgVL module, that is, a global-local perception unit (Adaptive Global-Local Perception Unit, AGLPU) module is set in the backbone network and / or the TgVL module. The AGLPU module fuses by separately extracting the local features and global semantic information of the input second image features, so as to achieve more accurate scene understanding.
[0084] Specifically, the AGLPU module mainly performs two parts of work: local feature extraction and global semantic understanding. The local feature extraction part focuses on local regions in the image, such as specific anomalies like road cracks and manhole cover damage, and can efficiently capture detailed information; while the global semantic understanding part focuses on the context information of the entire scene, including weather, lighting changes, and surrounding environmental features, in order to more comprehensively understand the background of abnormal events.
[0085] By fusing these two types of features, the AGLPU module can not only improve the model's sensitivity to detailed anomalies, but also more accurately judge the nature of anomalies in complex backgrounds. The introduction of this module helps to enhance the semantic perception ability of the anomaly detection model, thus effectively overcoming problems such as the ambiguity of classification boundaries, context dependence, and ambiguity of semantic understanding.
[0086] This combination of global and local perception not only improves the judgment accuracy of the anomaly detection model, but also further optimizes the scene understanding ability, enabling the anomaly detection model to more accurately and stably detect various facility anomalies in complex urban environments.
[0087] Convolutional neural networks (CNNs) can effectively extract local features such as edges and textures through convolutional layers, and reduce the number of parameters through the parameter sharing mechanism, reducing the risk of overfitting. Pooling layers help reduce the computational load and memory footprint, while translational invariance to translation enhances its generalization ability. However, the limitations of CNNs are also obvious: its global information capture ability is limited, it cannot handle long-range dependencies, and it has a large demand for training data, making it prone to overfitting when data is insufficient. In addition, the deep network structure increases the computational complexity.
[0088] The advantage of Transformer lies in its self-attention mechanism, which can capture dependencies at different positions in the sequence, achieve global context awareness, and can effectively capture long-range dependencies, especially suitable for sequence data such as natural language. Transformer improves the training and inference efficiency by processing sequences in parallel. However, the disadvantages of Transformer are high computational complexity, requiring a large amount of computing resources and memory, and usually relying on a large amount of data and computing resources for pre-training and fine-tuning. The complexity of the model also leads to poor interpretability, posing challenges to tasks that require interpretability.
[0089] Therefore, when designing a structure incorporating the AGLPU module, the advantages of CNN and Transformer are combined to give full play to their complementary characteristics. CNN is good at extracting local features, while Transformer captures global information and long-range dependencies through its self-attention mechanism. Combining the two can achieve better performance in urban facility anomaly detection tasks.
[0090] In some embodiments, the AGLPU module includes a convolutional layer, a local perception unit (LPU) branch, a global perception unit (GPU) branch, and a second fusion layer. The convolutional layer is responsible for integrating the input features and achieving channel dimensionality reduction, providing a more compact feature representation for subsequent processing. The LPU branch focuses on extracting local detail features in the image, such as specific anomaly locations and morphological details, and can capture fine information. The GPU branch is responsible for extracting global semantic information, enhancing the overall understanding of facility anomalies by considering the overall environmental background (such as lighting, weather, surrounding facilities, etc.). Finally, the second fusion layer effectively fuses the local and global features through an adaptive weighting mechanism, thereby enhancing the model's detection ability in complex scenarios.
[0091] Among them, the LPU branch uses a CNN architecture to extract local detailed features, and the GPU branch extracts global semantic information through a Transformer structure. However, in traditional methods, the two are usually simply connected in series or in parallel and added directly. This feature calculation method with equal weights is not conducive to the effective fusion of features. The reason is that in the initial stage of the network, the detailed features are more prominent. At this time, higher weights should be given to the semantic information to make up for the deficiency of the detailed features. As the network deepens, the high-level semantic information gradually dominates, and the detailed information gradually weakens. At this time, the weights of the detailed features need to be increased. If this difference is ignored, it will be difficult to balance the local detailed features and the global semantic information.
[0092] Therefore, in this embodiment, adaptive weights are introduced in the fusion process. These weight coefficients can be dynamically adjusted according to the feedback during the training process, so as to achieve a more flexible fusion of local detailed features and global semantic information, and improve the quality of feature representation and the performance of subsequent tasks.
[0093] In this embodiment, the AGLPU module extracts local features and global semantic information through the LPU branch and the GPU branch respectively, and adaptively weights and fuses the two. This module can flexibly adjust the ratio and structure of the CNN and the Transformer according to specific task requirements, so as to build a more efficient and powerful image processing model. That is, the AGLPU module can effectively solve problems such as blurred classification boundaries, semantic understanding ambiguities, and high context dependencies in urban facility anomaly detection, and improve the accuracy and robustness of the model.
[0094] In some embodiments, the LPU branch is composed of two bottleneck structures with residual connections connected in series, and is used to extract local detailed features. Each bottleneck structure includes multiple convolutional layers, which are used to capture subtle changes in the image, such as local information such as edges, textures, and shapes. The introduction of the residual connection ensures the smooth transmission of the information flow, avoids the problem of gradient disappearance, and at the same time accelerates the training process of the network. In this way, the LPU branch can efficiently extract key local features from the image, providing strong support for subsequent anomaly detection.
[0095] In some embodiments, the GPU branch consists of a Local Multi-Head Self-Attention (LMHSA) structure with residual connections and a Feed-Forward Neural Network (FFN) structure with residual connections in series to form the LMHSA structure, which is used to capture the global semantic information of the input features, effectively model the relationships between different positions through the multi-head self-attention mechanism, and enhance the model's ability to perceive long-range dependencies. The FFN structure further strengthens the non-linear transformation of the global features and improves the expression ability of semantic information. The introduction of residual connections not only helps to alleviate the vanishing gradient problem in the training of deep networks but also improves the stability and convergence speed of the model. Through this series structure, the GPU branch can efficiently extract global semantic information and enhance the overall performance of the model after fusing with the local features of the LPU branch.
[0096] In some embodiments, the AGLPU module can be set in the VEE branch of the TgVL module to optimize and improve the quality of the image features extracted from the image, thereby enhancing the expression ability of the features and the detection accuracy.
[0097] In some embodiments, the anomaly detection model sets the TgVL module in the neck network. The TgVL module enhances the image features through the VFE branch, learns the weights of the image features under the guidance of the text through the TgV branch, and then combines with the TVF detection head. The multi-modal fusion of text semantic information and the spatial features of the image can improve the understanding of complex scenes and more accurately and efficiently determine the detection results. In addition, the AGLPU module set in the backbone network and the TgVL module can use the LPU branch and the GPU branch to extract local features and global features and perform adaptive weighted fusion on the two, further improving the accuracy of urban facility anomaly detection.
[0098] Exemplarily, if the YOLO series model is used as the base model and the base model is improved through the TgVL module, the TVF detection head, and the AGLPU module, the resulting model can be referred to as Figure 2 This model can be called the Open-Vocabulary Multimodal Facility Anomaly Detection Algorithm based on YOLO Network (OV-MFS-YOLO network).
[0099] In some embodiments, the above anomaly detection model can adopt a two-stage training method that combines fast convergence and fine optimization during training. The training process is set for a total of N rounds, where the first n rounds are used for fast convergence. The main purpose is to quickly approach the region of the optimal solution through a relatively large learning rate to accelerate the initial training of the model. This stage focuses on improving the overall performance of the model, enabling the model's performance on the training set to improve rapidly. The remaining rounds then enter the fine optimization stage, using a smaller learning rate to further adjust the model's parameters to achieve more precise parameter adjustment and performance improvement. In this stage, the model can better capture the subtle differences in anomaly detection, reduce the risk of overfitting, and improve the generalization ability on the validation set. Through this step-by-step optimization method, the model can balance training efficiency and high precision to ensure the best results in complex urban facility anomaly detection tasks.
[0100] In some embodiments, the loss function of the above anomaly detection model mainly consists of two parts: regression loss and classification loss. When adopting the two-stage training method, the classification loss in both stages is the Binary Cross-Entropy Loss (BCE Loss), which is used to determine the specific category in the anchor box.
[0101] Among them, the regression loss in the first stage is the Complete Intersection over Union loss (CIOU Loss) and the Distribution Focal Loss (DFL loss), which are used to calculate the error between the predicted bounding box and the ground truth box. It is used for preliminary prediction, and the Adaptive Training Sample Selection (ATSS) static matching strategy is adopted for matching between positive and negative samples.
[0102] The regression loss in the second stage is the Skew Intersection over Union Loss (SIoU Loss) and the DFL loss, which are used to finely optimize the prediction box, and the DFL dynamic matching strategy is adopted for matching between positive and negative samples.
[0103] Specifically, the formula for BCE Loss is as follows:
[0104]
[0105] Among them, L BCE is the BCE Loss, nThe number of image samples with abnormal urban facilities is y i is the i true category of the ith image sample,
[0106] CIOU Loss is an improved version of common loss functions such as L1, L2, Intersection over Union (IoU), and GIoU. It adds a penalty term for the aspect ratio, which can better distinguish errors in different situations when the centers of the predicted bounding box and the ground truth bounding box coincide, and has scale invariance. The formula is as follows:
[0107]
[0108] Among them, L CIoU is the CIOU Loss, b and b gt are the centers of the predicted bounding box and the ground truth bounding box respectively; ρ is the Euclidean distance between the predicted bounding box and the ground truth bounding box; c is the distance of the diagonal of the closed area between the predicted bounding box and the ground truth bounding box; v is the consistency of the relative ratio between the predicted bounding box and the ground truth bounding box, IoU is the Intersection over Union of the predicted bounding box and the ground truth bounding box; α is the weight coefficient; w and w gt are the widths of the predicted bounding box and the ground truth bounding box respectively; h and h gt are the heights of the predicted bounding box and the ground truth bounding box respectively.
[0109] However, L CIoU it does not take into account the direction between the ground truth bounding box and the predicted bounding box, which often causes the network to "wander around" during training, with a slow convergence speed and low efficiency, and easily reduces the robustness of the model.
[0110] To solve this problem, the SIoU Loss adopted in the second stage takes into account the vector angle between the ground truth bounding box and the predicted bounding box and redefines the penalty metrics. It specifically includes four parts: Angle cost, Distance cost, Shape cost, and IoU cost.
[0111] 1. Angle cost
[0112] Refer to Figure 3, the formula for the angular loss Λ can be written as follows:
[0113]
[0114] Where, is the center point coordinate of the predicted bounding box, is the center point coordinate of the ground truth bounding box, σ is the center point distance between the ground truth bounding box and the predicted bounding box, d h is the center point height difference between the ground truth bounding box and the predicted bounding box, d w is the center point height difference between the ground truth bounding box and the predicted bounding box.
[0115] 2. Distance loss
[0116] Refer to Figure 4 , the formula for the distance loss Δ can be written as follows:
[0117]
[0118] c h and c w are the width and height of the minimum bounding rectangles of the ground truth bounding box and the predicted bounding box respectively.
[0119] 3. Shape loss
[0120] The formula for the shape loss Ω can be written as follows:
[0121]
[0122] w , h , w gt and h gt are the width and height of the predicted bounding box and the ground truth bounding box respectively, and θ is a hyperparameter that can control the degree of attention of the shape loss.
[0123] L SIoU Combining the losses in the above three aspects can improve the comprehensiveness of each part of the loss, making the convergence speed faster and the accuracy higher, which is better than L CioU . L SIoU can be written as follows:
[0124]
[0125] The formula for the DFL loss is as follows:
[0126]
[0127] Among them, S i is the cross - entropy loss between the true left border and the predicted border, S i+1 is the cross - entropy loss between the true right border and the predicted border.
[0128] The formula for the total training loss function in the first stage of the anomaly detection model is as follows:
[0129]
[0130] Among them, λ1 and λ2 are the classification and regression balance coefficients in the first stage.
[0131] The formula for the total training loss function in the second stage of the anomaly detection model is as follows:
[0132]
[0133] λ3 and λ4 are the classification and regression balance coefficients in the second stage.
[0134] Based on the two total loss functions of the anomaly detection model, two - stage training is performed on the anomaly detection model described in any of the foregoing embodiments, which can improve the model convergence speed and thus improve the efficiency of model training. By continuously iterating the anomaly detection model, the detection effect of the anomaly detection model on the anomalies of various urban facilities is better and tends to be stable. Finally, a robust version can be selected from multiple versions converged in the second stage as the trained anomaly detection model.
[0135] In some embodiments, in order to comprehensively and accurately measure the performance of each version of the anomaly detection model, it can be evaluated through a validation set. Specifically, the performance of the anomaly detection model on the validation set can be evaluated according to preset conditions. Exemplarily, these preset conditions may include performance indicators such as the intersection - over - union (IoU), detection accuracy, and recall rate of the anomaly detection model for urban facility anomalies.
[0136] IoU is an index for evaluating the overlapping degree between the predicted border and the true border, which calculates the ratio of the intersection to the union of the predicted border and the true border, and it plays a key role in determining whether it is a correct detection. The calculation formula of IoU can be written as:
[0137]
[0138] Precision (p), also known as the precision rate, refers to the proportion of correctly predicted positives among all predicted positives, as shown in the formula:
[0139]
[0140] The recall rate (Recall, R), also known as the completeness rate, refers to the proportion of correctly predicted positives among all actual positives, as shown in the formula:
[0141]
[0142] The average precision (Average Precision, AP) is calculated from the precision rate and the recall rate. A line graph of the precision rate is plotted based on the recall rate values, and the area under this line graph is calculated, as shown in the formula:
[0143]
[0144] The mean average precision (mean Average Precision, mAP) refers to the average value of the average precision AP for C different urban facility anomaly categories, as shown in the formula:
[0145]
[0146] That is to say, after verifying each version of the anomaly detection model through the validation set, the anomaly detection models of each version can be comprehensively evaluated based on the above several indicators, so as to determine the anomaly detection model with the best performance from each version as the trained anomaly detection model.
[0147] Based on the anomaly detection model in the foregoing embodiments, the present application proposes an urban facility anomaly detection method based on a multimodal model.
[0148] The urban facility anomaly detection method based on a multimodal model provided in the embodiments of the present application can be applied to electronic devices such as mobile phones, tablet computers, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). A trained facility detection module is configured on the electronic device, and the specific type of the electronic device is not limited in the embodiments of the present application.
[0149] To illustrate the technical solutions proposed in the present application, the following will describe each embodiment with an electronic device as the execution subject.
[0150] Figure 5 The schematic flowchart of the urban facility anomaly detection method based on a multimodal model provided by the present application is shown. The detection method includes:
[0151] Step 510: The electronic device extracts features from the image to be detected based on the backbone network of the pre-trained anomaly detection model, obtaining image features.
[0152] The image to be detected includes urban facilities.
[0153] Step 520: The electronic device fuses the image features based on the neck network of the anomaly detection model, obtaining the first fused feature.
[0154] Step 530: The electronic device detects the first fused feature based on the detection head of the anomaly detection model, obtaining the detection result of the anomaly of urban facilities in the image to be detected.
[0155] In this embodiment, due to the introduced TgVL module in the improved neck network, the electronic device can use the TgVL module to fully fuse the features of the two modalities of image and text, enabling the effective image features to be enhanced in expression under the guidance of text. That is, the TgVL module can make the model extract more expressive and richer effective features, namely the second fused feature, by using text information to guide image feature learning. The first fused feature determined based on the second fused feature is used for urban facility anomaly detection, which can make the anomaly detection model obtain stronger generalization ability and higher accuracy, effectively solving the problem of few-shot learning in the process of urban facility anomaly detection.
[0156] In some embodiments, for the TgVL module including a VFE branch and a TgV branch, the electronic device can perform the following operations:
[0157] Step A1: The electronic device processes the first image feature through the VFE branch to obtain an enhanced image feature.
[0158] Step A2: The electronic device sequentially performs a summation operation and a weighted activation operation on the preprocessed first image feature and text feature through the TgV branch to obtain a guiding learning weight.
[0159] Step A3: The electronic device multiplies the enhanced image feature and the guiding learning weight through the TVFF layer and performs a convolution operation to obtain the second fused feature corresponding to the first image feature and text feature.
[0160] In this embodiment, the electronic device processes the first image feature through the VFE branch to enhance its feature expression ability and generate an enhanced image feature; uses the TgV branch to guide the extraction of effective information in the first image feature with text features and learn the guiding weight; and performs an attention operation on the enhanced image feature based on the guiding learning weight through the TVFF layer to promote the fusion of image and text modality features, enabling the introduction of the TgVL module to effectively alleviate the challenges brought by few-shot learning.
[0161] In some embodiments, when the TgV branch can be implemented by an image feature preprocessing sub-branch, a text feature preprocessing sub-branch, an Einstein summation layer, and a MaxSigmoid activation function layer, the foregoing step A2 specifically includes:
[0162] Step A21: The electronic device sequentially performs a convolution operation and a reshaping operation on the first image feature through the image feature preprocessing sub-branch to obtain the preprocessed first image feature.
[0163] Step A22: The electronic device sequentially performs a linear operation and a reshaping operation on the text feature through the text feature preprocessing sub-branch to obtain the preprocessed text feature.
[0164] Step A23: The electronic device performs an Einstein summation operation on the preprocessed first image feature and the preprocessed first image feature through the Einstein summation layer to obtain a third fused feature.
[0165] Step A24: The electronic device performs a MaxSigmoid function activation operation on the third fused feature through the MaxSigmoid activation function layer to obtain a guiding learning weight.
[0166] The electronic device first uses a preset convolutional layer and a linear layer in the TgV branch to respectively perform a preliminary adjustment of the dimensions of the first image feature and the text feature. Then, the shapes of the adjusted first image feature and text feature are changed through a reshaping operation. Subsequently, an Einstein summation operation is implemented to integrate these features. Finally, this branch outputs MaxSigmoid weights, which are multiplied by the output of the VFE branch to play the role of an attention mechanism, emphasizing key information and improving the performance of the model.
[0167] Among them, the convolution operation can be implemented by a convolutional layer with a specified convolution kernel, such as a 3×3 convolutional layer. The reshaping operation refers to the Reshape operation, which can adjust the dimensions of the data by changing the shape of the tensor to adapt to the input requirements of different network layers. Through the Reshape operation, the model can flexibly convert the data format, thereby improving the adaptability and flexibility of the network structure.
[0168] Performing a linear operation on the text feature means mapping each word or character to a high-dimensional vector space through an embedding layer to obtain the representation of the text unit. Then, these embedding vectors are processed using a linear transformation (such as matrix multiplication), usually using a weight matrix to adjust the feature dimensions, thereby achieving feature compression or expansion. Finally, an activation function (such as ReLU, Sigmoid, etc.) is often applied to non-linearly activate the result of the linear transformation to enhance the expressive ability of the model. This linear operation helps to effectively learn text features during the training process and capture the potential semantic relationships of the text.
[0169] In some embodiments, the VFE branch mainly consists of two 1x1 convolutional layers and an AGLFU module. The first 1x1 convolution is used to reduce the channel dimension, the second 1x1 convolution is used to increase the channel dimension, and the goal of the AGLFU module is to optimize and improve the quality of the image features extracted from the image, thereby enhancing the feature expression ability and detection accuracy.
[0170] In some embodiments, the TVFF layer contains two 3x3 convolutional layers, and its main function is to perform deeper fusion and quality improvement on the input features. Through the processing of these two convolutional layers, the module can optimize the feature representation and enhance its discriminative power and effectiveness.
[0171] Exemplarily, Figure 6 The specific structural schematic diagram of the TgVL module (including the VFE branch, the TgV branch, and the TVFF layer) is shown. Based on Figure 6 the structure of the TgVL module in assuming the first image feature of the input is and the text feature X V and X T The processing of
[0172]
[0173] where, represents a convolution with a kernel size of 1x1; represents a convolution with a kernel size of 3x3; the Reshape operation represents the corresponding reshaping operation; Y V represents the enhanced high-quality image features; Linear represents the linear layer; einsum represents the Einstein summation operation; W TgV represents the weight for text-guided image learning; O represents the output of the entire TgVL module.
[0174] In some embodiments, the TVF detection head includes a regression branch, a classification branch, and a splicing layer. Step 530 specifically includes:
[0175] Step B1, the electronic device performs a depth convolution operation on the first fusion feature through the regression branch to obtain a regression result.
[0176] Step B2: The electronic device sequentially performs a summation operation and a scaling operation on the classified first fusion feature and the preprocessed text feature through the classification branch to obtain a classification result;
[0177] Step B3: The electronic device splices the regression result and the classification result through a splicing layer and activates them based on a preset activation function to obtain the visualized detection result.
[0178] Exemplarily, Figure 7 The network structure diagram of the TVF detection head (including a regression branch, a classification branch, and a splicing layer) is shown.
[0179] Among them, the electronic device enhances and refines the input first fusion feature through two consecutive 3x3 depthwise convolutions (DWConv) of the regression branch. The depthwise convolution operation can effectively extract local features while reducing the computational amount and the number of parameters by independently processing each channel in each convolution layer. This process helps to refine the image features, enhance the model's sensitivity to details, and provide a higher-quality feature representation for subsequent anomaly detection.
[0180] The electronic device also combines the processed first fusion feature with the input text feature through the classification branch. In this process, first, a depthwise convolution is applied to the first fusion feature for further processing, and then the stability is improved through regularization techniques (such as Normalize) to ensure the generalization ability of the model. For the text feature, the regularization technique is directly used to ensure its consistency and reliability. To efficiently perform the operations between features, the Einstein summation convention is introduced to optimize the calculation process. Finally, the result is scaled with the log_scale coefficient to achieve a more refined feature fusion, thereby enhancing the model's ability to process multimodal data and the accuracy of anomaly detection.
[0181] For the features obtained from the two branches, the electronic device first splices them and activates the splicing result through a preset activation function to form the final visualized detection result. The detection result is still an image, and there are annotation boxes for urban facility anomalies and the categories of urban facility anomalies in the image.
[0182] In some embodiments, an AGLPU module is provided in the backbone network and / or the TgVL module. The AGLPU module is used to extract and fuse the local features and global features of the input second image feature to obtain a fourth fusion feature.
[0183] In some embodiments, in order to fully and comprehensively extract the local and global features of the second image feature, the AGLPU module may include a convolutional layer, an LPU branch, a GPU branch, and a second fusion layer. Among them, the electronic device extracts local detailed features through the LPU branch based on the CNN architecture and extracts global semantic information through the GPU branch based on the Transformer structure. The specific extraction steps are as follows:
[0184] Step C1: The electronic device performs a convolution operation on the second image feature through the convolutional layer to obtain a convolutional feature.
[0185] Step C2: The electronic device extracts features from the convolutional feature through the local perception unit branch to obtain local features.
[0186] Step C3: The electronic device extracts features from the convolutional feature through the global perception unit to obtain global features.
[0187] Step C4: The electronic device adaptively weights and fuses the local features and global features through the second fusion layer to obtain a fourth fusion feature.
[0188] The convolutional layer is mainly used for feature integration and channel dimensionality reduction, such as a 1×1 convolutional layer. The convolutional features output by the convolutional layer can be respectively input into the LPU branch and the GPU branch to extract local detailed features and global semantic information. In order to fully balance local features and global features, the electronic device introduces adaptive weights in the fusion process through the second fusion layer, and these weight coefficients can be dynamically adjusted according to the feedback during the training process, so as to realize the flexible fusion of local detailed features and global semantic information and improve the quality of feature representation and the performance of subsequent tasks.
[0189] In this embodiment, based on the LPU branch and the GPU branch, the AGLPU module effectively solves problems such as blurred classification boundaries, semantic understanding ambiguities, and high context dependencies in urban facility anomaly detection, and improves the accuracy and robustness of the model.
[0190] In some embodiments, the LPU branch is composed of two bottleneck structures with residual connections in series; the aforementioned step C2 specifically includes:
[0191] Step C21: The electronic device sequentially performs a convolution operation, a depth convolution operation, and a convolution operation on the convolutional feature to obtain a local first sub-feature.
[0192] Step C22: The electronic device concatenates the convolutional feature and the local first sub-feature to obtain a first concatenated feature.
[0193] Step C23: The electronic device sequentially performs a convolution operation, a depth convolution operation, and a convolution operation on the first concatenated feature to obtain a local second sub-feature.
[0194] Step C24: The electronic device splices the first splicing feature and the local second sub-feature to obtain a local feature.
[0195] The bottleneck structure with residual connection consists of a convolutional layer, a depth convolutional layer, and a convolutional layer connected in series. The electronic device compresses the channels and integrates features through the first convolutional layer, extracts deep local features through the depth convolutional layer, and restores the channels through the last convolutional layer. The electronic device also alleviates the problem of gradient disappearance in the training of deep networks through the introduced residual connection, thereby improving the stability and convergence speed of the model. By connecting two bottleneck structures with residual connections in series, the electronic device can fully extract the local detail information in the convolutional features, thereby enhancing the model's ability to express local features.
[0196] In some embodiments, the GPU branch consists of a LMHSA structure with residual connection and a FFN structure with residual connection connected in series; the aforementioned step C3 specifically includes:
[0197] Step C31: The electronic device sequentially performs layer normalization operation, multi-head self-attention operation, and regularization operation on the convolutional features to obtain a global first sub-feature.
[0198] Step C32: The electronic device splices the convolutional features and the global first sub-feature to obtain a second splicing feature.
[0199] Step C33: The electronic device performs layer normalization operation, feature transformation operation, and regularization operation on the second splicing feature to obtain a global second sub-feature.
[0200] Step C34: The electronic device splices the second splicing feature and the global second sub-feature to obtain a global feature.
[0201] The LMHSA structure with residual connection includes a normalization layer, a Lite Multi-Head Self-Attention (LMHSA) module, a regularization layer, and a residual connection. The FFN structure with residual connection includes a normalization layer, a Feed-Forward Network (FFN), a regularization layer, and a residual connection; the two connected structures are used to extract global semantic information. Exemplarily, the regularization layer can be Dropout / DropPath.
[0202] The LMHSA module first performs a linear transformation on the convolutional features to generate queries (Q), keys (K), and values (V), and then performs self-attention operations. To reduce the computational complexity, we introduce k×k depth convolution to reduce the spatial size and add relative position bias in the self-attention calculation BTo enhance the position perception ability. The FFN consists of two fully connected (FC) layers + a GELU (Gaussian Error Linear Unit) activation function and a depth convolution layer. Among them, the first FC layer is used to reduce the feature dimension, the depth convolution layer is used to enhance the local feature modeling ability, and the second FC layer is used to expand the feature dimension, thereby improving the expression ability of the model.
[0203] Exemplarily, Figure 8 The schematic diagram of the network structure (including the convolution layer, LPU branch, GPU branch and the second fusion layer) is shown. Based on Figure 8 the structure of the AGLPU module in, assume that the input second image feature is , where H, W, and C respectively represent the height, width, and channels of the second image feature; specifically, the processing of the input X by the AGLPU module can be expressed as the following content:
[0204]
[0205]
[0206] Among them, represents the convolution feature, represents a convolution with a kernel size of 1×1 and an output channel of C / 2, which is used for channel compression and feature integration; represents a depth convolution with a kernel size of 3×3 and an output channel of C / 4; Y C1 represents the output of the first bottleneck structure with a residual connection in the CNN branch, i.e., LPU; similarly, Y C2 represents the output of the second bottleneck structure with a residual connection in the CNN branch, i.e., LPU; LN represents a normalization (LayerNorm) layer; LMHSA represents a lightweight multi-head self-attention module; FFN represents a feed-forward neural network, which is used to enhance the non-linearity ability of the network; Y T1 represents the output of the first LMHSA structure with a residual connection in the Transformer branch, i.e., GPU; similarly, Y T2 represents the output of the FFN structure with a residual connection in the Transformer branch, i.e., GPU; YRepresents the output of the entire AGLPU module, which is formed by weighted combination of the LPU branch and the GPU branch, where the weight α is a network hyperparameter learned through network training.
[0207] Combined with the detection methods shown in the above embodiments, the anomaly detection model of the present application sets the TgVL module in the neck network. The TgVL module enhances image features through the VFE branch, and learns the weights of image features under text guidance through the TgV branch. By applying attention to the enhanced image features with these weights, it can promote the feature fusion of the two modalities and effectively solve the problem of difficult small-sample learning. Combined with the TVF detection head, the multi-modal fusion of text semantic information and spatial features of images can improve the understanding of complex scenes and more accurately and efficiently determine the detection results. In addition, the AGLPU module set in the backbone network and / or the TgVL module can use the LPU branch and the GPU branch to extract local features and global features, and perform adaptive weighted fusion on the two, further improving the accuracy of urban facility anomaly detection.
[0208] In some embodiments, the urban facility anomaly detection method based on the OV-MFS-YOLO network integrates visual and text information, analyzes urban image data captured by large-scale vehicle-mounted cameras through deep learning, and automatically mines complex environmental features and detects anomalies. This method can accurately detect and classify various urban facility anomaly events without relying on preset rules, demonstrating a high degree of automation and intelligence. Its specific advantages are summarized as follows:
[0209] (1) Cross-modal fusion and generalization ability: The OV-MFS-YOLO network breaks through the limitations of traditional object detection methods, jointly trains text and image modalities, can not only accurately detect known categories, but also effectively detect and classify new and undefined anomaly categories. Its powerful generalization ability enables it to complete small-sample learning or even zero-sample detection tasks even when the sample size is limited.
[0210] (2) Incremental learning and knowledge retention: This algorithm supports continuous incremental learning, can effectively alleviate the problem of "catastrophic forgetting", ensures that while absorbing new data, it can still retain and optimize existing knowledge, and continuously improve the model performance.
[0211] (3) Efficient inference and flexible deployment: The OV-MFS-YOLO network supports online prompt vocabulary learning and offline vocabulary prediction, significantly improving the inference speed of the network while reducing the computational resource requirements, and is suitable for efficient applications in private deployment and actual scenarios.
[0212] Currently, the OV-MFS-YOLO network has broad application prospects in the fields of smart cities and smart urban management, and can be used for tasks such as urban facility anomaly detection, violation monitoring, and environmental governance. Its cross-modal fusion ability and incremental learning characteristics enable it to adapt to the dynamically changing urban environment and achieve accurate and real-time anomaly detection. This network can also be extended to task scenarios such as autonomous driving, intelligent security, and road inspection that require continuous incremental learning and few-shot learning.
[0213] Exemplarily, in autonomous driving, it can be used to detect road anomalies, traffic sign changes, and emergencies, improving driving safety.
[0214] Exemplarily, in intelligent security, it can monitor abnormal events such as illegal intrusion and facility damage, enhancing the intelligence level of the security system.
[0215] Exemplarily, in road inspection, it can assist traffic management departments in detecting road surface damage, obstacles, and traffic facility anomalies, improving the inspection efficiency and accuracy.
[0216] In summary, with the continuous optimization of this network, its application scope will be further expanded, promoting the in-depth application of multi-modal perception technology in the fields of smart city management and intelligent transportation.
[0217] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0218] Corresponding to the urban facility anomaly detection method based on a multi-modal model in the above embodiments, Figure 9 FIG. shows a structural block diagram of an urban facility anomaly detection device 9 provided by an embodiment of the present application. For the sake of illustration, only the parts related to the embodiments of the present application are shown.
[0219] Referring to Figure 9 , the urban facility anomaly detection device 9 based on a multi-modal model includes:
[0220] A feature extraction module 91, configured to extract features from an image to be detected based on the backbone network of a pre-trained anomaly detection model, obtaining image features; the image to be detected includes urban facilities;
[0221] A feature fusion module 92, configured to fuse the image features based on the neck network of the anomaly detection model, obtaining a first fused feature;
[0222] A detection module 93, configured to detect the first fused feature based on the detection network of the anomaly detection model, obtaining a detection result of anomalies of urban facilities in the image to be detected;
[0223] Among them, the neck network is provided with a TgVL module, and the TgVL module is used to fuse the input first image feature and the text feature of the preset guiding text to obtain a second fusion feature, and the first fusion feature is obtained based on the second fusion feature.
[0224] Optionally, the TgVL module includes an image feature enhancement branch, a text-guided image branch, and a first fusion layer; for the first image feature and the text feature:
[0225] Process the first image feature through the image feature enhancement branch to obtain an enhanced image feature;
[0226] Perform a summation operation and a weighted activation operation on the preprocessed first image feature and the text feature in sequence through the text-guided image branch to obtain a guided learning weight;
[0227] Multiply the enhanced image feature and the guided learning weight through the first fusion layer and perform a convolution operation to obtain a second fusion feature corresponding to the first image feature and the text feature.
[0228] Optionally, the text-guided image branch includes an image feature preprocessing sub-branch, a text feature preprocessing sub-branch, an Einstein summation layer, and a MaxSigmoid activation function layer; performing an Einstein summation operation and a weighted activation operation on the preprocessed first image feature and the text feature in sequence through the text-guided image branch to obtain a guided learning weight includes:
[0229] Perform a convolution operation and a reshaping operation on the first image feature in sequence through the image feature preprocessing sub-branch to obtain the preprocessed first image feature;
[0230] Perform a linear operation and a reshaping operation on the text feature in sequence through the text feature preprocessing sub-branch to obtain the preprocessed text feature;
[0231] Perform an Einstein summation operation on the preprocessed first image feature and the preprocessed first image feature through the Einstein summation layer to obtain a third fusion feature;
[0232] Perform a MaxSigmoid function activation operation on the third fusion feature through the MaxSigmoid activation function layer to obtain a guided learning weight.
[0233] Optionally, the detection network includes a regression branch, a classification branch, and a splicing layer, and the detection module 93 is used for:
[0234] Perform a depth convolution operation on the first fusion feature through the regression branch to obtain a regression result;
[0235] Performing a summation operation and a scaling operation on the classified first fusion feature and the preprocessed text feature in sequence through a classification branch to obtain a classification result;
[0236] Concatenating the regression result and the classification result through a concatenation layer and activating based on a preset activation function to obtain the visualized detection result.
[0237] Optionally, an AGLPU module is set in the backbone network and / or the TgVL module, and the AGLPU module is used to extract and fuse the local features and the global features of the input second image feature to obtain a fourth fusion feature.
[0238] Optionally, the AGLPU module includes a convolutional layer, a local perception unit branch, a global perception unit branch, and a second fusion layer. For each second image feature:
[0239] Performing a convolution operation on the second image feature through the convolutional layer to obtain a convolution feature;
[0240] Performing feature extraction on the convolution feature through the local perception unit branch to obtain local features;
[0241] Performing feature extraction on the convolution feature through the global perception unit to obtain global features;
[0242] Performing adaptive weighted fusion on the local features and the global features through the second fusion layer to obtain a fourth fusion feature.
[0243] Optionally, the local perception unit branch is composed of two bottleneck structures with residual connections in series; performing feature extraction on the convolution feature through the local perception unit branch to obtain local features;
[0244] Performing a convolution operation, a depth convolution operation, and a convolution operation on the convolution feature in sequence to obtain a local first sub-feature;
[0245] Concatenating the convolution feature and the local first sub-feature to obtain a first concatenated feature;
[0246] Performing a convolution operation, a depth convolution operation, and a convolution operation on the first concatenated feature in sequence to obtain a local second sub-feature;
[0247] Concatenating the first concatenated feature and the local second sub-feature to obtain local features.
[0248] Optionally, the global perception unit branch is composed of an LMHSA structure with residual connections and an FFN structure with residual connections in series; performing feature extraction on the convolution feature through the global perception unit to obtain global features, including:
[0249] Perform layer normalization operations, multi-head self-attention operations, and regularization operations on the convolutional features in sequence to obtain the first global sub-feature;
[0250] Concatenate the convolutional features and the first global sub-feature to obtain the second concatenated feature;
[0251] Perform layer normalization operations, feature transformation operations, and regularization operations on the second concatenated feature to obtain the second global sub-feature;
[0252] Concatenate the second concatenated feature and the second global sub-feature to obtain the global feature.
[0253] It should be noted that the content such as information interaction and execution process between the above-mentioned devices / units, due to being based on the same concept as the method embodiment of the present application, for its specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details will not be elaborated here.
[0254] Figure 10 This is a schematic structural diagram of the physical layer of an electronic device provided by an embodiment of the present application. As Figure 10 shown, the electronic device 10 in this embodiment includes: at least one processor 100 ( Figure 10 only one processor is shown in the figure), a memory 101, and a computer program 102 stored in the memory 101 and executable on at least one processor 100. When the processor 100 executes the computer program 102, it implements the steps in any of the above method embodiments of the urban facility anomaly detection method based on a multi-modal model, such as Figure 5 the steps 510-530 shown.
[0255] The processor 100 may be a central processing unit (CPU), and the processor 100 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0256] The memory 101 may be an internal storage unit of the electronic device 10 in some embodiments, such as the hard disk or memory of the electronic device 10. The memory 101 may also be an external storage device of the electronic device 10 in other embodiments, such as a plug-in hard disk equipped on the electronic device 10, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.
[0257] Furthermore, the memory 101 may also include both the internal storage unit of the electronic device 10 and the external storage device. The memory 101 is used to store an operating device, application programs, a BootLoader, data, and other programs, such as the program code of a computer program. The memory 101 may also be used to temporarily store data that has been output or will be output.
[0258] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the above device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments and will not be elaborated here.
[0259] The embodiment of this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the foregoing method embodiments can be implemented.
[0260] The embodiment of this application provides a computer program product. When the computer program product runs on an electronic device, the electronic device can execute the steps in the foregoing method embodiments.
[0261] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of this application, a computer program can be used to instruct the relevant hardware to complete. The above computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the above computer program includes computer program code, and the above computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The above computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / electronic device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc.
[0262] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0263] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0264] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the above-mentioned module or unit division is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0265] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0266] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An abnormal detection method for urban facilities based on a multimodal model, characterized in that Including: Performing feature extraction on the image to be detected based on the backbone network of the anomaly detection model that has been pre-trained to obtain image features; The image to be detected includes urban facilities; Fusing the image features based on the neck network of the anomaly detection model to obtain a first fused feature; Detecting the first fused feature based on the detection head of the anomaly detection model to obtain the detection result of the anomaly of urban facilities in the image to be detected; Wherein, the neck network is provided with a TgVL module, and the TgVL module is used to fuse the input first image feature and the text feature of the preset guiding text to obtain a second fused feature, and the first fused feature is obtained based on the second fused feature; The TgVL module includes an image feature enhancement branch, a text-guided image branch and a first fusion layer; for the first image feature and the text feature: Processing the first image feature through the image feature enhancement branch to obtain an enhanced image feature; Performing a summation operation and a weighted activation operation on the pre-processed first image feature and the pre-processed text feature in sequence through the text-guided image branch to obtain a guiding learning weight; Multiplying the enhanced image feature and the guiding learning weight through the first fusion layer and performing a convolution operation to obtain the second fused feature corresponding to the first image feature and the text feature; The text-guided image branch includes an image feature pre-processing sub-branch, a text feature pre-processing sub-branch, an Einstein summation layer and a MaxSigmoid activation function layer; the summation operation includes an Einstein summation operation, and performing the summation operation and the weighted activation operation on the pre-processed first image feature and the pre-processed text feature in sequence through the text-guided image branch to obtain a guiding learning weight, including: Performing a convolution operation and a reshaping operation on the first image feature in sequence through the image feature pre-processing sub-branch to obtain the pre-processed first image feature; Performing a linear operation and a reshaping operation on the text feature in sequence through the text feature pre-processing sub-branch to obtain the pre-processed text feature; Performing an Einstein summation operation on the pre-processed first image feature and the pre-processed first image feature through the Einstein summation layer to obtain a third fused feature; Performing a MaxSigmoid function activation operation on the third fused feature through the MaxSigmoid activation function layer to obtain the guiding learning weight.
2. The detection method according to claim 1, wherein The detection head includes a regression branch, a classification branch and a splicing layer, and detecting the first fused feature based on the detection head of the anomaly detection model to obtain the detection result of the anomaly of urban facilities in the image to be detected, including: Performing a depth convolution operation on the first fused feature through the regression branch to obtain a regression result; Performing a summation operation and a scaling operation on the classified first fused feature and the pre-processed text feature in sequence through the classification branch to obtain a classification result; The splicing layer splices the regression result and the classification result and activates them based on a preset activation function to obtain the visualized detection result.
3. The detection method according to claim 1 or 2, characterized in that, An AGLPU module is provided in the backbone network and / or the TgVL module, and the AGLPU module is used to extract and fuse the local features and global features of the input second image features to obtain a fourth fused feature.
4. The detection method according to claim 3, characterized in that, The AGLPU module includes a convolutional layer, a local perception unit branch, a global perception unit branch, and a second fusion layer. For each of the second image features: The convolutional layer performs a convolution operation on the second image feature to obtain a convolutional feature; The local perception unit branch extracts features from the convolutional feature to obtain the local feature; The global perception unit extracts features from the convolutional feature to obtain the global feature; The second fusion layer adaptively weights and fuses the local feature and the global feature to obtain the fourth fused feature.
5. The detection method according to claim 4, characterized in that The local perception unit branch is composed of two bottleneck structures with residual connections in series; the step of extracting features from the convolutional feature through the local perception unit branch to obtain the local feature includes: Performing a convolution operation, a depth convolution operation, and a convolution operation on the convolutional feature in sequence to obtain a local first sub-feature; Splicing the convolutional feature and the local first sub-feature to obtain a first spliced feature; Performing a convolution operation, a depth convolution operation, and a convolution operation on the first spliced feature in sequence to obtain a local second sub-feature; Splicing the first spliced feature and the local second sub-feature to obtain the local feature.
6. The detection method according to claim 4, characterized in that The global perception unit branch is composed of an LMHSA structure with residual connections and an FFN structure with residual connections in series; the step of extracting features from the convolutional feature through the global perception unit to obtain the global feature includes: Performing a layer normalization operation, a multi-head self-attention operation, and a regularization operation on the convolutional feature in sequence to obtain a global first sub-feature; Splicing the convolutional feature and the global first sub-feature to obtain a second spliced feature; Performing a layer normalization operation, a feature transformation operation, and a regularization operation on the second spliced feature to obtain a global second sub-feature; Splicing the second spliced feature and the global second sub-feature to obtain the global feature.
7. An urban facility anomaly detection device based on a multi-modal model, characterized in that, Including: A feature extraction module for extracting features from an image to be detected based on the backbone network of a pre-trained anomaly detection model to obtain image features; the image to be detected includes urban facilities; A feature fusion module for fusing the image features based on the neck network of the anomaly detection model to obtain a first fused feature; A detection module for detecting the first fused feature based on the detection head of the anomaly detection model to obtain a detection result of anomalies in urban facilities in the image to be detected; Among them, the neck network is provided with a TgVL module, and the TgVL module is used to fuse the input first image feature and the text feature of a preset guiding text to obtain a second fusion feature, and the first fusion feature is obtained based on the second fusion feature; The TgVL module includes an image feature enhancement branch, a text-guided image branch, and a first fusion layer; for the first image feature and the text feature: The first image feature is processed by the image feature enhancement branch to obtain an enhanced image feature; The sum operation and the weighted activation operation are sequentially performed on the preprocessed first image feature and the preprocessed text feature through the text-guided image branch to obtain a guided learning weight; The enhanced image feature and the guided learning weight are multiplied by the first fusion layer and a convolution operation is performed to obtain the second fusion feature corresponding to the first image feature and the text feature; The text-guided image branch includes an image feature preprocessing sub-branch, a text feature preprocessing sub-branch, an Einstein summation layer, and a MaxSigmoid activation function layer; the sum operation includes an Einstein summation operation, and the sum operation and the weighted activation operation are sequentially performed on the preprocessed first image feature and the preprocessed text feature through the text-guided image branch to obtain a guided learning weight, including: The first image feature is sequentially subjected to a convolution operation and a reshaping operation through the image feature preprocessing sub-branch to obtain the preprocessed first image feature; The text feature is sequentially subjected to a linear operation and a reshaping operation through the text feature preprocessing sub-branch to obtain the preprocessed text feature; An Einstein summation operation is performed on the preprocessed first image feature and the preprocessed first image feature through the Einstein summation layer to obtain a third fusion feature; The MaxSigmoid function activation operation is performed on the third fusion feature through the MaxSigmoid activation function layer to obtain the guided learning weight.
8. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for detecting anomalies in urban facilities based on a multimodal model according to any one of claims 1 to 6.
9. A computer program product, the computer program product storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for detecting anomalies in urban facilities based on a multimodal model according to any one of claims 1 to 6.
Citation Information
Patent Citations
YOLO-based abnormal city facility detection method and apparatus, and electronic device
CN118865002A