A repeated urban management case identification method based on multi-modal data

CN121302270BActive Publication Date: 2026-09-25XIAN DINGCHENG CHENGWEI BIG DATA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511512775.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-09-25
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

在案件上报过程中,由于不同上报人员可能针对同一案件分别发现并拍摄上传,导致大量重复案件的产生,不仅增加了数据存储和管理的负担,还可能影响案件的后续委派和处理进度,降低管理系统的运行效率

Benefits of technology

[0015]相对于现有技术,本申请具有以下有益效果:本申请融合案件图像、文本描述与地理位置信息多模态数据,构建深度学习模型,实现对重复城管案件的高效识别;该方法有效克服了传统人工去重方式存在的准确率低、人工成本高等问题,不仅显著提升了重复案件识别的准确性和处理效率,还增强了系统对复杂场景的适应能力,具有较强的实用性和可扩展性,为智慧城管案件管理提供了一种智能化、自动化的解决方案。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121302270B_ABST
    Figure CN121302270B_ABST
Patent Text Reader

Abstract

The application relates to a repeated urban management case recognition method based on multi-modal data, and belongs to the field of pattern recognition. The method comprises the following steps: constructing a high-frequency repeated case library; screening cases matching the type of a to-be-recognized case from the high-frequency repeated case library as to-be-compared cases; extracting feature information of the to-be-recognized case and the to-be-compared cases by using a feature extraction network, wherein the feature information comprises a case occurrence image feature, a text description feature and a geographical position feature; fusing the three single-modal features to obtain comprehensive features; calculating the similarity of the two cases according to the comprehensive features of the to-be-recognized case and the to-be-compared cases; and inputting the similarity into a binary classification network to determine whether the to-be-recognized case is a repeated case. The application not only significantly improves the accuracy and processing efficiency of repeated case recognition, but also enhances the adaptability of the system to complex scenes, and has strong practicability and expandability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of pattern recognition, specifically to a method for identifying recurring urban management cases based on multimodal data. Background Technology

[0002] With the continuous advancement of smart city construction and the increasing informatization of urban management, the collection and processing of urban management cases are increasingly reliant on multimodal data such as images, text, and location information uploaded by mobile terminals. To improve case processing efficiency, many urban management systems allow frontline personnel to collect information on-site using mobile phones and other devices and upload it to the urban management case processing platform. However, during the case reporting process, different reporting personnel may discover and upload photos of the same case separately, leading to a large number of duplicate cases. This not only increases the burden of data storage and management but may also affect the subsequent assignment and processing progress of cases, reducing the operational efficiency of the management system.

[0003] Currently, the deduplication of cases largely relies on manual identification and screening by staff under the urban management case processing platform. However, this method suffers from problems such as low identification accuracy and high labor costs, which seriously restricts the efficiency of daily case processing in the city. Summary of the Invention

[0004] To overcome at least one deficiency in the prior art, this application provides a method for identifying duplicate urban management cases based on multimodal data.

[0005] Firstly, a method for identifying recurring urban management cases based on multimodal data is provided, including: A high-frequency duplicate case database is constructed, which includes cases that are repeatedly reported more than a set threshold within a set time window; cases that match the type of the case to be identified are selected from the high-frequency duplicate case database as cases to be compared. A feature extraction network was used to extract feature information of the case to be identified and the case to be compared, including features of the crime scene image, text description features, and geographical location features. By integrating the crime scene image features, text description features, and geographical location features of the case to be identified, a comprehensive feature of the case to be identified is obtained; by integrating the crime scene image features, text description features, and geographical location features of the case to be compared, a comprehensive feature of the case to be compared is obtained. Based on the comprehensive characteristics of the case to be identified and the comprehensive characteristics of the case to be compared, the similarity results of the two cases are calculated. The similarity results are input into a binary classification network to determine whether the case to be identified and the case to be compared are duplicate cases.

[0006] In one embodiment, the feature extraction network includes EfficientNet, a BERT model, and a geolocation coding network; EfficientNet is used to extract crime scene image features based on case images, and the BERT model is used to extract text description features based on text descriptions. Geolocation coding networks consist of an input layer, a hidden layer, and an output layer, used to extract geolocation features based on latitude and longitude coordinates.

[0007] In one embodiment, the geolocation coding network is a trained network, and the loss function used during training is:

[0008] in, For loss function, The number of sample pairs Indicates the first If two cases in a sample pair are duplicate cases, then... , if not, , The Euclidean distance between the latitude and longitude coordinates of the two cases is given. This is a constant that controls the minimum distance between similar points.

[0009] In one embodiment, the crime scene image features, text description features, and geographic location features of the case to be identified are fused to obtain comprehensive features of the case to be identified, including: Geographic location features and crime scene image features are concatenated to generate a joint vector; The text description features are mapped to higher dimensions using a linear layer to obtain a transformed vector with the same dimension as the joint vector; The transformed vector is used as the query vector, and the joint vector is used as the key vector and value vector. Through cross attention, the enhanced text features are obtained. The joint vector is used as the query vector, and the transformation vector is used as the key vector and value vector. Through cross attention, the enhanced joint features are obtained. The enhanced text features and the enhanced joint features are added together to obtain the comprehensive features of the case to be identified.

[0010] In one embodiment, the similarity result between the two cases is calculated based on the comprehensive characteristics of the case to be identified and the comprehensive characteristics of the case to be compared, using the following formula:

[0011] in, For the comprehensive characteristics of case A to be identified, To compare the comprehensive characteristics of case B, The similarity results are for case A to be identified and case B to be compared. It represents the absolute value.

[0012] In one embodiment, the binary classification network includes a fully connected network and a Softmax layer; the fully connected network includes three fully connected layers with 256, 128, and 2 neurons respectively.

[0013] In one embodiment, constructing a high-frequency recurring case database includes: Cases that are repeatedly reported within a set time window and reach a set threshold are compiled into a high-frequency duplicate case database; the high-frequency duplicate case database is updated at a set period.

[0014] Secondly, a device for identifying duplicate urban management cases based on multimodal data is provided to implement the aforementioned method for identifying duplicate urban management cases based on multimodal data.

[0015] Compared with existing technologies, this application has the following advantages: This application integrates multimodal data of case images, text descriptions and geographic location information to construct a deep learning model, thereby achieving efficient identification of duplicate urban management cases; This method effectively overcomes the problems of low accuracy and high labor costs of traditional manual deduplication methods, which not only significantly improves the accuracy and processing efficiency of duplicate case identification, but also enhances the system's adaptability to complex scenarios, and has strong practicality and scalability, providing an intelligent and automated solution for smart urban management case management. Attached Figure Description

[0016] This application can be better understood by referring to the description given below in conjunction with the accompanying drawings, which, together with the detailed description below, are incorporated in and form part of this specification. In the drawings: Figure 1 A flowchart of a method for identifying duplicate urban management cases based on multimodal data is shown. Figure 2 A schematic diagram of the process for obtaining cases to be compared is shown; Figure 3 A schematic diagram of a geolocation coding network is shown; Figure 4 A schematic diagram of the feature fusion module is shown; Figure 5 A schematic diagram of a binary classification network is shown. Detailed Implementation

[0017] Exemplary embodiments of the present application will be described below with reference to the accompanying drawings. For clarity and brevity, not all features of the actual embodiments are described in the specification. However, it should be understood that many embodiment-specific decisions can be made in the development of any such actual embodiment to achieve the developer’s specific objectives, and these decisions may vary as the embodiments differ.

[0018] It should also be noted that, in order to avoid obscuring this application with unnecessary details, only the device structure closely related to the solution according to this application is shown in the accompanying drawings, while other details that are not closely related to this application are omitted.

[0019] It should be understood that this application is not limited to the described embodiments by virtue of the following description with reference to the accompanying drawings. In this document, embodiments may be combined with each other, features may be substituted or borrowed between different embodiments, and one or more features may be omitted in one embodiment, where feasible.

[0020] This application provides a method for identifying duplicate urban management cases based on multimodal data. Figure 1 A flowchart illustrating a method for identifying duplicate urban management cases based on multimodal data is shown. (See attached image) Figure 1 The method mainly includes the following steps: Step S1: Construct a high-frequency duplicate case database, which includes cases that are reported more than a set threshold within a set time window; and select cases in the high-frequency duplicate case database that match the type of the case to be identified as cases to be compared.

[0021] To ensure the real-time nature of duplicate case identification, a high-frequency duplicate case database is constructed to filter out cases that are reported repeatedly within a set time window. When a new case is uploaded, it is only compared with cases of the same category in the high-frequency duplicate case database, i.e., cases to be compared, thereby reducing the amount of computation for comparing new cases, improving comparison efficiency, and speeding up the judgment process.

[0022] Figure 2 The diagram illustrates the process of acquiring cases for comparison. The high-frequency duplicate case database is built based on a set time window (N days). Cases that are repeatedly reported at or above a threshold m within this time frame are counted and included in the high-frequency duplicate case database to improve the speed of comparing new cases. Furthermore, the database is updated every X days, recalculating the case duplication situation over the past N days, adding new high-frequency duplicate cases that meet the criteria, and removing outdated cases to maintain the database's timeliness.

[0023] In the process of counting the number of repeated cases, we first iterate through the historical case data from the past N days, and record it as follows: , Indicates the first The case, regarding the first Case Using the duplicate case identification network designed in subsequent steps, it is sequentially linked with its subsequent... Case Perform a comparison. If the model determines... and For duplicate cases, then Mark as duplicated, and The number of repetitions is incremented by one. After the traversal is complete, for each case... The number of times each case appeared within N days was recorded, and this information was then used for subsequent screening of high-frequency repeat cases.

[0024] When a new case is uploaded, cases matching that category are selected from the high-frequency duplicate case database. These are the cases to be compared. There may be one or more cases matching the category. The following steps determine whether the case to be identified is a duplicate of each case to be compared. The case category is selected by the person reporting the case when uploading it. Examples of pre-defined categories include gas pipelines under public utilities, green space obstruction in urban environments, and tree pits in landscaping.

[0025] Step S2: Use a feature extraction network to extract feature information of the case to be identified and the case to be compared. The feature information includes features of the crime scene image, text description features, and geographical location features.

[0026] Each case includes information such as a case image, text description, and latitude and longitude coordinates. In duplicate case identification, the case image provides crucial visual information for recognition. To better extract effective features from case images from complex backgrounds, this embodiment uses EfficientNet to extract image features from the case image, obtaining the crime scene image features. X i ; This embodiment selects the BERT (Bidirectional Encoder Representations from Transformers) model to extract text description features based on the text description. X t The BERT model is a pre-trained language representation model that uses a bidirectional Transformer structure. This allows the model to utilize the contextual information before and after each word when generating its representation. BERT also uses a multi-head self-attention mechanism to capture the interdependencies between different positions in the sequence. This mechanism can be computed in parallel and can model the dependencies between distant words in long text sequences.

[0027] Geographic location information is provided in latitude and longitude format. For two different cases, although the cases themselves are different, the coordinate text format may be very similar, making it impossible to effectively distinguish case characteristics simply by relying on coordinate values. Further analysis reveals that for the same case, the latitude and longitude coordinates should be relatively close in actual space; while for different cases, the located latitude and longitude coordinates should be relatively far apart. Therefore, based on this characteristic, this embodiment designs a geo-embedded network model to extract geographic location features from latitude and longitude coordinates, aiming to map each geographic coordinate to a high-dimensional vector space. The model's goal is to ensure that spatially close geographic coordinates maintain similar feature vectors in the high-dimensional space, while spatially distant coordinates should also be far apart in the high-dimensional space. In this way, the network can effectively learn the relative relationships between geographic locations, improving its ability to distinguish case similarity in subsequent tasks.

[0028] Figure 3 A schematic diagram of a geolocation coding network is shown; see [link / reference]. Figure 3 The geolocation coding network consists of an input layer, hidden layers, and an output layer. The input data for the input layer is the latitude and longitude coordinates of each case. The hidden layer contains two fully connected layers with 128 and 256 neurons respectively, and a ReLU activation function is applied after each hidden layer to enhance the non-linear expressive power of the model. The output layer has a dimension of 128, meaning that each geographic coordinate point is mapped to a 128-dimensional feature vector. X g .

[0029] Step S3: Integrate the crime scene image features, text description features, and geographical location features of the case to be identified to obtain the comprehensive features of the case to be identified; Integrate the crime scene image features, text description features, and geographical location features of the case to be compared to obtain the comprehensive features of the case to be compared.

[0030] After extracting the crime scene image features, text description features, and geographic location features of the case, the three unimodal features are fused to obtain the comprehensive features of the case. The case description text typically contains two main parts: "the location of the crime" and "a description of the case's issues." Therefore, geographic location information is associated with the "the location of the crime" part of the case description text, while image information is associated with the "a description of the case's issues" part. Thus, in the multimodal feature fusion algorithm, the geographic location features and case image features are first concatenated, and then the concatenated features are fused with the text description features to generate a unified multimodal feature representation of the case.

[0031] Specifically, Figure 4A schematic diagram of the feature fusion module is shown; see [link / reference]. Figure 4 The fusion process for the case to be identified and the case to be compared is the same. The text description features, case scene image features, and geographic location features are denoted as follows: X t , X i and X g Taking the fusion process of cases to be identified as an example, the fusion process includes: Geographical location features X g and crime scene image features X i Concatenate the vectors to generate a joint vector. X gi ; Text description features X t By using a linear layer for dimensionality-upgrading mapping, a transformation vector with the same dimension as the joint vector is obtained. ; Using transformation vectors As the query vector Q, a joint vector is used. X gi Using key vector K and value vector V, enhanced text features are obtained through cross-attention. ; Using joint vectors X gi As the query vector Q, the transformation vector is used. Using the key vector K and value vector V, enhanced joint features are obtained through cross-attention. ; Enhanced text features and enhanced joint features The addition operation is performed to obtain the comprehensive characteristics of the case to be identified.

[0032] Step S4: Calculate the similarity result between the two cases based on the comprehensive characteristics of the case to be identified and the comprehensive characteristics of the case to be compared. This can be achieved using a similarity calculation module.

[0033] To more comprehensively reflect the differences between case characteristics, the similarity calculation method designed in this embodiment measures the similarity between two cases (case A to be identified and case B to be compared) by calculating the element-wise difference of their feature vectors. Specifically, the feature vectors of case A and case B are subtracted element-wise and the absolute value is taken to obtain the difference vector between the two cases, i.e., the similarity result. The calculation formula is as follows:

[0034] in, For the comprehensive characteristics of case A to be identified, To compare the comprehensive characteristics of case B, The similarity results are for case A to be identified and case B to be compared. It represents the absolute value.

[0035] Step S5: Input the similarity results into the binary classification network to determine whether the case to be identified and the case to be compared are duplicate cases.

[0036] The similarity (difference vector) of comprehensive case features determines whether two cases are duplicates. This embodiment designs a deep learning-based binary classification network, which includes a fully connected network and a Softmax layer. The fully connected network consists of three layers with 256, 128, and 2 neurons respectively. Figure 5 A schematic diagram of a binary classification network is shown.

[0037] The similarity results of two cases are first input into a fully connected network. After feature transformation, a softmax layer is used to calculate the final classification probability to determine whether the two cases are duplicates. If the classification probability is close to 1 (e.g., reaching 0.9), the two cases are considered duplicates; otherwise, they are not. This step determines the classification probability between the case to be identified and each case to be compared, thus determining whether the two cases are duplicates.

[0038] In this embodiment, the feature extraction network, feature fusion module, similarity calculation module, and binary classification network constitute a duplicate case identification network. Here, the duplicate case identification network is the trained network.

[0039] The training process is as follows: First, a duplicate case identification dataset is constructed, which includes multiple positive sample pairs and multiple negative sample pairs. Positive sample pairs consist of images, text descriptions, and latitude / longitude coordinates of two duplicate cases, labeled 1. Negative sample pairs consist of images, text descriptions, and latitude / longitude coordinates of two non-duplicate cases, labeled 0. Then, the duplicate case identification network is trained based on the duplicate case identification dataset to obtain the trained duplicate case identification network.

[0040] To ensure that the coordinates of duplicate cases are mapped relatively close to each other in the feature space, while the coordinates of non-duplicate cases are mapped relatively far, a contrastive loss is used as the training objective when training the geolocation encoding network model. Contrastive loss optimizes the distance structure in the embedding space by minimizing the distance between similar samples and maximizing the distance between dissimilar samples. The loss function used during training is:

[0041] in, For loss function, The number of sample pairs Indicates the first If two cases in a sample pair are duplicate cases, then... , if not, , The Euclidean distance between the latitude and longitude coordinates of the two cases is given. To control the minimum distance between similar points, m is set to 0.5.

[0042] Binary cross-entropy loss is used as the loss function for the duplicate case identification network. The difference between the predicted value and the actual label is calculated to train the duplicate case identification network, and finally a deep neural network with duplicate discrimination ability is formed to identify duplicate cases reported by urban management inspectors.

[0043] The trained duplicate case identification network is then converted into the ONNX (Open Neural Network Exchange) format for deployment.

[0044] To verify the effectiveness of the duplicate case identification method based on multimodal data proposed in this application, relevant experiments were conducted. Specifically, the experiments included the following seven groups: 1. Identifying duplicate cases using only image, text, and geographic location features; 2. Identifying duplicate cases using a combination of image and text features, removing geographic location information; 3. Identifying duplicate cases using a combination of image and geographic location features, removing text features; 4. Identifying duplicate cases using a combination of text and geographic location features, removing image features; 5. Identifying duplicate cases using a combination of image, text, and geographic location features. The experimental results are shown in Table 1.

[0045] Table 1 Comparison of recognition performance with different feature combinations

[0046] The results show that when integrating case images, text, and geographic location information, the model achieves an accuracy of 95.64% and an F1 score of 97.13%, outperforming combinations of single-modality and dual-modality approaches. This demonstrates that the combined use of multimodal features can fully exploit the information potential of each modality, overcome the limitations of a single modality, and improve overall performance. Images provide intuitive problem information, text supplements detailed descriptions and location information, and geographic location provides spatial reference; their synergistic effect improves the model's accuracy in recurring case identification tasks.

[0047] This application also provides a device for identifying duplicate urban management cases based on multimodal data, used to implement the aforementioned method for identifying duplicate urban management cases based on multimodal data.

[0048] The device for identifying duplicate urban management cases based on multimodal data in this embodiment has the same inventive concept as the method for identifying duplicate urban management cases based on multimodal data described above. Therefore, the specific implementation of this device can be found in the embodiment section of the method for identifying duplicate urban management cases based on multimodal data described above, and its technical effects correspond to the technical effects of the above method, so it will not be repeated here.

[0049] The above descriptions are merely various embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for identifying duplicate urban management cases based on multimodal data, characterized in that, include: A high-frequency duplicate case database is constructed, which includes cases that are repeatedly reported more than a set threshold within a set time window; cases that match the type of the case to be identified are selected from the high-frequency duplicate case database as cases to be compared. A feature extraction network is used to extract feature information of the case to be identified and the case to be compared, respectively. The feature information includes crime scene image features, text description features, and geographical location features. By fusing the crime scene image features, text description features, and geographical location features of the case to be identified, a comprehensive feature of the case to be identified is obtained; by fusing the crime scene image features, text description features, and geographical location features of the case to be compared, a comprehensive feature of the case to be compared is obtained. Based on the comprehensive characteristics of the case to be identified and the comprehensive characteristics of the case to be compared, the similarity result of the two cases is calculated; The similarity results are input into a binary classification network to determine whether the case to be identified and the case to be compared are duplicate cases; The fusion process for the cases to be identified and the cases to be compared is the same; By integrating the crime scene image features, text description features, and geographical location features of the case to be identified, a comprehensive feature of the case to be identified is obtained, including: The geographic location features and the crime scene image features are concatenated to generate a joint vector; The text description features are subjected to dimensionality-upgrading mapping using a linear layer to obtain a transformation vector with the same dimension as the joint vector; Using the transformation vector as the query vector and the joint vector as the key and value vectors, enhanced text features are obtained through cross attention. Using the joint vector as the query vector and the transformation vector as the key vector and value vector, the enhanced joint feature is obtained through cross attention; The enhanced text features and the enhanced joint features are added together to obtain the comprehensive features of the case to be identified.

2. The method as described in claim 1, characterized in that, The feature extraction network includes EfficientNet, a BERT model, and a geolocation coding network; EfficientNet is used to extract crime scene image features based on case images, and the BERT model is used to extract text description features based on text descriptions. The geolocation coding network includes an input layer, a hidden layer, and an output layer, used to extract geolocation features based on latitude and longitude coordinates.

3. The method as described in claim 2, characterized in that, The geolocation coding network is a trained network, and the loss function used during training is: in, For loss function, The number of sample pairs Indicates the first If two cases in a sample pair are duplicate cases, then... , if not, , The Euclidean distance between the latitude and longitude coordinates of the two cases is given. This is a constant that controls the minimum distance between similar points.

4. The method as described in claim 1, characterized in that, in, Based on the comprehensive characteristics of the case to be identified and the comprehensive characteristics of the case to be compared, the similarity result of the two cases is calculated using the following formula: in, For the comprehensive characteristics of case A to be identified, To compare the comprehensive characteristics of case B, The similarity results are for case A to be identified and case B to be compared. It represents the absolute value.

5. The method as described in claim 1, characterized in that, The binary classification network includes a fully connected network and a Softmax layer; the fully connected network includes three fully connected layers with 256, 128, and 2 neurons respectively.

6. The method as described in claim 1, characterized in that, in, Constructing a database of frequently repeated cases includes: Cases that are repeatedly reported within a set time window and reach a set threshold are compiled into a high-frequency repeat case database; the high-frequency repeat case database is updated at a set period.

7. A device for identifying repeated urban management cases based on multimodal data, characterized in that, This method is used to implement the method for identifying duplicate urban management cases based on multimodal data as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Anti-fraud detection method and device, computer equipment and storage medium

    CN112232971A

  • Road damage detection method and device and application

    CN114267003A