AI-based intelligent city patrol system and methods

By using an AI-based intelligent city patrol system that combines optical and BeiDou satellite data for multimodal patrols, it achieves efficient and accurate hazard prediction for urban roads, buildings and ancillary facilities. This solves the problems of low patrol efficiency and large data errors in existing technologies and provides intelligent hazard early warning capabilities.

CN121073739BActive Publication Date: 2026-01-30SHAOXING ZHONGDAO ELECTRONIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511605435.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-01-30
Estimated Expiration
2045-11-05

AI Technical Summary

Technical Problem

Existing technologies for urban patrols are inefficient, prone to errors, and data is easily distorted, making it difficult to detect real problems in a timely manner. They also suffer from duplicate reporting and exaggeration of minor issues.

Method used

An AI-based intelligent city patrol system is adopted, which collects optical frame data and Beidou satellite signal data through multimodal patrol units, performs preprocessing and 3D reconstruction by combining visual analysis and Beidou positioning units, performs point cloud fusion by multimodal data fusion units, and performs deep reasoning by AI-driven units to achieve the prediction of hidden dangers in urban roads, buildings and ancillary facilities.

Benefits of technology

It improves the accuracy and intelligence of patrols, reduces monitoring blind spots, enables timely detection and early warning of potential hazards, and provides a multi-dimensional and multi-modal data foundation to ensure comprehensive coverage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121073739B_ABST
    Figure CN121073739B_ABST
Patent Text Reader

Abstract

An AI-based intelligent urban patrol system and method includes: a multimodal patrol unit, composed of acoustic sensors and optical components, for collecting optical frame data and acoustic echo data of the foundation pit construction area; a visual analysis unit, for preprocessing the optical frame data to extract an optical depth map containing local deformation information and pixel depth values; a BeiDou positioning unit, for preprocessing and 3D reconstruction of the acoustic echo data to determine the 3D point cloud of the foundation pit edge and internal structure; a multimodal data fusion unit, for point cloudification of the optical depth map to obtain an optical point cloud, and for fusing the optical point cloud with the 3D point cloud to obtain a fused point cloud and a corresponding 3D mesh model; and an AI-driven unit, for extracting feature tensors based on the 3D mesh model and performing depth inference based on the feature tensors to output the hazard prediction results for each slice, reducing blind spots in foundation pit monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and urban patrol technology, and more specifically, relates to an AI-based intelligent urban patrol system and method. Background Technology

[0002] With the continuous development of urban construction in my country, the defects of urban roads, buildings and ancillary facilities need to be discovered and dealt with in a timely manner. The original method of relying on manual visual inspection is inefficient, has large errors, and the data is easily distorted. It is also prone to problems such as difficulty in discovering real problems during urban inspections, repeated reporting of the same problems, and exaggeration of minor problems. Summary of the Invention

[0003] To address the shortcomings of existing technologies, the present invention aims to overcome the aforementioned deficiencies and propose an AI-based intelligent city patrol system and method.

[0004] The present invention adopts the following technical solution.

[0005] The first aspect of this invention discloses an AI-based intelligent city patrol system, the system comprising: a multimodal patrol unit, the multimodal patrol unit being composed of a BeiDou positioning component and an optical component, used to collect optical frame data and BeiDou satellite signal data of urban roads, buildings and ancillary facilities.

[0006] A visual analysis unit is used to preprocess the optical frame data to extract an optical depth map containing local deformation information and pixel depth values ​​from the optical frame data.

[0007] The Beidou positioning unit is used to preprocess and reconstruct the acoustic echo data in order to determine the three-dimensional point cloud of the pit edge and the internal structure of the pit.

[0008] The multimodal data fusion unit is used to perform point cloud processing on the optical depth map to obtain an optical point cloud, and to fuse the optical point cloud with the three-dimensional point cloud to obtain a fused point cloud and a corresponding three-dimensional mesh model.

[0009] The AI-driven unit is used to extract feature tensors based on the three-dimensional mesh model and perform deep inference based on the feature tensors to output the hazard prediction results for each slice.

[0010] The second aspect of this invention discloses an AI-based intelligent city patrol method, implemented through the AI-based intelligent city patrol system described in the first aspect. The method includes: acquiring video data of urban roads, buildings, and ancillary facilities, and BeiDou satellite signal data of the vehicle's location using a vehicle-mounted AI camera and a BeiDou positioning unit; preprocessing the optical frame data to extract an optical depth map containing local deformation information and pixel depth values; preprocessing and 3D reconstruction of the BeiDou satellite signal data to construct a 3D point cloud of the edges and internal structures of urban roads, buildings, and ancillary facilities; converting the optical depth map into an optical point cloud, and fusing and reconstructing the optical point cloud and the 3D point cloud to obtain a fused point cloud and its corresponding 3D mesh model; slicing the 3D mesh model to extract feature tensors from the 3D mesh model and the fused point cloud using a convolutional neural network model, and using a deep learning model to predict and output hazard prediction results for each slice based on the feature tensors.

[0011] Furthermore, the acquisition of video data of urban roads, buildings, and ancillary facilities, and BeiDou satellite signal data of the vehicle's location via an in-vehicle AI camera and BeiDou positioning unit includes: performing three-dimensional mapping of the edges of urban roads, buildings, and ancillary facilities using a total station and LiDAR to obtain the edge lengths and line shapes of urban roads, buildings, and ancillary facilities; calculating equal spacing based on the edge lengths and line shapes of urban roads, buildings, and ancillary facilities to obtain a deployment planning strategy for BeiDou positioning components and optical components; and setting up rigid supports based on the deployment planning strategy to integrate multimodal inspection units. An HDR camera is mounted at the top of the rigid support, filter components are mounted on both sides, a light projector is mounted at the bottom, and an ultrasonic array is embedded in the middle. Optical frame data and BeiDou satellite signal data are collected through the multimodal inspection units, and the optical frame data and BeiDou satellite signal data are synchronized with the GPS clock. At the same time, heartbeat detection is performed on each multimodal inspection unit. Based on the clock-synchronized optical frame data and BeiDou satellite signal data, a time-series optical frame sequence and an acoustic echo sequence are constructed respectively, and the optical frame sequence and acoustic echo sequence are fed back to the central server for verification.

[0012] Furthermore, the preprocessing of the optical frame data to extract an optical depth map containing local deformation information and pixel depth values ​​includes: acquiring image data with different polarization directions through the HDR camera to obtain different intensity matrices, and calculating polarization contrast based on the different intensity matrices to generate a polarization contrast map; projecting light patterns onto the walls of urban roads, buildings, and ancillary facilities using the light projector according to a preset strobe coding mode and a preset period, and simultaneously capturing the distorted image of the light patterns after projection onto the wall using the HDR camera to calculate the phase shift angle difference, which is used to calculate a depth map sequence containing local deformation information; using the RANSAC algorithm to perform inter-frame alignment of all frames in the depth map sequence based on adjacent frames in the depth map sequence, and calling linear weighted fusion to perform weighted temporal fusion on the inter-frame aligned depth map to obtain a fused depth map.

[0013] Furthermore, the preprocessing of the optical frame data to extract an optical depth map containing local deformation information and pixel depth values ​​further includes: using the polarization contrast map as a guide map, performing guided filtering on the fused depth map to generate a dehazed enhanced depth map, and performing median filtering and bilateral filtering on the dehazed enhanced depth map to obtain an initial optical depth map; marking positions in the polarization contrast map whose absolute values ​​exceed a set threshold as high-reflection regions, and extracting depth values ​​from the pixel positions corresponding to the high-reflection regions in the optical depth map to construct a high-reflection depth sub-map; performing neighborhood depth interpolation repair on the high-reflection depth sub-map, and using the repaired depth map as the first channel and the polarization contrast map as the second channel to construct a dual-channel fusion map, and generating the enhanced optical depth map based on the dual-channel fusion map according to a preset format encoding.

[0014] Furthermore, the preprocessing and 3D reconstruction of the BeiDou satellite signal data to construct a 3D point cloud of the edges and internal structures of urban roads, buildings, and ancillary facilities includes: calculating the time delay between different BeiDou positioning components based on the distance difference between each BeiDou positioning component, and generating a focused acoustic signal using a weighted delay summation algorithm; performing full-wave rectification and Hilbert transform on the focused acoustic signal to extract the envelope signal, and segmenting the envelope signal according to a time window to obtain a time-series mapping; calculating the maximum and minimum acoustic echoes within each distance dimension interval in the time-series mapping to normalize the BeiDou satellite signal data, and performing median filtering on the normalized BeiDou satellite signal data to obtain preprocessed BeiDou satellite signal data; calling a first-order difference algorithm to detect the echo peaks in the preprocessed BeiDou satellite signal data, and converting the time delay corresponding to the echo peaks into spatial distances to generate polar coordinates, which are used to calculate planar coordinates to construct a 3D point cloud of the edges and internal structures of urban roads, buildings, and ancillary facilities.

[0015] Furthermore, the step of converting the optical depth map into an optical point cloud and fusing and reconstructing the optical point cloud and the 3D point cloud to obtain a fused point cloud and its corresponding 3D mesh model includes: projecting the depth value corresponding to each pixel in the optical depth map into a 3D point to identify outlier 3D points outside a preset range and outputting the optical point cloud; extracting multiple sets of landmark point pairs from the 3D point cloud and the optical point cloud, and calculating the least-squares eccentric transformation parameters of each set of landmark point pairs to perform coarse registration of the optical point cloud; iteratively updating the optical point cloud based on the 3D point cloud using rigid transformation to obtain a registered optical point cloud; merging the registered optical point cloud with the 3D point cloud to obtain a fused point cloud, and generating a closed triangular mesh based on the fused point cloud using a Poisson surface reconstruction algorithm, while simultaneously performing Lapss smoothing on the closed triangular mesh to construct the 3D mesh model.

[0016] Furthermore, the step of slicing the 3D mesh model to extract feature tensors from the 3D mesh model and the fused point cloud using a convolutional neural network model, and then using a deep learning model to predict and output the hazard prediction results corresponding to each slice based on the feature tensors, includes: dividing the 3D mesh model into multiple voxel blocks, extracting vertices and faces from each voxel block to obtain multiple slices, and normalizing the coordinates of each slice; projecting the vertices of the slices onto the voxel mesh to generate a binary occupancy voxel map, calculating the normal vector and local curvature of each vertex, and mapping the normal vector and local curvature to a unified voxel mesh; concatenating the binary occupancy voxel map, normal vector, and local curvature to obtain the feature tensor, and using a 3D convolutional neural network to perform semantic segmentation on the feature tensor to output the hazard category probability and risk score of each slice; based on the hazard prediction results of each slice, clustering and merging the slices according to the hazard category probability, and extracting the boundaries of the merged regions to output a list of hazard regions, while prioritizing the hazard regions in the hazard region list according to the risk score.

[0017] A third aspect of the present invention discloses a terminal, including a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the method described in the second aspect.

[0018] The fourth aspect of the present invention discloses a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the second aspect.

[0019] The beneficial effects of the present invention are as follows: (1) By rationally planning and deploying multimodal patrol units on the complex urban roads, buildings and ancillary facilities terrain, the present invention can ensure that the Beidou positioning components and optical components uniformly and comprehensively cover the complex urban roads, buildings and ancillary facilities areas. At the same time, the data acquisition combines Beidou satellite signal data and optical frame data, providing a multi-directional and multimodal data foundation for the monitoring of urban roads, buildings and ancillary facilities areas, and reducing monitoring blind spots to a certain extent.

[0020] (2) Based on the collection of BeiDou satellite signal data and optical frame data of urban roads, buildings and ancillary facilities, this invention continuously registers the BeiDou satellite signal data and optical frame data, and then fuses the continuously registered acoustic point cloud and optical point cloud. A three-dimensional mesh model is constructed according to the fused point cloud, which makes the constructed three-dimensional mesh model more accurate and closer to the shape of urban roads, buildings and ancillary facilities. Finally, based on the three-dimensional mesh model, deep neural networks and deep learning models are called to predict the hidden dangers of urban roads, buildings and ancillary facilities. This not only further reduces the blind spots of urban roads, buildings and ancillary facilities monitoring, but also enables early warning of related hidden dangers based on monitoring data, with a high degree of intelligence. Attached Figure Description

[0021] Figure 1 This is a schematic diagram of the structure of the AI-based intelligent city patrol system provided by the present invention.

[0022] Figure 2 This is a flowchart illustrating the AI-based intelligent city patrol method provided by the present invention. Detailed Implementation

[0023] The present application will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention, and should not be construed as limiting the scope of protection of the present application.

[0024] like Figure 1 As shown, in one embodiment, an AI-based intelligent city patrol system includes: a multimodal patrol unit, which is composed of a Beidou positioning component and an optical component, used to collect optical frame data and Beidou satellite signal data of urban roads, buildings and ancillary facilities.

[0025] The visual analysis unit is used to preprocess the optical frame data to extract an optical depth map containing local deformation information and pixel depth values ​​from the optical frame data.

[0026] The BeiDou positioning unit is used to preprocess and reconstruct the acoustic echo data in order to determine the three-dimensional point cloud of the pit edge and the internal structure of the pit.

[0027] The multimodal data fusion unit is used to perform point cloud processing on the optical depth map to obtain an optical point cloud, and then fuse the optical point cloud with the 3D point cloud to obtain a fused point cloud and a corresponding 3D mesh model.

[0028] The AI-driven unit is used to extract feature tensors based on a 3D mesh model and perform deep inference based on the feature tensors to output the hazard prediction results for each slice.

[0029] like Figure 2As shown, in one embodiment, an AI-based smart city patrol method is implemented through the aforementioned AI-based smart city patrol system, including the following steps: Step S110, acquiring video data of urban roads, buildings and ancillary facilities and BeiDou satellite signal data of the vehicle's location through an in-vehicle AI camera and a BeiDou positioning unit.

[0030] In some embodiments, the AI-based smart city patrol method provided by the present invention includes the following steps in step S110: Step S111, using a total station and lidar to perform three-dimensional mapping of the edges of urban roads, buildings and ancillary facilities to obtain the length and line shape of the edges of urban roads, buildings and ancillary facilities, and calculating equal spacing based on the length and line shape of the edges of urban roads, buildings and ancillary facilities to obtain the deployment planning strategy of Beidou positioning components and optical components.

[0031] Step S112: Based on the deployment planning strategy, a rigid bracket is set up to integrate a multimodal inspection unit. An HDR camera is set at the top of the rigid bracket, filter components are set on both sides, a light projector is set at the bottom, and an ultrasonic array is embedded in the middle.

[0032] Step S113: Optical frame data and BeiDou satellite signal data are collected through the multimodal inspection unit, and the optical frame data and BeiDou satellite signal data are synchronized with the GPS clock. At the same time, heartbeat detection is performed on each multimodal inspection unit.

[0033] Step S114: Based on the clock-synchronized optical frame data and BeiDou satellite signal data, construct a time-series optical frame sequence and an acoustic echo sequence respectively, and feed the optical frame sequence and acoustic echo sequence back to the central server for verification.

[0034] In a specific embodiment, the AI-based smart city patrol method provided by the present invention includes steps 1 to 5: Step 1, deployment of multimodal patrol units and synchronous data collection.

[0035] It includes the following sub-steps: Sub-step 1.1, surveying and planning of the construction area.

[0036] Specifically, a total station or lidar is used to perform 3D mapping of the edges of urban roads, buildings, and ancillary facilities, obtaining edge lengths (in meters) and curve shapes to determine the number of inspection units. The unit spacing is then calculated based on the principle of equidistant spacing. Finally, a layout plan map with latitude, longitude, or planar coordinates is generated based on the formula results. This step rationally plans the unit locations on complex terrain, ensuring uniform coverage of subsequent sensors and effectively avoiding blind spots.

[0037] Sub-step 1.2: Assembly of device structure and hardware integration.

[0038] Specifically, a rigid aluminum alloy bracket is fabricated to ensure overall height and shock resistance requirements. An HDR camera is mounted at the top of the bracket, a controllable polarization filter is fixed to the side, a structured light projector is installed at the bottom, and an ultrasonic transducer array is embedded in the middle of the bracket. Next, a precision synchronization clock module is connected to all sensors via a dedicated coaxial cable, and an industrial Ethernet interface and power interface are configured on the bracket base. Each component is then packaged into a unit according to its serial number and labeled according to the coordinate diagram. This step achieves hardware integration, ensuring that the optical and BeiDou positioning components are deployed on a shared platform, facilitating transport and maintenance.

[0039] Sub-step 1.3, on-site setup and physical installation.

[0040] Specifically, the units are transported to the site, and the brackets are securely fixed to the edges of urban roads, buildings, and ancillary facilities' guardrails or safety platforms using expansion bolts or magnetic bases. Power lines and industrial Ethernet cables are laid to each unit's base, and waterproof sealing is performed. Protective covers and sunshades are installed to prevent dust and rainwater from directly corroding the sensor surface. Finally, the verticality and horizontality of each unit are checked according to their numbers, with an error not exceeding +2°. This process ensures that the units remain stable and immobile in the field environment, and that power and network connectivity are maintained, providing a reliable foundation for data acquisition.

[0041] Sub-step 1.4, Synchronization clock calibration and network configuration.

[0042] Specifically, a GPS reference clock is deployed in the central control room, the IEEE 1588 precision clock synchronization protocol is enabled, and a synchronization client is run on each unit to ensure that the local clock deviates from the reference clock by ≤±10 nanoseconds. Then, industrial Ethernet VLAN and QoS policies are configured to ensure that camera and ultrasonic data streams are transmitted through separate pipelines. Simultaneously, heartbeat packet detection confirms a packet loss rate of ≤0.1% and network latency of ≤5 milliseconds. This step achieves precise temporal alignment of optical and acoustic data, providing timing consistency for multimodal fusion.

[0043] Sub-step 1.5: Initial data acquisition and connectivity verification.

[0044] Specifically, all units are simultaneously triggered to perform a structured light projection and acoustic pulse emission. Each unit records HDR images and ultrasonic echoes within the same time window (calibrated by a synchronization clock). Afterward, the first batch of data is automatically uploaded to the central server, where a data integrity and timing alignment verification script is run. If the verification passes, the unit status is marked as "online"; otherwise, the process returns to sub-step 1.4 for recalibration. This step verifies the success of deployment and synchronization, ensuring that subsequent continuous inspection data possesses multimodal, high-quality, and time-consistent characteristics, laying the foundation for the next stage of data processing.

[0045] Step S120: Preprocess the optical frame data to extract an optical depth map containing local deformation information and pixel depth values ​​from the optical frame data.

[0046] In some embodiments, the AI-based smart city patrol method provided by the present invention includes the following steps in step S120: Step S121, acquiring image data with different polarization directions through an HDR camera to obtain different intensity matrices, and calculating polarization contrast based on the different intensity matrices to generate a polarization contrast map.

[0047] Step S122: A light projector projects light patterns onto the walls of urban roads, buildings and ancillary facilities according to a preset strobe coding mode and a preset cycle. An HDR camera is used to simultaneously capture the distorted image of the light patterns after projection onto the wall to calculate the phase shift angle difference. The phase shift angle difference is used to calculate a depth mapping sequence containing local deformation information.

[0048] Step S123: The RANSAC algorithm is used to perform inter-frame alignment of all frames in the depth map sequence based on adjacent frames in the depth map sequence, and linear weighted fusion is called to perform weighted temporal fusion of the inter-frame aligned depth map to obtain the fused depth map.

[0049] In some embodiments, the AI-based smart city patrol method provided by the present invention further includes the following steps in step S120: Step S124, using a polarization comparison map as a guide map, performing guided filtering on the fused depth map to generate a dehazing enhanced depth map, and performing median filtering and bilateral filtering on the dehazing enhanced depth map to obtain an initial optical depth map.

[0050] Step S125: Mark the positions in the polarization comparison map where the absolute value exceeds the set threshold as high reflectivity regions, and extract the depth values ​​at the corresponding pixel positions in the optical depth map to construct a high reflectivity depth sub-map.

[0051] Step S126: Perform neighborhood depth interpolation repair on the high-reflectivity depth sub-map, and use the repaired depth map as the first channel and the polarization contrast map as the second channel to construct a dual-channel fusion map, so as to generate an enhanced optical depth map based on the dual-channel fusion map according to a preset format encoding.

[0052] In a specific embodiment, the AI-based smart city patrol method provided by the present invention includes step 2: optical signal preprocessing and anti-interference enhancement.

[0053] It includes the following sub-steps: Sub-step 2.1, polarization image decomposition and transmission / reflection suppression.

[0054] Specifically, the camera acquires images at polarization angles of 0° and 90° respectively, and obtains the intensity matrix. and Calculate polarization contrast The expression is: In the formula, This is the polarization contrast matrix. and This represents the light intensity value of the same pixel at polarization directions of 0° and 90°.

[0055] After that, Normalization is performed, and threshold filtering removes invalid regions close to zero, generating a polarization contrast image. This step utilizes polarization differences to suppress local overexposure and "black holes" caused by diffuse reflection and metallic luster, restoring the true surface morphology information.

[0056] Sub-step 2.2: Structured light coding and decoding and depth extraction.

[0057] Specifically, the projector projects periodic light patterns onto the walls of urban roads, buildings, and ancillary facilities using a preset stroboscopic coding pattern (such as stripes, dot matrix, or grid). Simultaneously, the camera captures the distorted image of the coded light patterns projected onto the wall. Using the camera's in-camera calibration parameters, the phase shift of the stripes is decoded pixel by pixel, and the phase shift angle difference is calculated. Then, based on the mapping relationship between the phase shift difference and the calibration distance, the relative distance of each pixel to the projector is calculated, and these are stitched together to form a preliminary depth mapping matrix. This step, even under dust and stroboscopic interference, allows structured light coding to still provide high-precision deformation information, supplementing the depth details lacking in polarization decomposition.

[0058] Sub-step 2.3: Temporal frame fusion and dehazing / reflection enhancement.

[0059] Specifically, for the current frame and the frames before and after it, a rigid registration method based on feature points is used to extract the set of edge key points (such as SIFT features or ORB features) from each frame. The optimal affine transformation matrix is ​​calculated using the RANSAC algorithm, which maps the frames before and after the current frame to the coordinate system of the current frame to obtain the aligned depth map (for the frames before and after the current frame).

[0060] Then, a three-frame linear weighted fusion is used, expressed as: In the formula, For the generated fusion depth map, The first frame after alignment Current frame and the aligned subsequent frames The corresponding weights, when added together, equal 1. These are the pixel coordinates. The weights can be dynamically adjusted according to the scene: if the flash amplitude is large, the weight of the current frame can be reduced and the weight of adjacent frames can be increased to generate a fused depth map, effectively smoothing out depth anomalies caused by flashes or short-term occlusions.

[0061] Subsequently, using the aforementioned polarization comparison image as a guide image, guided filtering is applied to the generated fused depth map. That is, the guide image mean and variance are calculated for each pixel's neighborhood window, and then the linear coefficient and offset are calculated to finally generate a dehazing enhanced depth map. The depth map maintains depth jumps at the edge positions and removes haze interference in flat areas.

[0062] Finally, the dehazed and enhanced depth map is first subjected to median filtering (3×3 window size) to remove isolated noise points, and then bilateral filtering (spatial domain parameters and depth domain parameters) is performed. That is, while preserving the edges, it smooths out small depth fluctuations and outputs the final clear depth map.

[0063] Sub-step 2.4, surface reflection and geometric depth fusion enhancement.

[0064] Specifically, firstly, locations in the polarization contrast map where the absolute value exceeds a set threshold (e.g., 0.4) are marked as high-reflectivity regions. Depth values ​​are then extracted from the corresponding pixel locations in the final clear depth map to construct a high-reflectivity depth sub-map. Next, for depth distortion points within the high-reflectivity region set caused by light spots, a k×k neighborhood (e.g., 5×5) is selected around the high-reflectivity region set. All pixels within the high-reflectivity region set are excluded, and the neighborhood depth values ​​are weighted and averaged (based on the Euclidean distance from the center). The depth at the center location is then replaced with the weighted average to form the restored depth map.

[0065] Next, the repaired depth map is used as channel 1, and the polarization contrast map is scaled to [0,1] and used as channel 2 to construct a two-channel fusion map. Each channel in the two-channel fusion map needs to be multiplied by its corresponding reflection feature enhancement coefficient. Then, the two-channel fusion map is encoded according to a predefined format (such as 16-bit depth + 8-bit reflection) to output an enhanced optical depth map.

[0066] Finally, the error difference between the enhanced optical depth map and the preliminary depth map obtained in sub-step 2.2 is calculated. If the mean square error is reduced by more than 20%, the enhancement is deemed effective. Otherwise, the threshold or neighborhood size is adjusted to finally output a stable depth map without blind spots.

[0067] This step precisely locates and repairs highly reflective "dead zones," eliminates specular reflection artifacts, and simultaneously preserves geometric depth information and surface reflection features to form a multi-dimensional perception map. Through closed-loop parameter tuning based on quality assessment, it ensures that high-quality depth maps can be output under different optical interference conditions, providing a reliable foundation for subsequent acoustic-optical fusion and AI recognition.

[0068] Step S130 involves preprocessing and 3D reconstruction of the BeiDou satellite signal data to construct a 3D point cloud of the edges and internal structures of urban roads, buildings, and ancillary facilities.

[0069] In some embodiments, the AI-based smart city patrol method provided by the present invention includes the following steps in step S130: Step S131, calculating the time delay between different BeiDou positioning components based on the distance difference between each BeiDou positioning component, and generating a focused acoustic signal using a weighted delay summation algorithm.

[0070] Step S132: Perform full-wave rectification and Hilbert transform on the focused acoustic signal to extract the envelope signal, and segment the envelope signal according to the time window to obtain the time-series mapping.

[0071] Step S133: Calculate the maximum and minimum acoustic echoes within each distance dimension interval in the time-series mapping to normalize the BeiDou satellite signal data, and perform median filtering on the normalized BeiDou satellite signal data to obtain preprocessed BeiDou satellite signal data.

[0072] Step S134: Call the first-order difference algorithm to detect the echo peak in the preprocessed BeiDou satellite signal data, and convert the time delay corresponding to the echo peak into spatial distance to generate polar coordinates. The polar coordinates are used to calculate the planar coordinates to construct a three-dimensional point cloud of the edges and internal structures of urban roads, buildings and ancillary facilities.

[0073] In a specific embodiment, the AI-based intelligent city patrol method provided by the present invention includes step 3: acoustic signal preprocessing and three-dimensional reconstruction.

[0074] It includes the following sub-steps: Sub-step 3.1, board carrier beam formation and time delay compensation.

[0075] Specifically, acoustic transducers are arranged at equal intervals, with the spacing less than half the wavelength of a sound wave. The array is mounted in the center of the support frame and connected to a synchronization clock module via a coaxial cable. The time delay (the ratio of the geometric distance difference to the speed of sound (approximately 343 m / s)) is calculated based on the geometric distance difference from each unit to the target detection direction. Then, a weighted delay summation method is used to generate a focused acoustic signal. The expression is: In the formula, The number of acoustic transducers. Let i be the weight of the i-th acoustic transducer. The signal after time delay compensation. For time delay.

[0076] Sub-step 3.2, echo intensity normalization and noise suppression.

[0077] Specifically, the focused acoustic signal is rectified using full-wave rectification and Hilbert transform is applied to extract the envelope signal. The extracted envelope signal is then segmented according to time windows and mapped to the corresponding distance dimension. Next, the maximum and minimum echo values ​​are calculated within each distance interval. The envelope signal is then normalized based on the maximum and minimum echo values, and finally, median filtering (with a window length of 5) is applied to the normalized envelope signal to remove isolated noise points, resulting in a normalized echo intensity sequence and the corresponding emission angle.

[0078] Sub-step 3.3: Single scan planar point cloud extraction.

[0079] Specifically, a first-order difference algorithm is used to detect local maxima on the normalized echo intensity sequence to determine the number of peaks and echo peaks. Then, the time delay of each maximum is converted into distance, and polar coordinates are output in combination with the current array elevation angle.

[0080] Sub-step 3.4: Multi-angle scanning and stitching and two-dimensional point cloud generation.

[0081] Specifically, the polar coordinates of each scan are converted into planar coordinates (i.e., the product of each coordinate and its corresponding pitch angle trigonometric function), and then all planar coordinates (three-dimensional coordinates) under all scan angles are merged into the same point cloud. At the same time, the original point cloud is downsampled into a voxel grid (voxel size 5cm), and low-intensity echo points are removed.

[0082] Step S140: Convert the optical depth map into an optical point cloud, and fuse and reconstruct the optical point cloud and the 3D point cloud to obtain the fused point cloud and its corresponding 3D mesh model.

[0083] In some embodiments, the AI-based intelligent city patrol method provided by the present invention includes the following steps in step S140: Step S141, projecting the depth value corresponding to each pixel in the optical depth map into a three-dimensional point to extract outlier three-dimensional points outside the preset range and outputting an optical point cloud.

[0084] Step S142: Extract multiple sets of landmark point pairs from the 3D point cloud and the optical point cloud, and calculate the least squares radiative transformation parameters of each set of landmark point pairs to perform coarse registration of the optical point cloud, and perform iterative updates of the optical point cloud based on the 3D point cloud through rigid transformation to obtain the registered optical point cloud.

[0085] Step S143: The registered optical point cloud and the 3D point cloud are merged to obtain a fused point cloud. A closed triangular mesh is generated based on the fused point cloud using the Poisson surface reconstruction algorithm. At the same time, the closed triangular mesh is smoothed by Lapss smoothing to construct a 3D mesh model.

[0086] In a specific embodiment, the AI-based intelligent city patrol method provided by the present invention includes step 4: multimodal data registration and fusion reconstruction.

[0087] The process includes the following sub-steps: Sub-step 4.1, optical depth map point cloudification.

[0088] Specifically, the depth value of each pixel is back-projected into a 3D point according to the camera parameters, and the back-mapping operation is repeated for all valid depth pixels to remove outliers beyond 0.5-10m and output an optical point cloud set.

[0089] Sub-step 4.2, initial coarse registration based on landmark matching.

[0090] Specifically, N pairs of corresponding landmark points are extracted from both the acoustic and optical point clouds. Landmarks can be corners, feature block centers, or artificially placed targets. The parameters of the least-squares affine transformation (including rotation and translation) are calculated. The expression is: In the formula, It is a 3×3 rotation matrix. For a 3D translation vector, the landmark point pairs are: .

[0091] Subsequently, based on the least-squares affine transformation parameters (including rotation and translation), each optical point cloud in the optical point cloud geometry is... Perform a coarse transformation to obtain the coarsely registered optical point cloud. : .

[0092] Sub-step 4.3, fine registration based on the iterative nearest point algorithm and ICP.

[0093] Specifically, the following operations are repeated for the coarsely registered optical and acoustic point clouds: Find the nearest neighbor points in the acoustic point cloud set for the coarsely registered optical point cloud to form a corresponding set; recalculate the optimal rigid transformation to minimize the corresponding error; and then update the optical point cloud. This operation is iterated until the mean square error converges or the number of iterations reaches the upper limit (e.g., 50 times). Finally, record the final updated optical point cloud, which is the finely registered optical point cloud.

[0094] Sub-step 4.4, point cloud fusion and network reconstruction.

[0095] Specifically, the finely registered optical point cloud and acoustic point cloud are merged, and duplicate points (e.g., those less than 2cm apart) are eliminated by removing the density of the merged points. Voxel mesh downsampling (5cm voxel side length) is applied to balance density and detail. A closed triangular mesh is generated on the fused point cloud using the Poisson surface reconstruction algorithm, and the reconstruction depth level is set (e.g., 8 levels). The reconstructed mesh is then smoothed using Laplacian smoothing (3 iterations) to remove artifact spikes and fill small holes (maximum aperture 10cm).

[0096] Step S150: Slice the 3D mesh model into slices, call the convolutional neural network model to extract the feature tensors in the 3D mesh model and the fused point cloud, and call the deep learning model to predict the hazard prediction results corresponding to each slice based on the feature tensors.

[0097] In some embodiments, the AI-based smart city patrol method provided by the present invention includes the following steps in step S150: Step S151, dividing the three-dimensional mesh model into multiple voxel blocks, extracting vertices and faces in each voxel block to obtain multiple slices, and normalizing the coordinates of each slice.

[0098] Step S152: Project the vertices of the slice onto a voxel grid to generate a binary occupied voxel map, and calculate the normal vector and local curvature of each vertex, and map the normal vector and local curvature onto a uniform voxel grid.

[0099] Step S153: Channel concatenation of the binary occupancy voxel map, normal vector, and local curvature is performed to obtain a feature tensor. A 3D convolutional neural network is then used to perform semantic segmentation on the feature tensor to output the hazard category probability and risk score for each slice.

[0100] Step S154: Based on the hazard prediction results of each slice, the slices are clustered and merged according to the probability of hazard category, and the boundaries of the merged areas are extracted to output a list of hazard areas. At the same time, the hazard areas in the list of hazard areas are prioritized according to the risk score.

[0101] In a specific embodiment, the AI-based intelligent city patrol method provided by the present invention includes step 5, AI-driven structural hazard identification and location.

[0102] The process includes the following sub-steps: Sub-step 5.1, 3D model preprocessing and mesh slicing.

[0103] Specifically, the reconstructed mesh model is divided into several voxel blocks of a fixed voxel size (e.g., 0.5 m3). The internal vertices and faces of each voxel block are extracted to form slices. The coordinates of each slice are then centered and scaled to a uniform range (within a unit cube).

[0104] Sub-step 5.2, 3D feature extraction and point cloud encoding.

[0105] Specifically, the vertices in the slice are projected onto the voxel mesh (32 3 The process involves generating a binary occupancy voxel map, calculating the normal vector and local curvature for each vertex, and mapping these values ​​to the same voxel mesh to form additional channels. Then, the occupancy map, the normal vector channel (3 channels), and the curvature channel are concatenated to obtain the feature tensor. This step encodes 3D geometric information into multi-channel voxel features, enabling deep networks to simultaneously utilize multi-dimensional features such as structural occupancy, normal vectors, and curvature.

[0106] Sub-step 5.3, Deep network reasoning and hazard classification.

[0107] Specifically, a 3D convolutional neural network (such as 3D-UNet) is used for semantic segmentation. The network input is a set of feature tensors, and the output is a category probability map of the same size. The direction of the maximum value of the output probability map is then taken to determine the main hidden danger category of the slice. The mean or maximum probability of the category is used as the risk score.

[0108] Sub-step 5.4, spatial positioning and boundary fitting.

[0109] Specifically, similar hazards in adjacent slices are merged according to their spatial adjacency to form preliminary regions. The outer surface boundary vertices of the occupying voxels within these merged regions are then extracted. Next, the Ramer-Douglas-Peucker algorithm is used to refine the boundary curves and reduce redundant vertices. Finally, the boundary vertex coordinates are restored from slice-normalized coordinates to the model's global coordinate system. This step groups the slice-level results output by the network into continuous hazard regions, accurately identifying their spatial location and shape boundaries, facilitating on-site localization.

[0110] Sub-step 5.5: Generation and priority sorting of the hazard list.

[0111] Specifically, the hazard categories, area boundaries, and risk scores from each hazard area list are compiled into a list of items, sorted from highest to lowest risk score, to generate the final list. Finally, the list is encoded into JSON or a database record, with fields including hazard ID, category, score, array of boundary vertex coordinates, and recommended treatment level. This step, by outputting a structured, interactive hazard report and prioritizing it according to risk, provides clear priority guidance for inspectors or maintenance systems, thereby completing AI-driven structural hazard identification and location.

[0112] The applicant of this invention has provided a detailed description of the embodiments of the invention in conjunction with the accompanying drawings. However, those skilled in the art should understand that the above embodiments are merely preferred embodiments of the invention. The detailed description is only intended to help readers better understand the spirit of the invention and is not intended to limit the scope of protection of the invention. On the contrary, any improvements or modifications made based on the inventive spirit of the invention should fall within the scope of protection of the invention.

Claims

1. An AI-based intelligent city patrol system, characterized by, The system comprises: a multi-modal patrol unit composed of an acoustic sensor and an optical assembly, for collecting optical frame data and acoustic echo data of a foundation pit construction area; a visual analysis unit for preprocessing the optical frame data to extract an optical depth map containing local deformation information and pixel depth values in the optical frame data; a Beidou positioning unit for preprocessing and three-dimensional reconstruction of the acoustic echo data to determine a three-dimensional point cloud of the foundation pit edge and internal structure; a multi-modal data fusion unit for point cloud processing of the optical depth map to obtain an optical point cloud, and fusing the optical point cloud with the three-dimensional point cloud to obtain a fused point cloud and a corresponding three-dimensional mesh model; an AI driving unit for extracting a feature tensor based on the three-dimensional mesh model, and performing deep reasoning according to the feature tensor to output hazard prediction results of each slice, specifically comprising: dividing the three-dimensional mesh model into a plurality of voxel blocks to extract vertices and facets in each voxel block, obtaining a plurality of slices, and performing normalization processing on the coordinates of each slice; projecting the vertices of the slice to the voxel grid to generate a binary occupancy voxel map, and calculating the normal vector and local curvature of each vertex, and mapping the normal vector and local curvature to a unified voxel grid; channel splicing the binary occupancy voxel map, the normal vector and the local curvature to obtain the feature tensor, and performing semantic segmentation on the feature tensor using a 3D convolutional neural network to output a hazard class probability and a risk score of each slice; based on the hazard prediction results of each slice, clustering and merging the slices according to the hazard class probability, and extracting the boundary of the merged region to output a hazard region list, and simultaneously prioritizing the hazard regions in the hazard region list according to the risk score.

2. An AI-based intelligent city patrol method, characterized by, The AI-based intelligent city patrol system of claim 1 is implemented, and the method comprises: acquiring optical frame data and acoustic echo data of a foundation pit construction area through an acoustic sensor and an optical assembly; preprocessing the optical frame data to extract an optical depth map containing local deformation information and pixel depth values in the optical frame data; preprocessing and three-dimensional reconstruction of the acoustic echo data to construct a three-dimensional point cloud of the foundation pit edge and the internal structure of the foundation pit; converting the optical depth map into an optical point cloud, and fusing and reconstructing the optical point cloud and the three-dimensional point cloud to obtain a fused point cloud and a corresponding three-dimensional mesh model; performing mesh slicing on the three-dimensional mesh model to call a convolutional neural network model to extract feature tensors in the three-dimensional mesh model and the fused point cloud, and call a deep learning model to predict and output hidden danger prediction results corresponding to each slice based on the feature tensors, specifically comprising: dividing the three-dimensional mesh model into a plurality of voxel blocks to extract vertices and facets in each voxel block to obtain a plurality of slices, and performing normalization processing on the coordinates of each slice; projecting the vertices of the slice to the voxel grid to generate a binary occupancy voxel map, and calculating the normal vector and local curvature of each vertex, and mapping the normal vector and local curvature to a unified voxel grid; channel splicing the binary occupancy voxel map, normal vector and local curvature to obtain the feature tensors, and performing semantic segmentation on the feature tensors using a 3D convolutional neural network to output hidden danger class probability and risk score of each slice; based on the hidden danger prediction results of each slice, clustering and merging the slices according to the hidden danger class probability, and extracting the boundary of the merged region to output a hidden danger region list, and simultaneously prioritizing the hidden danger regions in the hidden danger region list according to the risk score. 3.The AI-based smart city patrolling method of claim 2, wherein, The AI-based intelligent city patrol system of claim 1 is implemented, and the method comprises: acquiring optical frame data and acoustic echo data of a foundation pit construction area through an acoustic sensor and an optical assembly; preprocessing the optical frame data to extract an optical depth map containing local deformation information and pixel depth values in the optical frame data; preprocessing and three-dimensional reconstruction of the acoustic echo data to construct a three-dimensional point cloud of the foundation pit edge and the internal structure of the foundation pit; converting the optical depth map into an optical point cloud, and fusing and reconstructing the optical point cloud and the three-dimensional point cloud to obtain a fused point cloud and a corresponding three-dimensional mesh model; performing mesh slicing on the three-dimensional mesh model to call a convolutional neural network model to extract feature tensors in the three-dimensional mesh model and the fused point cloud, and call a deep learning model to predict and output hidden danger prediction results corresponding to each slice based on the feature tensors, specifically comprising: dividing the three-dimensional mesh model into a plurality of voxel blocks to extract vertices and facets in each voxel block to obtain a plurality of slices, and performing normalization processing on the coordinates of each slice; projecting the vertices of the slice to the voxel grid to generate a binary occupancy voxel map, and calculating the normal vector and local curvature of each vertex, and mapping the normal vector and local curvature to a unified voxel grid; channel splicing the binary occupancy voxel map, normal vector and local curvature to obtain the feature tensors, and performing semantic segmentation on the feature tensors using a 3D convolutional neural network to output hidden danger class probability and risk score of each slice; based on the hidden danger prediction results of each slice, clustering and merging the slices according to the hidden danger class probability, and extracting the boundary of the merged region to output a hidden danger region list, and simultaneously prioritizing the hidden danger regions in the hidden danger region list according to the risk score. 4.The AI-based smart city patrolling method of claim 3, wherein, The preprocessing of the optical frame data to extract the optical depth map containing local deformation information and pixel depth values in the optical frame data comprises: acquiring image data of different polarization directions by the HDR camera to obtain different intensity matrices, and calculating a polarization contrast based on the different intensity matrices to generate a polarization contrast map; projecting a striae on the pit wall surface according to a preset stroboscopic coding mode and a preset period by the light projector, and synchronously capturing a distortion image of the striae after projection on the wall surface by the HDR camera to calculate a phase shift angle difference value, which is used to calculate a depth mapping sequence containing local deformation information; using a RANSAC algorithm to perform inter-frame alignment on all frames in the depth mapping sequence based on adjacent frames in the depth mapping sequence, and calling a linear weighted fusion to perform weighted time series fusion on the inter-frame aligned depth mapping to obtain a fused depth map. 5.The AI-based smart city patrolling method of claim 4, wherein, The preprocessing of the optical frame data to extract the optical depth map containing local deformation information and pixel depth values in the optical frame data further comprises: using the polarization contrast map as a guide map to guide filtering of the fused depth map to generate a defogging enhanced depth map, and performing median filtering and bilateral filtering on the defogging enhanced depth map to obtain an initial optical depth map; marking positions with absolute values exceeding a set threshold in the polarization contrast map as high reflection areas, and extracting depth values at pixel positions corresponding to the high reflection areas in the optical depth map to construct a high reflection depth submap; performing neighborhood depth interpolation repair on the high reflection depth submap, and using the repaired depth map as a first channel and the polarization contrast map as a second channel to construct a dual-channel fusion map, so as to encode the enhanced optical depth map based on the dual-channel fusion map in a preset format. 6.The AI-based smart city patrolling method of claim 2, wherein, The preprocessing and three-dimensional reconstruction of the acoustic echo data to construct a three-dimensional point cloud of the pit edge and the pit internal structure comprises: calculating a time delay between different acoustic sensors according to a distance difference between the acoustic sensors, and generating a focused acoustic signal by using a weighted delay sum algorithm; performing full-wave rectification and Hilbert transform on the focused acoustic signal to extract an envelope signal, and segmenting the envelope signal according to a time window to obtain a time sequence mapping; calculating a maximum acoustic echo and a minimum acoustic echo in each distance dimension interval in the time sequence mapping to normalize the acoustic echo data, and performing median filtering on the normalized acoustic echo data to obtain preprocessed acoustic echo data; calling a first-order difference algorithm to detect echo peak values in the preprocessed acoustic echo data, and converting time delays corresponding to the echo peak values into spatial distances to generate polar coordinates, which are used to calculate plane coordinates to construct a three-dimensional point cloud of the pit edge and the pit internal structure. 7.The AI-based smart city patrolling method of claim 2, wherein, The converting the optical depth map into an optical point cloud and fusing the optical point cloud and the three-dimensional point cloud to reconstruct a fused point cloud and a corresponding three-dimensional mesh model comprises: projecting a depth value corresponding to each pixel in the optical depth map into a three-dimensional point to propose an outlying three-dimensional point outside a preset range, and outputting an optical point cloud; extracting a plurality of groups of landmark point pairs from the three-dimensional point cloud and the optical point cloud, and calculating least square radiation transformation parameters of each group of landmark point pairs to coarsely register the optical point cloud, and iteratively updating the optical point cloud based on the three-dimensional point cloud through a rigid transformation to obtain a registered optical point cloud; merging the registered optical point cloud and the three-dimensional point cloud to obtain a fused point cloud, and generating a closed triangular mesh based on the fused point cloud through a Poisson surface reconstruction algorithm, and simultaneously performing Laplace smoothing processing on the closed triangular mesh to construct the three-dimensional mesh model. 8.A terminal, comprising a processor and a storage medium; characterized in that: The storage medium is configured to store instructions, and the processor is configured to operate according to the instructions to perform steps of the method according to any one of claims 2-7.

9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program, when executed by the processor, implements the steps of the method according to any one of claims 2-7.

Citation Information

Patent Citations

  • Unmanned aerial vehicle routing inspection line adaptive obstacle detection method and system based on monocular camera

    CN119672577A

  • Urban surveying and mapping system and method thereof

    TW202531137A