City road online monitoring method and system based on image recognition and internet of things

By using multipath physical modeling and image recognition technology to complete the outline of the occluded area, a holographic traffic situation map is generated, and a temporal elasticity metric space is constructed for sequence alignment. This solves the problem of insufficient accuracy in the inversion of target attributes in blind spots in urban road monitoring, and improves the accuracy of traffic situation perception and control efficiency.

CN122493399APending Publication Date: 2026-07-31SICHUAN KANGJISHENG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN KANGJISHENG TECHNOLOGY CO LTD
Filing Date
2026-07-01
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies for urban road monitoring suffer from insufficient accuracy in inverting target attributes in blind spots and reduced physical reliability of fused data. In particular, it is difficult to perform accurate attribute inversion using the physical characteristics of radio frequency signals in severely obstructed scenarios, and there are timing deviations when synchronizing multi-source data.

Method used

By extracting pure reflected wave features through multipath physical modeling and combining them with image recognition models to complete the contour of the occluded area, a holographic traffic situation map is generated. A temporal elasticity metric space is constructed for sequence alignment, and physical kinematic constraints are used to determine the spatiotemporal topological mismatch state, thereby realizing data uploading and traffic control.

Benefits of technology

It improves the accuracy of target perception in blind spots in non-line-of-sight scenarios, enhances the physical reliability of traffic situation maps and the efficiency of traffic control, and reduces the bandwidth consumption of the Internet of Things.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493399A_ABST
    Figure CN122493399A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for online monitoring of urban roads based on image recognition and the Internet of Things (IoT). It relates to the field of image recognition technology and includes: acquiring real-time visible light video streams and radio frequency (RF) sensing data of the urban road monitoring area; performing multipath physical modeling on the RF sensing data; extracting pure reflected wave features caused by dynamic targets; converting the pure reflected wave features into RF semantic vectors representing the physical properties of targets in blind zones; inputting the RF semantic vectors into a pre-trained image recognition model; combining the vehicle edge contours within the visible area to complete the contours of occluded areas, generating a holographic traffic situation map; and constructing a temporal elasticity metric space and aligning the sequences based on the holographic traffic situation map by extracting image feature vector sequences and combining them with the RF semantic vector sequences. This invention reduces IoT bandwidth consumption while improving the perception accuracy of the holographic traffic situation map and the execution efficiency of traffic control commands.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition technology, and in particular to a method and system for online monitoring of urban roads based on image recognition and the Internet of Things. Background Technology

[0002] Urban road monitoring technology based on the fusion of computer vision and millimeter-wave radar has been widely applied in intelligent transportation systems. Existing technologies typically utilize deep learning algorithms to perform target detection on video streams captured by cameras, supplemented by radar point cloud data for feature fusion to improve perception robustness in complex environments. These solutions often focus on 2D or 3D reconstruction of visible areas, while perception of non-line-of-sight areas still relies primarily on mathematical extrapolation based on historical trajectories, lacking in-depth analysis of the physical characteristics of electromagnetic signal propagation. This results in blind spots in situational awareness under severely obstructed conditions. Existing solutions struggle to accurately invert the attributes of targets in blind spots using the physical characteristics of radio frequency signals, and often employ rigid timestamps for multi-source data synchronization, ignoring the propagation differences between optical imaging and electromagnetic wave reflection in different media. This leads to phase discrepancies between image features and radio frequency features in the time dimension, consequently reducing the physical reliability of the fused data and impacting the efficiency of subsequent traffic control decisions. Summary of the Invention

[0003] In view of the aforementioned existing problems, the present invention is proposed.

[0004] Therefore, this invention provides an online monitoring method for urban roads based on image recognition and the Internet of Things to solve the problems of insufficient accuracy in inverting target attributes in blind spots and decreased physical reliability of fused data caused by the lack of physical feature constraints and time-series flexible alignment mechanisms in existing solutions.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides an online monitoring method for urban roads based on image recognition and the Internet of Things, which includes: collecting real-time visible light video streams and radio frequency sensing data of the urban road monitoring area; performing multipath physical modeling on the radio frequency sensing data; extracting pure reflection wave features caused by dynamic targets; and converting the pure reflection wave features into radio frequency semantic vectors of the physical properties of blind zone targets. Radio frequency semantic vectors are input into a pre-trained image recognition model, and the contours of the occluded areas are completed by combining the vehicle edge contours within the visible area to generate a holographic traffic situation map. Based on the holographic traffic situation map, image feature vector sequences are extracted and combined with radio frequency semantic vector sequences to construct a temporal elasticity metric space and align the sequences. The geodesic distance of the aligned sequences in the high-dimensional topological manifold space is calculated, and the spatiotemporal topological mismatch state is determined based on the geodesic distance. Based on the determined spatiotemporal topological mismatch state, the data corresponding to the holographic traffic situation map is uploaded to the cloud, and corresponding traffic control operations are performed according to the data corresponding to the holographic traffic situation map.

[0006] As a preferred embodiment of the urban road online monitoring method based on image recognition and the Internet of Things described in this invention, wherein: the radio frequency semantic vector of the physical attributes of the blind zone target includes, Using prior environmental knowledge, a multipath physical modeling based on propagation path decomposition is performed on the electromagnetic propagation channel response corresponding to real-time visible light video streams and radio frequency sensing data. Time-varying topological manifold filtering is used to extract the time-varying scattering component caused by the dynamic target by nonlinear manifold unwrapping of the dynamic micro-Doppler component in the phase space, generating multipath physical modeling results. Fractional Fourier transform is performed on the multipath physics modeling results to separate the linear frequency modulated signals of different scatterers on the rotating time-frequency plane. The spatial topology of the reflected wave is reconstructed based on the correlation matrix of the scattering points to remove the residual fixed clutter and extract the features of the pure reflected wave caused by the dynamic target. Phase difference sequences are extracted by Riemannian geometric phase unwrapping of pure reflected wave characteristics, and a physical feature fingerprint database is established by combining the dielectric constant response characteristics of electromagnetic waves in different materials. By matching geodesics on the manifold, the phase difference sequence is mapped to the target's position coordinates and physical attribute classification labels, and the pure reflected wave features are converted into radio frequency semantic vectors of the target's physical attributes in the blind zone.

[0007] As a preferred embodiment of the urban road online monitoring method based on image recognition and the Internet of Things described in this invention, the holographic traffic situation map includes: The physical attribute labels of the radio frequency semantic vector are mapped to the joint embedding features of the spatial location of the vehicle edge contour within the visible area by using the physical constraint loss function. The spatial location of the vehicle edge contour is jointly embedded into the cross-modal attention layer of the pre-trained image recognition model, thereby modulating the response strength of the pre-trained image recognition model to the blind zone region features indicated by the radio frequency semantic vector. The vehicle edge contours within the visible area are converted into a spatial location mask matrix. The feature map response range of the pre-trained image recognition model is constrained by element-wise multiplication to obtain the mask-constrained pre-trained image recognition model. An adaptive probability extension decoder is constructed based on the physical properties indicated by the radio frequency semantic vector. During the training phase, the adaptive probability extension decoder learns the mapping relationship between large metal targets and rectangular probability distributions, and between pedestrians and cylindrical probability distributions. The contour of the occluded area is completed by an adaptive probabilistic extension decoder, and the completed occluded area contour is obtained. The visible area image and the completed occluded area contour are fused together, and the spatial position, velocity and physical attributes of each target are marked in the fusion result to generate a holographic traffic situation map.

[0008] As a preferred embodiment of the urban road online monitoring method based on image recognition and the Internet of Things described in this invention, the geodesic distance includes: Extract image feature vector sequences carrying target physical attributes from holographic traffic situation maps, and extract radio frequency semantic vector sequences that match the target physical attributes from radio frequency semantic vectors; A time elasticity metric space is constructed based on the physical kinematic constraints corresponding to the target's physical properties. A dynamic time warping algorithm constrained by physical kinematic constraints is used to perform nonlinear time alignment between the matched image feature vector sequence and the radio frequency semantic vector sequence. The aligned sequence is mapped to a high-dimensional topological manifold space. In the high-dimensional topological manifold space, the integral path of the geodesic distance is adjusted using the target physical properties as the intrinsic curvature parameter of the manifold. The geodesic distance between the adjusted image feature vector sequence and the radio frequency semantic vector sequence is calculated.

[0009] As a preferred embodiment of the urban road online monitoring method based on image recognition and the Internet of Things described in this invention, the determination of the spatiotemporal topological mismatch state includes, Based on the geodesic distance between the adjusted image feature vector sequence and the radio frequency semantic vector sequence, a causal constraint of physical motion is constructed, and the spatiotemporal topological mismatch state is determined according to the causal constraint of physical motion.

[0010] As a preferred embodiment of the urban road online monitoring method based on image recognition and the Internet of Things described in this invention, the step of uploading the data corresponding to the holographic traffic situation map to the cloud includes: Based on the spatiotemporal topological mismatch state determined by the judgment, the data corresponding to the holographic traffic situation map is classified and encapsulated according to the target's physical attributes. Based on the target physical attributes carried by the data corresponding to the encapsulated holographic traffic situation map, select the corresponding transmission channel and use the selected transmission channel as the carrier for uploading the data corresponding to the holographic traffic situation map. The data corresponding to the encapsulated holographic traffic situation map is uploaded to the cloud through the selected transmission channel.

[0011] As a preferred embodiment of the urban road online monitoring method based on image recognition and the Internet of Things described in this invention, the execution of the corresponding traffic control operation includes: The data uploaded to the cloud is parsed to distinguish between structured semantic information and abnormal data packets, and the corresponding traffic control plan is matched based on the parsing results; Based on the matched traffic control plan, a handling instruction is generated and sent to the execution unit via the Internet of Things to carry out traffic control operations.

[0012] In a second aspect, the present invention provides an online monitoring system for urban roads based on image recognition and the Internet of Things, including an extraction module that collects real-time visible light video streams and radio frequency sensing data of the urban road monitoring area, performs multipath physical modeling on the radio frequency sensing data, extracts pure reflection wave features caused by dynamic targets, and converts the pure reflection wave features into radio frequency semantic vectors of the physical properties of blind zone targets. The completion module inputs the radio frequency semantic vector into the pre-trained image recognition model, and combines the vehicle edge contours within the visible area to complete the contours of the occluded area, generating a holographic traffic situation map. The alignment module extracts image feature vector sequences from holographic traffic situation maps and combines them with radio frequency semantic vector sequences to construct a temporal elasticity metric space and align the sequences. It then calculates the geodesic distance of the aligned sequences in a high-dimensional topological manifold space and determines the spatiotemporal topological mismatch state based on the geodesic distance. The execution module, based on the determined spatiotemporal topological mismatch state, uploads the data corresponding to the holographic traffic situation map to the cloud, and performs corresponding traffic control operations based on the data corresponding to the holographic traffic situation map.

[0013] Thirdly, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein the computer program, when executed by the processor, implements any step of the online monitoring method for urban roads based on image recognition and the Internet of Things as described in the first aspect of the present invention.

[0014] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the online urban road monitoring method based on image recognition and the Internet of Things as described in the first aspect of the present invention.

[0015] The beneficial effects of this invention are as follows: By extracting pure reflected wave features through multipath physical modeling and Riemannian geometric phase unwrapping, and converting them into radio frequency semantic vectors containing position coordinates and physical properties, the problem of inaccurate inversion of target attributes in blind spots under non-line-of-sight scenarios is solved, providing strong physical-level prior constraints for visual completion. By constructing a time elasticity metric space and adjusting the geodesic integral path using the target physical properties as manifold curvature parameters, nonlinear alignment and physical reliability metric quantification of images and radio frequency sequences are achieved, overcoming the spatiotemporal misalignment caused by rigid synchronization. Combined with dynamic data upload based on physical motion causality verification, the perception accuracy of holographic traffic situation maps and the execution efficiency of traffic control instructions are improved while reducing IoT bandwidth consumption. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart of an online monitoring method for urban roads based on image recognition and the Internet of Things.

[0018] Figure 2 This is a schematic diagram of an online urban road monitoring system based on image recognition and the Internet of Things. Detailed Implementation

[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0020] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0021] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0022] Reference Figures 1-2 As one embodiment of the present invention, this embodiment provides a method for online monitoring of urban roads based on image recognition and the Internet of Things, including the following steps: S1. Collect real-time visible light video streams and radio frequency sensing data of urban road monitoring areas, perform multipath physical modeling on radio frequency sensing data, extract pure reflection wave features caused by dynamic targets, and convert pure reflection wave features into radio frequency semantic vectors of physical properties of blind zone targets.

[0023] S1.1. Using prior environmental knowledge, multipath physical modeling based on propagation path decomposition is performed on the electromagnetic propagation channel response corresponding to real-time visible light video stream and radio frequency sensing data. Time-varying topological manifold filtering is used to extract the time-varying scattering component caused by the dynamic target by nonlinear manifold unwrapping of the dynamic micro-Doppler component in the phase space, generating multipath physical modeling results.

[0024] Furthermore, prior environmental knowledge is used to model the propagation loss, obstacle distribution, and electromagnetic propagation characteristics of each reflection path in real-time visible light video streams and radio frequency sensing data. A multipath physical model including static and dynamic multipaths is constructed. Time-varying topological manifold filtering is used to nonlinearly unwrap the dynamic micro-Doppler components in phase space. By adjusting the curvature on the manifold, the time-varying scattering components caused by the dynamic target and the static background scattering components are separated to generate multipath physical modeling results. Prior environmental knowledge constrains the parameter space of multipath physical modeling, reducing modeling complexity and improving the matching degree between the model and the real scene. Time-varying topological manifold filtering adapts to the micro-Doppler changes of the dynamic target by dynamically adjusting the manifold structure. Nonlinear manifold unwrapping avoids the truncation error of time-varying components by traditional linear filtering. The extracted time-varying scattering components accurately reflect the true reflection characteristics of the dynamic target.

[0025] S1.2 Perform fractional Fourier transform on the multipath physical modeling results to separate the linear frequency modulated signals of different scatterers on the rotating time-frequency plane, reconstruct the spatial topology of the reflected wave based on the correlation matrix of the scattering points, remove the residual fixed clutter, and extract the features of the pure reflected wave caused by the dynamic target.

[0026] Furthermore, a fractional Fourier transform is performed on the multipath physics modeling results. The linear frequency modulated (LFM) signals of different scatterers are projected onto the optimal rotation angle on the rotating time-frequency plane, achieving separation of the LFM signals from noise and other interference signals. The spatial topology of the reflected wave is reconstructed based on the scattering point correlation matrix. The rank analysis of the correlation matrix distinguishes between dynamic and fixed scattering points, eliminates fixed clutter residues, and extracts the pure reflected wave features caused by dynamic targets. The operation of the fractional Fourier transform on the rotating time-frequency plane breaks through the fixed basis function limitation of the traditional Fourier transform, is more suitable for the time-frequency characteristics of LFM signals, and improves the signal separation accuracy. The rank analysis of the spatial topology of the scattering point correlation matrix accurately identifies fixed scattering points, avoiding the erroneous elimination of weak dynamic signals by the traditional threshold method. The extracted pure reflected wave features retain the complete reflection information of the dynamic target.

[0027] S1.3. Extract the phase difference sequence by Riemann geometric phase unwrapping of the pure reflected wave characteristics, and establish a physical feature fingerprint database by combining the dielectric constant response characteristics of electromagnetic waves in different materials.

[0028] Furthermore, Riemannian geometric phase unwrapping is performed on the pure reflected wave characteristics. The phase entanglement effect is eliminated by geodesic integration on the Riemann manifold, and the phase difference sequence reflecting the target's physical properties is extracted. Combined with the dielectric constant response characteristics of electromagnetic waves in different materials, a physical feature fingerprint library containing dielectric constant characteristics of typical materials such as metals, water bodies, and pedestrians is constructed. The geodesic integration on the Riemann manifold avoids the cumulative error of traditional phase unwrapping. The phase difference sequence more accurately characterizes the dielectric properties of the target. By integrating the dielectric constant response characteristics of different materials, the physical feature fingerprint library establishes a precise correspondence between the phase difference sequence and the target's physical properties.

[0029] S1.4. By matching geodesics on the manifold, the phase difference sequence is mapped to the target's position coordinates and physical attribute classification labels, and the pure reflected wave features are converted into radio frequency semantic vectors of the target's physical attributes in the blind zone.

[0030] Furthermore, by comparing the phase difference sequence with features in the physical feature fingerprint database through geodesic matching on the manifold, the geodesic distance between the phase difference sequence and the features in the fingerprint database is obtained in the manifold space. The feature corresponding to the minimum geodesic distance is selected as the target's position coordinates and physical attribute classification label. The pure reflected wave features are converted into radio frequency semantic vectors of the physical attributes of the blind zone target, which include position coordinates and physical attribute classification labels. Geodesic matching on the manifold utilizes the geometric structure characteristics of the manifold space to measure the similarity of high-dimensional phase difference sequences more accurately than Euclidean distance, accurately mapping the target's position coordinates and physical attribute classification labels. The converted radio frequency semantic vectors of the physical attributes of the blind zone target integrate the target's physical attributes and spatial location information.

[0031] S2. Input the radio frequency semantic vector into the pre-trained image recognition model, and combine it with the vehicle edge contours in the visible area to complete the contours of the occluded area, generating a holographic traffic situation map.

[0032] S2.1. The physical attribute labels of the radio frequency semantic vector are mapped to the spatial location joint embedding features of the vehicle edge contour within the visible area through the physical constraint loss function.

[0033] Furthermore, the physical attribute labels of the radio frequency semantic vector are mapped to the joint embedding features of the spatial location of the vehicle edge contour within the visible area through the physical constraint loss function. The physical constraint loss function applies geometric consistency constraints on the target physical attributes and spatial location during the mapping process, so that the physical attribute labels of the radio frequency semantic vector and the vehicle edge contour within the visible area form an associated embedding representation in spatial location. The associated embedding representation contains the physical attribute information of the radio frequency semantic vector and the spatial location information of the vehicle edge contour within the visible area.

[0034] Specifically, the physical constraint loss function forces the physical attribute labels of the radio frequency semantic vector to be spatially aligned with the vehicle edge contours within the visible area through geometric consistency constraints. This avoids the problem of physical attributes and spatial positions being disconnected in traditional mapping methods. The spatial position joint embedding feature integrates the physical attribute information of the radio frequency semantic vector with the contour position information of the visible area, providing a precise feature association basis for the cross-modal attention layer. This enables the cross-modal attention layer to accurately modulate the feature response based on the association between physical attributes and spatial positions.

[0035] It should be noted that the pre-trained image recognition model adopts a cross-modal fusion architecture, including an image feature extraction module, a cross-modal feature fusion module, and a contour generation module. During the training phase, training samples are composed of visible light images, vehicle edge contour annotations, spatial location annotations, and corresponding radio frequency semantic vectors. The physical constraint loss function and contour reconstruction error are used as optimization objectives to complete model training. The visible area vehicle edge contour features are extracted by the image feature extraction module, and the physical attribute labels of the radio frequency semantic vectors are mapped to the spatial location joint embedding features corresponding to the vehicle edge contours using the physical constraint loss function. Subsequently, the spatial location joint embedding features are input into the cross-modal attention layer and fused with the image features to learn the correlation between the spatial location of the vehicle edge contours and the radio frequency semantic vectors. The vehicle edge contours are converted into spatial location mask matrices. The response range of the model feature map is constrained by element-wise multiplication, retaining only the effective features of the visible area. The constrained features are input into the adaptive probability extension decoder to learn the mapping relationship between large metal targets and rectangular probability distributions, and between pedestrians and cylindrical probability distributions. The contours of the occluded areas are then recovered based on the learned probability distributions. Once the image recognition model has been trained and meets the preset convergence conditions, the parameters of the trained image recognition model are fixed and used as a pre-trained image recognition model for the online monitoring phase.

[0036] S2.2. Input the spatial location of the vehicle edge contour into the cross-modal attention layer of the pre-trained image recognition model to modulate the response intensity of the pre-trained image recognition model to the blind zone features indicated by the radio frequency semantic vector.

[0037] Furthermore, the joint embedding features of the spatial location of the vehicle edge contour are input into the cross-modal attention layer of the pre-trained image recognition model. The cross-modal attention layer uses the joint embedding features of the spatial location of the vehicle edge contour as the query vector and the radio frequency semantic vector as the key-value pair to obtain attention weights and modulate the response intensity of the pre-trained image recognition model to the blind area features indicated by the radio frequency semantic vector, thereby enhancing the model's attention to the blind area features.

[0038] Specifically, the cross-modal attention layer uses the spatial location of the vehicle edge contour to jointly embed features to guide the allocation of attention weights, enabling the model to focus on the blind zone indicated by the radio frequency semantic vector, avoiding interference from irrelevant features. The modulated response intensity accurately reflects the importance of blind zone features, improving the model's sensitivity to blind zone targets, overcoming the defect of scattered feature responses in traditional cross-modal fusion, and providing an accurate feature response basis for occluded region contour completion.

[0039] S2.3. Convert the vehicle edge contours within the visible area into a spatial position mask matrix. Constrain the feature map response range of the pre-trained image recognition model by element-wise multiplication to obtain the mask-constrained pre-trained image recognition model.

[0040] Furthermore, the vehicle edge contours within the visible area are converted into a spatial location mask matrix. The spatial location mask matrix assigns a value of 1 to the visible area and a value of 0 to the occluded area at the feature map scale. The spatial location mask matrix is ​​multiplied with the feature map of the pre-trained image recognition model through element-wise multiplication, thus constraining the feature map response range of the pre-trained image recognition model to be limited to the visible area, resulting in a mask-constrained pre-trained image recognition model.

[0041] Specifically, the spatial location mask matrix precisely constrains the response range of the feature map through element-wise multiplication, avoiding invalid feature extraction of the occluded area by the model. The pre-trained image recognition model after mask constraint focuses on the effective features within the visible area, reducing computational redundancy and improving feature extraction efficiency, while ensuring that feature completion of the occluded area is based on reliable contour information of the visible area.

[0042] S2.4 Construct an adaptive probability extension decoder based on the physical attributes indicated by the radio frequency semantic vector. The adaptive probability extension decoder learns the mapping relationship between large metal targets and rectangular probability distributions, and between pedestrians and cylindrical probability distributions during the training phase.

[0043] Furthermore, an adaptive probability extension decoder is constructed based on the physical properties indicated by the radio frequency semantic vector. During the training phase, the adaptive probability extension decoder learns the mapping relationship between large metal targets and rectangular probability distributions, and between pedestrians and cylindrical probability distributions. The decoder parameters are adjusted through supervised learning so that the decoder can output the corresponding probability distribution shape according to the physical properties of the input radio frequency semantic vector.

[0044] Specifically, the adaptive probabilistic extension decoder learns specific probability distribution mappings for different physical properties. The setting of rectangular probability distribution for large metal targets and cylindrical probability distribution for pedestrians conforms to the physical shape characteristics of the targets, making the contour completion of occluded areas closer to the shape of the real targets. This overcomes the generalization problem of traditional decoders using a uniform shape for all targets, improves the physical rationality of the completed contours, and provides shape guidance that conforms to physical properties for contour completion of occluded areas.

[0045] S2.5. The contour of the occluded area is completed by an adaptive probabilistic extension decoder to obtain the completed occluded area contour. The visible area image and the completed occluded area contour are fused together. The spatial position, velocity and physical attributes of each target are marked in the fusion result to generate a holographic traffic situation map.

[0046] Furthermore, the contour of the occluded area is completed by an adaptive probabilistic extension decoder. The decoder outputs the corresponding probability distribution shape according to the physical attributes indicated by the radio frequency semantic vector, generates contour candidates that match the physical attributes in the occluded area, and obtains the completed occluded area contour. The visible area image and the completed occluded area contour are fused, and the spatial position, velocity and physical attributes of each target are marked in the fusion result to generate a holographic traffic situation map.

[0047] Specifically, each target refers to vehicles, pedestrians, and infrastructure in the visible area of ​​the urban road monitoring area, as well as dynamic and static objects in the blind area obtained through radio frequency semantic vector inversion, which have spatial location, speed, and physical attribute information. The adaptive probabilistic extension decoder outputs the matching contour shape based on the physical attributes. The contour of the occluded area after completion is consistent with the physical attributes of the target. After fusing the visible area image and the contour of the occluded area after completion, the holographic traffic situation map contains the precise contour of the visible area and the physical reasonable contour of the blind area. The labeled spatial location, speed, and physical attribute information comprehensively reflect the traffic situation.

[0048] S3. Based on the holographic traffic situation map, extract the image feature vector sequence and combine it with the radio frequency semantic vector sequence to construct a time elasticity metric space and align the sequence. Calculate the geodesic distance of the aligned sequence in the high-dimensional topological manifold space and determine the spatiotemporal topological mismatch state based on the geodesic distance.

[0049] S3.1 Extract the image feature vector sequence carrying the target's physical attributes from the holographic traffic situation map, and extract the radio frequency semantic vector sequence that matches the target's physical attributes from the radio frequency semantic vector.

[0050] Furthermore, image feature vector sequences carrying target physical attributes are extracted from the holographic traffic situation map. These image feature vector sequences contain information on the target's spatial location, speed, and physical attributes. Radio frequency semantic vector sequences matching the target's physical attributes are extracted from the radio frequency semantic vectors. These radio frequency semantic vector sequences contain blind zone target reflection features corresponding to the target's physical attributes. The matching process is based on the consistency of the target's physical attribute category.

[0051] Specifically, the image feature vector sequence and radio frequency semantic vector sequence are screened by matching the consistency of the target physical attributes. This avoids the alignment error caused by the misalignment of physical attributes in traditional sequence matching. The image feature vector sequence carrying the target physical attributes and the matched radio frequency semantic vector sequence provide input with consistent physical attributes for the subsequent construction of the temporal elasticity metric space. This ensures that the temporal alignment process is based on target features with the same physical attributes, thereby improving alignment accuracy.

[0052] S3.2. Construct a time elasticity metric space based on the physical kinematic constraints corresponding to the target physical properties, and use a dynamic time warping algorithm constrained by physical kinematic constraints to perform nonlinear time alignment between the matched image feature vector sequence and the radio frequency semantic vector sequence.

[0053] Furthermore, a time elasticity metric space is constructed based on the physical kinematic constraints corresponding to the target's physical properties. The physical kinematic constraints limit the target's maximum acceleration and velocity variation range. A dynamic time warping algorithm constrained by physical kinematic constraints is used to perform nonlinear time alignment on the matched image feature vector sequence and the radio frequency semantic vector sequence, and the time warping path is dynamically adjusted to conform to the physical kinematic constraints.

[0054] Specifically, physical kinematic constraints limit the search path of the dynamic time warping algorithm, eliminate temporal matching that does not conform to the laws of physical motion, avoid temporal misalignment caused by the lack of constraints in the traditional dynamic time warping algorithm, and nonlinear time alignment makes the image feature vector sequence and the radio frequency semantic vector sequence conform to the actual motion law of the target in the time dimension.

[0055] S3.3 Map the aligned sequence to a high-dimensional topological manifold space. In the high-dimensional topological manifold space, adjust the integral path of the geodesic distance using the target physical properties as the intrinsic curvature parameter of the manifold. Calculate the geodesic distance between the adjusted image feature vector sequence and the radio frequency semantic vector sequence.

[0056] Furthermore, the aligned sequence is mapped to a high-dimensional topological manifold space. In the high-dimensional topological manifold space, the integral path of the geodesic distance is adjusted using the target physical properties as the intrinsic curvature parameters of the manifold. The target physical properties affect the direction of the geodesics by changing the local curvature of the manifold. The geodesic distance between the adjusted image feature vector sequence and the radio frequency semantic vector sequence is calculated.

[0057] Specifically, by using the target physical properties as the intrinsic curvature parameters of the manifold, the geometric structure of the high-dimensional topological manifold space is dynamically adjusted according to the target physical properties. The integral path of the geodesic distance is more in line with the spatial distribution characteristics corresponding to the target physical properties, overcoming the defect of traditional Euclidean distance that ignores the differences in physical semantics of feature dimensions, improving the sensitivity of geodesic distance to differences in physical properties, and accurately quantifying the physical consistency between sequences.

[0058] The expression for geodesic distance is: ; in, The aligned image feature vector sequence With radio frequency semantic vector sequence Geodesic distance in a high-dimensional topological manifold space For connections in high-dimensional topological manifold space and Geodesic path, For target physical properties, Geodesic path For parameters The derivative, It is a sequence of image feature vectors. It is a sequence of radio frequency semantic vectors. For a high-dimensional topological manifold space, the metric tensor is... Geodesic path The curve parameters.

[0059] S3.4. Based on the geodesic distance between the adjusted image feature vector sequence and the radio frequency semantic vector sequence, construct the causal constraints of physical motion, and determine the spatiotemporal topological mismatch state according to the causal constraints of physical motion.

[0060] Furthermore, based on the geodesic distance between the adjusted image feature vector sequence and the radio frequency semantic vector sequence, a causal constraint for physical motion is constructed. This constraint requires that the target's trajectory must conform to physical causality when the geodesic distance exceeds a threshold. The spatiotemporal topological mismatch state is determined based on this constraint. When the geodesic distance exceeds the threshold and does not conform to physical causality, it is determined to be a spatiotemporal topological mismatch. The causal constraint for physical motion combines geodesic distance and motion causal logic, avoiding misjudgments caused by the traditional threshold method relying solely on distance values. The determination of the spatiotemporal topological mismatch state considers both feature similarity and the rationality of physical motion, thus improving the accuracy of anomaly detection.

[0061] S4. Based on the spatiotemporal topological mismatch state, upload the data corresponding to the holographic traffic situation map to the cloud, and execute the corresponding traffic control operations according to the data corresponding to the holographic traffic situation map.

[0062] S4.1 Based on the spatiotemporal topological mismatch state of the judgment, the data corresponding to the holographic traffic situation map is classified and encapsulated according to the target physical attributes.

[0063] Furthermore, based on the determined spatiotemporal topological mismatch state, the data corresponding to the holographic traffic situation map is classified and encapsulated according to the target's physical attributes. Data carrying the physical attributes of large metallic targets and data carrying the physical attributes of pedestrians are encapsulated into independent data packets. The physical attributes of large metallic targets specifically refer to the classification label obtained by matching the physical feature fingerprint database in the radio frequency semantic vector, which characterizes the target material as metal and has a large volume or size, and its associated electromagnetic reflection feature parameters. The target physical attribute label and the spatiotemporal topological mismatch state identifier are attached to the header of the data packet. By classifying and encapsulating the data according to the target physical attributes, the data packets carry clear physical semantic information, improve the targeting of data transmission and the convenience of cloud processing, and provide a data foundation for differentiated transmission strategies.

[0064] S4.2. Based on the target physical attributes carried by the data corresponding to the encapsulated holographic traffic situation map, select the corresponding transmission channel and use the selected transmission channel as the carrier for uploading the data corresponding to the holographic traffic situation map.

[0065] Furthermore, data packets corresponding to the physical attributes of large metal targets are transmitted through high-bandwidth channels, while data packets corresponding to the physical attributes of pedestrians are transmitted through low-bandwidth channels. The selected transmission channels serve as the carriers for uploading data corresponding to the holographic traffic situation map. By selecting transmission channels based on the physical attributes of the targets, dynamic allocation of bandwidth resources is achieved. High-bandwidth channels transmit high-priority data, while low-bandwidth channels transmit low-priority data. This avoids bandwidth waste or congestion associated with traditional fixed channel allocation, improves transmission efficiency and resource utilization, ensures priority transmission of urgent and abnormal data, and reduces transmission latency.

[0066] S4.3 Upload the data corresponding to the encapsulated holographic traffic situation map to the cloud through the selected transmission channel.

[0067] Furthermore, the encapsulated holographic traffic situation map data is uploaded to the cloud via the selected transmission channel. During transmission, the target physical attribute label and spatiotemporal topology mismatch status identifier attached to the data packet header are kept intact to ensure that the data received by the cloud contains complete physical semantics and abnormal status information. Uploading data through a transmission channel with matching physical attributes ensures the reliability and timeliness of data transmission, avoids network congestion or loss of critical data caused by traditional indiscriminate uploading, and ensures that the cloud obtains accurate data corresponding to the holographic traffic situation map in a timely manner, providing real-time data support for subsequent traffic control decisions.

[0068] S4.4. Analyze the data uploaded to the cloud to distinguish between structured semantic information and abnormal data packets, and match the corresponding traffic control plan based on the analysis results.

[0069] Furthermore, the data uploaded to the cloud is parsed to distinguish between structured semantic information and abnormal data packets. The structured semantic information includes the target's spatial location, velocity, and physical attributes, while the abnormal data packets contain spatiotemporal topology mismatch status identifiers. Based on the parsing results, corresponding traffic control plans are matched. The spatiotemporal topology mismatch status identifier corresponds to an abnormal traffic control plan, while the structured semantic information corresponds to a regular traffic control plan. By parsing the data to distinguish between structured semantic information and abnormal data packets, accurate matching of traffic control plans is achieved, avoiding plan matching errors caused by traditional mixed parsing. This improves the accuracy and response speed of traffic control decisions, ensures that abnormal events are handled in a timely manner, and that regular traffic flows are rationally diverted.

[0070] S4.5. Generate handling instructions based on the matched traffic control plan, and send the handling instructions to the execution unit via the Internet of Things to carry out traffic control operations. Furthermore, based on the matched traffic control plan, disposal instructions are generated. These instructions include control measures corresponding to the target's spatial location, speed, and physical attributes. The disposal instructions are then sent to the execution unit via the Internet of Things (IoT). The execution unit adjusts traffic signal timings, issues road condition warnings, or intercepts specific targets according to the disposal instructions, thus performing traffic control operations. Targeted disposal instructions are generated based on the traffic control plan and sent to the execution unit via the IoT to achieve precise control of traffic flow. This avoids the poor control effect of traditional general instructions, improves the flexibility and effectiveness of traffic control, ensures the rapid recovery of traffic stability, and reduces the impact of abnormal events on traffic flow.

[0071] This embodiment also provides an online urban road monitoring system based on image recognition and the Internet of Things, including: The extraction module collects real-time visible light video streams and radio frequency sensing data from urban road monitoring areas, performs multipath physical modeling on the radio frequency sensing data, extracts the pure reflected wave features caused by dynamic targets, and converts the pure reflected wave features into radio frequency semantic vectors of the physical properties of targets in the blind zone. The completion module inputs the radio frequency semantic vector into the pre-trained image recognition model, and combines the vehicle edge contours within the visible area to complete the contours of the occluded area, generating a holographic traffic situation map. The alignment module extracts image feature vector sequences from holographic traffic situation maps and combines them with radio frequency semantic vector sequences to construct a temporal elasticity metric space and align the sequences. It then calculates the geodesic distance of the aligned sequences in a high-dimensional topological manifold space and determines the spatiotemporal topological mismatch state based on the geodesic distance. The execution module, based on the determined spatiotemporal topological mismatch state, uploads the data corresponding to the holographic traffic situation map to the cloud, and performs corresponding traffic control operations based on the data corresponding to the holographic traffic situation map.

[0072] This embodiment also provides a computer device applicable to the online monitoring method for urban roads based on image recognition and the Internet of Things, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the online monitoring method for urban roads based on image recognition and the Internet of Things as proposed in the above embodiment.

[0073] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0074] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the online urban road monitoring method based on image recognition and the Internet of Things as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0075] In summary, this invention extracts pure reflected wave features through multipath physical modeling and Riemannian geometric phase unwrapping, converting them into radio frequency semantic vectors containing position coordinates and physical properties. This solves the problem of inaccurate inversion of target attributes in blind spots in non-line-of-sight scenarios, providing strong physical-level prior constraints for visual completion. By constructing a temporal elasticity metric space and adjusting the geodesic integral path using target physical properties as manifold curvature parameters, it achieves nonlinear alignment and physical reliability quantification of images and radio frequency sequences, overcoming the spatiotemporal misalignment caused by rigid synchronization. Combined with dynamic data upload based on physical motion causality verification, this invention reduces IoT bandwidth consumption while improving the perception accuracy of holographic traffic situation maps and the execution efficiency of traffic control commands.

[0076] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for online monitoring of urban roads based on image recognition and the Internet of Things, characterized in that: include, Real-time visible light video streams and radio frequency sensing data of urban road monitoring areas are collected. Multipath physical modeling is performed on the radio frequency sensing data, and the pure reflected wave features caused by dynamic targets are extracted. The pure reflected wave features are converted into radio frequency semantic vectors of the physical properties of targets in the blind zone. Radio frequency semantic vectors are input into a pre-trained image recognition model, and the contours of the occluded areas are completed by combining the vehicle edge contours within the visible area to generate a holographic traffic situation map. Based on the holographic traffic situation map, image feature vector sequences are extracted and combined with radio frequency semantic vector sequences to construct a temporal elasticity metric space and align the sequences. The geodesic distance of the aligned sequences in the high-dimensional topological manifold space is calculated, and the spatiotemporal topological mismatch state is determined based on the geodesic distance. Based on the determined spatiotemporal topological mismatch state, the data corresponding to the holographic traffic situation map is uploaded to the cloud, and corresponding traffic control operations are performed according to the data corresponding to the holographic traffic situation map.

2. The urban road online monitoring method based on image recognition and the Internet of Things as described in claim 1, characterized in that: The radio frequency semantic vector of the physical attributes of the blind zone target includes, Using prior environmental knowledge, a multipath physical modeling based on propagation path decomposition is performed on the electromagnetic propagation channel response corresponding to real-time visible light video streams and radio frequency sensing data. Time-varying topological manifold filtering is used to extract the time-varying scattering component caused by the dynamic target by nonlinear manifold unwrapping of the dynamic micro-Doppler component in the phase space, generating multipath physical modeling results. Fractional Fourier transform is performed on the multipath physics modeling results to separate the linear frequency modulated signals of different scatterers on the rotating time-frequency plane. The spatial topology of the reflected wave is reconstructed based on the correlation matrix of the scattering points to remove the residual fixed clutter and extract the features of the pure reflected wave caused by the dynamic target. Phase difference sequences are extracted by Riemannian geometric phase unwrapping of pure reflected wave characteristics, and a physical feature fingerprint database is established by combining the dielectric constant response characteristics of electromagnetic waves in different materials. By matching geodesics on the manifold, the phase difference sequence is mapped to the target's position coordinates and physical attribute classification labels, and the pure reflected wave features are converted into radio frequency semantic vectors of the target's physical attributes in the blind zone.

3. The urban road online monitoring method based on image recognition and the Internet of Things as described in claim 2, characterized in that: The holographic traffic situation map includes, The physical attribute labels of the radio frequency semantic vector are mapped to the joint embedding features of the spatial location of the vehicle edge contour within the visible area by using the physical constraint loss function. The spatial location of the vehicle edge contour is jointly embedded into the cross-modal attention layer of the pre-trained image recognition model, thereby modulating the response strength of the pre-trained image recognition model to the blind zone region features indicated by the radio frequency semantic vector. The vehicle edge contours within the visible area are converted into a spatial location mask matrix. The feature map response range of the pre-trained image recognition model is constrained by element-wise multiplication to obtain the mask-constrained pre-trained image recognition model. An adaptive probability extension decoder is constructed based on the physical properties indicated by the radio frequency semantic vector. During the training phase, the adaptive probability extension decoder learns the mapping relationship between large metal targets and rectangular probability distributions, and between pedestrians and cylindrical probability distributions. The contour of the occluded area is completed by an adaptive probabilistic extension decoder, and the completed occluded area contour is obtained. The visible area image and the completed occluded area contour are fused together, and the spatial position, velocity and physical attributes of each target are marked in the fusion result to generate a holographic traffic situation map.

4. The urban road online monitoring method based on image recognition and the Internet of Things as described in claim 3, characterized in that: The geodesic distance includes, Extract image feature vector sequences carrying target physical attributes from holographic traffic situation maps, and extract radio frequency semantic vector sequences that match the target physical attributes from radio frequency semantic vectors; A time elasticity metric space is constructed based on the physical kinematic constraints corresponding to the target's physical properties. A dynamic time warping algorithm constrained by physical kinematic constraints is used to perform nonlinear time alignment between the matched image feature vector sequence and the radio frequency semantic vector sequence. The aligned sequence is mapped to a high-dimensional topological manifold space. In the high-dimensional topological manifold space, the integral path of the geodesic distance is adjusted using the target physical properties as the intrinsic curvature parameter of the manifold. The geodesic distance between the adjusted image feature vector sequence and the radio frequency semantic vector sequence is calculated.

5. The urban road online monitoring method based on image recognition and the Internet of Things as described in claim 4, characterized in that: The determination of spatiotemporal topological mismatch includes... Based on the geodesic distance between the adjusted image feature vector sequence and the radio frequency semantic vector sequence, a causal constraint of physical motion is constructed, and the spatiotemporal topological mismatch state is determined according to the causal constraint of physical motion.

6. The urban road online monitoring method based on image recognition and the Internet of Things as described in claim 5, characterized in that: Uploading the data corresponding to the holographic traffic situation map to the cloud includes, Based on the spatiotemporal topological mismatch state determined by the judgment, the data corresponding to the holographic traffic situation map is classified and encapsulated according to the target's physical attributes. Based on the target physical attributes carried by the data corresponding to the encapsulated holographic traffic situation map, select the corresponding transmission channel and use the selected transmission channel as the carrier for uploading the data corresponding to the holographic traffic situation map. The data corresponding to the encapsulated holographic traffic situation map is uploaded to the cloud through the selected transmission channel.

7. The urban road online monitoring method based on image recognition and the Internet of Things as described in claim 6, characterized in that: The execution of the corresponding traffic control operations includes, The data uploaded to the cloud is parsed to distinguish between structured semantic information and abnormal data packets, and the corresponding traffic control plan is matched based on the parsing results; Based on the matched traffic control plan, a handling instruction is generated and sent to the execution unit via the Internet of Things to carry out traffic control operations.

8. An online urban road monitoring system based on image recognition and the Internet of Things, based on the online urban road monitoring method based on image recognition and the Internet of Things as described in any one of claims 1 to 7, characterized in that: include, The extraction module collects real-time visible light video streams and radio frequency sensing data from urban road monitoring areas, performs multipath physical modeling on the radio frequency sensing data, extracts the pure reflected wave features caused by dynamic targets, and converts the pure reflected wave features into radio frequency semantic vectors of the physical properties of targets in the blind zone. The completion module inputs the radio frequency semantic vector into the pre-trained image recognition model, and combines the vehicle edge contours within the visible area to complete the contours of the occluded area, generating a holographic traffic situation map. The alignment module extracts image feature vector sequences from holographic traffic situation maps and combines them with radio frequency semantic vector sequences to construct a temporal elasticity metric space and align the sequences. It then calculates the geodesic distance of the aligned sequences in a high-dimensional topological manifold space and determines the spatiotemporal topological mismatch state based on the geodesic distance. The execution module, based on the determined spatiotemporal topological mismatch state, uploads the data corresponding to the holographic traffic situation map to the cloud, and performs corresponding traffic control operations based on the data corresponding to the holographic traffic situation map.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the urban road online monitoring method based on image recognition and the Internet of Things as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the urban road online monitoring method based on image recognition and the Internet of Things as described in any one of claims 1 to 7.