Intelligent video system application method and system suitable for smart city

By constructing a unified time reference and model version view in the smart city video system, the problem of lack of model version and operating environment labeling in algorithm results is solved, realizing the temporal consistency and quality traceability of video events, and supporting the quantitative evaluation of model upgrade effects.

CN121982633APending Publication Date: 2026-05-05ZHONGRUI COMM PLANNING & DESIGN
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHONGRUI COMM PLANNING & DESIGN
Filing Date
2025-12-31
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

In existing smart city video systems, algorithm results lack model version and operating environment labels, making it difficult to achieve traceable and comparative evaluation of effects, difficult to reconstruct events across cameras, difficult to objectively compare the effects of model upgrades, and difficult to quickly locate the causes of false alarms and missed alarms.

Method used

By constructing urban video acquisition data frames, performing field-of-view geometric modeling and regional coding, establishing a unified time reference, performing edge-side preprocessing and image quality measurement, constructing a model version view and a traceable evaluation index set, realizing an algorithm running data frame set, and constructing a unified detection result service view and a traceable backtracking interface.

Benefits of technology

It significantly reduces the risk of temporal misalignment in the reconstruction process of multi-source video events, ensures the consistency of the timeline of monitoring images across devices and regions, realizes continuous quantification and hierarchical labeling of channel quality, supports quantitative evaluation of model upgrade effects, and provides traceable detection result analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121982633A_ABST
    Figure CN121982633A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent video system application method and system suitable for a smart city, and relates to the technical field of visual inspection. The invention discloses an intelligent video system application method and system suitable for a smart city. The method comprises the following steps: S1, collecting multi-source video frame data to construct a city video collection data frame; s2, carrying out edge side preprocessing and picture quality measurement, and outputting a video detection input set; s3, constructing an algorithm operation data frame and a traceable evaluation index set; and S4, constructing a unified detection result service view and an original scene reconstruction interface. According to the method, the accuracy of target detection in a multi-source city video scene and the adaptability to low-quality and degraded channel pictures are effectively improved, and the problems that an algorithm result lacks model version and operation environment marks, and effect traceability and comparison evaluation are difficult to achieve are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of visual inspection technology, specifically to an application method and system for intelligent video systems suitable for smart cities. Background Technology

[0002] With the continuous expansion of smart city video surveillance and the increasing frequency of visual algorithm iterations, existing technologies have exposed problems in multi-source video acquisition, algorithm deployment, and effect evaluation, such as inconsistent time bases, missing spatiotemporal indexes, untraceable algorithm versions and operating environments, and difficulty in quantifying the impact of channel image quality degradation on detection results. A large number of front-end cameras and edge gateways are provided by different manufacturers and independently maintain their own local clocks. Uploaded data only carries a single device time field, while the cloud side uses database entry time as the sorting criterion, leading to inconsistencies in the time order of the same event across multiple tables and channels.

[0003] For example, invention patent CN109271554B discloses an intelligent video recognition system and its application, including: a front-end access device, a video image intelligent analysis system, a video big data analysis system, a video cloud platform, and a comprehensive application system. This invention enables networking of existing cameras, unified application of integrated resources, broad coverage, rich functionality, a large number of monitorable features, and comprehensive monitoring elements; simultaneously, it builds video surveillance front-ends based on the internet, offering high speed and low cost; and it selects high-value locations to deploy intelligent front-ends, enabling the capture and alarm control of people and vehicles.

[0004] For example, invention patent CN103475870B discloses a distributed, scalable intelligent video surveillance system, including: a decoding management module, a polling scheduling module, a combined analysis module, a resource management module, a screen display module, an alarm push module, and an external communication module. This invention addresses the new characteristics of modern video surveillance systems—digitalization, intelligence, and distribution—and through flexible configuration, can simultaneously meet the application needs of both small-scale and large-scale intelligent video surveillance systems, while also possessing smooth scalability.

[0005] In existing technologies, existing systems typically record detection results using only a single timestamp and device number. They lack a spatiotemporal index view of urban video acquisition based on a unified time base, lack joint modeling of channel quality health index and detection process, and lack a linkage mechanism between model version view and operating environment view with algorithm running data frames as the core. This leads to difficulties in reconstructing cross-camera events, difficulty in objectively comparing the effects of model upgrades, and difficulty in quickly locating the causes of false alarms and missed alarms.

[0006] Therefore, in response to the above problems, there is an urgent need for an intelligent video system application method and system suitable for smart cities. Summary of the Invention

[0007] Technical problems to be solved

[0008] To address the shortcomings of existing technologies, this invention provides an application method and system for intelligent video systems suitable for smart cities, solving the problem that algorithm results lack model version and operating environment markings, making it difficult to achieve traceable and comparative evaluation of effects.

[0009] Technical solution

[0010] To achieve the above objectives, the present invention provides the following technical solution: an application method and system for an intelligent video system applicable to smart cities, comprising: S1, constructing urban video acquisition data frames from multi-source video frame data, performing field-of-view geometric modeling and region coding to construct a spatiotemporal index view of urban video acquisition; S2, performing edge-side preprocessing and image quality measurement on the urban video acquisition data frame sequence, and performing channel quality time-series evaluation and quality-driven frame scheduling to output a video detection input set; S3, constructing an algorithm running data frame set based on the video detection input set, and constructing a model version view and a traceable evaluation index set; S4, constructing a unified detection result service view based on the algorithm running data frame set, and constructing a traceable backtracking and original scene reconstruction interface.

[0011] Furthermore, the specific process of constructing urban video acquisition data frames from the collected multi-source video frame data is as follows: Input multi-source video frame data collected by fixed cameras, PTZ cameras, pan-tilt cameras, vehicle-mounted terminals, ship-mounted terminals, and cameras mounted on drones; construct a unified time base and schedule channel sampling for the multi-source video frame data to form an urban video acquisition data frame sequence; during the acquisition process, a unified time source calibrated by a unified time service is used as the reference, and a unified sampling timestamp is assigned to each video sample; sampling records from different video terminals are sorted according to the sampling timestamp and assigned to the corresponding sampling period; the sampling timestamps are normalized using a unified time unit and starting reference point, and the local time of different devices is mapped to a relative time index under a unified time base; when generating a unified acquisition frame sequence, the sampling timestamp assigned by the unified time source is used as the main time axis, and the video sampling records of each channel on the same time axis are time-aligned, outputting an urban video acquisition data frame sequence arranged in a unified sampling time order.

[0012] Furthermore, the specific process of constructing a spatiotemporal index view for urban video acquisition through field-of-view geometric modeling and regional coding is as follows: Inputting a basic urban geographic information dataset consisting of a base topographic map, road and building vector data, administrative division and functional zoning data, and a sequence of urban video acquisition data frames; performing geometric modeling and regional coding on the installation location and field of view of each video terminal; mapping the urban video acquisition data frame sequence to a unified spatiotemporal coordinate system; and constructing a spatiotemporal index view for urban video acquisition. This involves mapping the installation point of each fixed camera to the unified urban coordinate system through coordinate transformation and querying the urban... The basic geographic information dataset centralizes the administrative division boundaries, road centerlines, and park boundaries that intersect with the installation point and surrounding areas. The administrative division code and road or park identifier to which the installation point belongs are used as regional codes to establish a one-to-one or one-to-many mapping relationship, forming a time-varying set of device trajectory points. The regional codes and time series are bound to the frame-level globally unique identifier, device unique identifier, channel identifier, and unified analysis time in the urban video acquisition data frame sequence, generating a frame-level spatiotemporal association record for each frame of urban video acquisition data frame sequence and writing it into the video spatiotemporal index table. The urban video acquisition spatiotemporal index view is then output.

[0013] Further, the specific process of edge-side preprocessing and image quality measurement of the urban video acquisition data frame sequence is as follows: Input the urban video acquisition data frame sequence and the urban video acquisition spatiotemporal index view; perform edge-side preprocessing and image quality measurement on the urban video acquisition data frame sequence to form a frame-level quality feature vector sequence; in the edge computing node, configure a circular frame buffer queue for each path, and write the urban video acquisition data frame sequence into the buffer according to a unified analysis time order, so that the buffer always retains the most recent continuous frame sequence; perform frame-level image quality feature vector operation on each frame: select multiple image quality evaluation operators to analyze the image, combine them in order to form a frame-level quality feature vector, and record the physical meaning, value range, and suggested threshold for each dimension feature in the quality feature dictionary; based on the frame-level quality feature vector, obtain the quality score and quality level label for each frame; write the frame-level quality feature vector, quality score, and quality level label into the frame-level quality feature table to form a city video quality feature vector sequence arranged in time order; output the frame-level quality feature vector sequence and the quality level label set.

[0014] Further, the specific process of performing channel quality time series evaluation and quality-driven frame scheduling to output the video detection input set is as follows: Input the frame-level quality feature vector sequence, the urban video acquisition data frame sequence, and the urban video acquisition spatiotemporal index view. Perform channel quality time series evaluation on the image quality changes of multi-source video channels in the time dimension, and execute quality-driven frame filtering and scheduling accordingly, reconstructing a quality-aware video detection input set. Using the device and channel combination identifier (device unique identifier and channel identifier) ​​as channel indexes, sort the frame-level quality records belonging to the same channel according to a unified analysis time, constructing a channel quality time series, and statistically analyzing quality features within a sliding time window. Construct a channel quality health index, and while generating the channel quality health index, write the channel quality status and change process into the channel quality status table and quality anomaly log table. Based on the channel quality status and frame-level quality level, construct a quality-driven video detection frame scheduling strategy: output a quality-aware video detection input set.

[0015] Furthermore, the specific process of constructing the algorithm execution data frame set based on the video detection input set is as follows: Input the video detection input set, uniformly collect and structurally encapsulate the detection results, model configuration, and runtime environment characteristics generated during the execution of each detection task to construct the algorithm execution data frame set; Based on the task entries in the video detection input set, select the model version and runtime configuration that match the task type, scene category, and quality conditions; On the inference node side, locate and load the corresponding preprocessed image frame in the city video acquisition data frame table based on the frame-level globally unique identifier, input the image into the specified model version for inference, and structurally encapsulate the detection results output by the model to obtain the basic record of the detection results; During the inference execution process, collect the runtime feature information of this inference call in real time, quantify the inference environment and resource consumption into runtime feature vectors; Construct the algorithm execution data frame, write the algorithm execution data frame record into the algorithm execution data frame table, and output the algorithm execution data frame set.

[0016] Furthermore, the specific process of constructing the model version view and traceable evaluation index set is as follows: Input the set of algorithm running data frames, the sequence of urban video acquisition data frames, the spatiotemporal index view of urban video acquisition, the sequence of frame-level quality feature vectors, and the channel quality status record. Aggregate and compare the detection results and running characteristics of different model versions on a unified data basis to construct a model version view and traceable evaluation index set for version management: Using model identifier and model version number as the main dimensions, group and aggregate the set of algorithm running data frames to construct the model version view; Statistically analyze samples under different quality levels and different channel quality states to distinguish the performance of the same model version under ideal, general, and degraded image quality conditions. For scenarios with multiple model versions configured for parallel detection, use the same frame-level globally unique identifier as the association key to compare the detection results given by multiple model versions for the same frame and the same target area one-to-one to construct a cross-version comparison sub-view; Output the model version view and evaluation index set.

[0017] Furthermore, the specific process of constructing a unified detection result service view based on the algorithm running data frame set is as follows: Input the algorithm running data frame set, model version view and evaluation index set, as well as the urban video acquisition data frame sequence, urban video acquisition spatiotemporal index view, frame-level quality feature vector sequence and channel quality status record. Perform field pruning, semantic merging and service encapsulation on the multi-source algorithm running data frames to construct a unified detection result service view: Based on the algorithm running data frame set, perform hierarchical extraction and semantic merging of fields, and structurally decouple the internally used running process fields from the externally served detection result fields: In the outer service view, only the technical fields directly related to the upper-layer system are retained, and the version health information and recommendation level from the model version view and evaluation index set are injected into each detection result record, and the unified detection result service view is output.

[0018] Furthermore, the specific process of constructing the traceable backtracking and original scene reconstruction interface is as follows: Input a unified detection result service view and an algorithm running data frame set; organize and index the correlation between detection results, model version, image quality, running environment, and the original video acquisition process; construct a traceable backtracking interface that supports combined queries and original scene reconstruction; design a combined condition query mechanism: users can specify model version, time interval, area range, image quality conditions, channel quality status characteristics, and running anomaly marker conditions; under the support of multi-level indexing in the service view, quickly filter matching detection result records to obtain a candidate result set; after obtaining the candidate detection result records and the corresponding running data frame identifiers, reconstruct the original technical context layer by layer through the source index: using the running data frame identifier as the key, restore the model identifier and model version number used in this detection; using the frame identifier and the globally unique frame-level identifier as the key, obtain the unified analysis time, spatial coordinate position, and field of view coverage area of ​​the frame; restore the image quality change and channel health index time curve within the time window before and after the frame, providing a unified traceable backtracking interface for the upper-layer system.

[0019] Furthermore, a second aspect of the present invention provides an intelligent video system application system applicable to smart cities, and an application method for an intelligent video system applicable to smart cities, comprising: a video acquisition module, used to acquire multi-source video frame data to construct urban video acquisition data frames, perform field-of-view geometric modeling and region coding to construct a spatiotemporal index view of urban video acquisition; a quality perception module, used to perform edge-side preprocessing and image quality measurement on the urban video acquisition data frame sequence, and perform channel quality time series evaluation and quality-driven frame scheduling to output a video detection input set; a visual detection module, used to construct an algorithm running data frame set based on the video detection input set, construct a model version view and a traceable evaluation index set; and a view and traceability evaluation module, used to construct a unified detection result service view based on the algorithm running data frame set, and construct a traceable backtracking and original scene reconstruction interface.

[0020] Beneficial effects

[0021] The present invention has the following beneficial effects: (1) This invention estimates and corrects the time deviation of multi-source video terminals in the acquisition link online, constructs a unified analysis time, significantly reduces the risk of time sequence misalignment of multi-source video events in the reconstruction process, ensures the physical consistency of monitoring images across devices and regions on the time axis, and realizes combined retrieval by time interval, spatial region and channel identifier.

[0022] (2) This invention constructs a time-series curve of the channel quality health index as a function of time by jointly modeling the frame-level image quality features and the channel operation status, and continuously quantifies and classifies the channel image jitter, blur, occlusion and noise problems, so as to realize the automatic identification and trend warning of the monitoring channel and provide a quantitative basis for model evaluation and operation and maintenance decision-making.

[0023] (3) This invention constructs a model version view and a quality stratification evaluation index set based on the algorithm running data frame, and respectively counts the hit rate, false negative rate and running stability index of different model versions on different quality level screens such as A, B, C, and D. It also supports comparative analysis of multiple versions on the same video sample, thereby realizing that the model upgrade effect can be quantified and the robustness of low quality screens can be verified, avoiding model replacement based solely on experience.

[0024] (4) This invention constructs a unified detection result service view, which outputs only simplified fields such as target category, location, quality level, model version summary and traceability identifier related to the upper-level system. At the same time, it retains the foreign key relationship between the algorithm running data frame, video acquisition frame, quality features and running environment internally, so that the upper-level business system can consume the detection results in a light manner and return to the complete technical context with one click based on the traceability identifier when needed.

[0025] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0026] Figure 1 This is a flowchart of an application method for an intelligent video system suitable for smart cities according to the present invention; Figure 2 This is a system framework diagram of an intelligent video system applicable to smart cities according to the present invention; Figure 3 This is a time-series trend chart of the channel quality health index of the present invention; Figure 4 This is a comparison chart of the detection accuracy of multiple model versions of the present invention; Figure 5 The algorithm of this invention is used to generate an ER graph showing the relationships between data frame tables. Detailed Implementation

[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] Please see Figures 1-5 This invention provides a technical solution: an application method and system for an intelligent video system applicable to smart cities, comprising: S1, constructing urban video acquisition data frames from multi-source video frame data, performing field-of-view geometric modeling and region coding to construct a spatiotemporal index view of urban video acquisition; S2, performing edge-side preprocessing and image quality measurement on the urban video acquisition data frame sequence, and performing channel quality time-series evaluation and quality-driven frame scheduling to output a video detection input set; S3, constructing an algorithm running data frame set based on the video detection input set, and constructing a model version view and a traceable evaluation index set; S4, constructing a unified detection result service view based on the algorithm running data frame set, and constructing a traceable backtracking and original scene reconstruction interface.

[0029] Specifically, the process of constructing urban video acquisition data frames from collected multi-source video frame data is as follows: Inputting multi-source video frame data collected by fixed cameras, PTZ cameras, pan-tilt cameras, vehicle-mounted terminals, ship-mounted terminals, and drone-mounted cameras; constructing a unified time base and scheduling channel sampling for the multi-source video frame data to form an urban video acquisition data frame sequence; the input fixed camera video dataset includes: continuous image frames and their bitstream parameters collected by surveillance cameras installed at fixed locations such as road intersections, park entrances / exits, and building entrances at a set frame rate, used to characterize the temporal changes of people, vehicles, and objects in a fixed scene; the input PTZ camera and pan-tilt camera video dataset includes: image frame sequences collected by cameras with pan-tilt rotation capabilities under preset position cruise or manual control, as well as corresponding pan-tilt azimuth, pitch, and zoom state operation data, used to characterize the temporal trajectory of the visible area; the input vehicle-mounted terminal and ship-mounted video dataset includes: image frame sequences output by video acquisition terminals installed on vehicles or ships, along with their accompanying odometer, speed, and... Directional running status parameters are used to provide dynamic information about road or water scenes from a moving perspective. The input UAV video dataset includes: image frame sequences collected by the UAV-mounted camera during the execution of the mission route, along with their corresponding flight altitude, attitude angles, and positioning coordinates, used to provide scene observation under high-altitude views or specific inspection trajectories. To ensure that the upper-layer algorithm can trace the source of video data, a video resource description table is constructed in the access gateway or edge node. For each video terminal, the terminal type, manufacturer model, access protocol type, physical installation location or initial deployment location, field of view direction, recommended sampling frame rate, and encoding format are recorded. During the initialization phase, the corresponding protocol adaptation module is automatically loaded according to the video resource description table. The protocol adaptation module uniformly encapsulates the underlying access logic of RTSP, ONVIF, GB, T28181, or the manufacturer's private SDK. After successfully reading each frame of video output, the original bitstream is decoded into standardized image frames, and device identifiers, channel identifiers, and initial data quality markers are added to form a single-channel video sampling record with a clear source identifier.

[0030] During the acquisition process, a unified time source calibrated by a unified time service is used as the benchmark. A unified sampling timestamp is assigned to each frame of video sampling. The unified time source is synchronized with the upper-level standard clock via the NTP protocol. The time synchronization accuracy of each node is preferably controlled at the millisecond level, and time synchronization commands and deviation evaluation tasks are periodically issued by the unified time service. When the clock deviation between a certain access gateway or edge node and the unified time source exceeds the deviation threshold, forced time synchronization is triggered and a time anomaly mark is written into the acquisition record. When the deviation is within the safe range, a sliding time window is used to smoothly compensate for clock drift, avoiding timing jitter introduced by frequent time synchronization, thus forming a unified time benchmark system with accuracy constraints and fault tolerance strategies. Sampling records from different video terminals are sorted by sampling timestamp and assigned to the corresponding sampling period. A unified time is used for the sampling timestamps. Unit and starting reference point normalization maps the local time of different devices to a relative time index under a unified time base, eliminating the impact of clock deviations and time start differences between different devices on video alignment and event reconstruction. To address the issue of inconsistent sampling frame rates for different types of video channels, a multi-channel sampling scheduler is built in the access gateway or edge node: high frame rate sampling tasks are configured for cameras in key road sections and important locations, while medium frame rate sampling tasks are configured for cameras in ordinary areas. The sampling period for each type of sampling task is configured by the multi-channel sampling scheduler based on the nominal frame rate of the device, the time resolution requirements of the target detection and tracking tasks, and the current bandwidth and computing power constraints. The unified sampling time axis can be discretized according to the basic time step configured by the system, and the sampling period of each channel is expressed in integer multiples, which facilitates alignment and interpolation of different frame rate channels on the unified time axis.

[0031] During operation, vehicle-mounted and drone mobile terminals are configured with adaptive frame rate sampling tasks. The channel sampling scheduler triggers the execution of each sampling task's protocol adaptation interface under a unified time base. When generating a unified acquisition frame sequence, the sampling timestamp assigned by a unified time source is used as the main time axis. The video sampling records of each channel on the same time axis are time-aligned. For channels with short-term frame drops or jitter, a complete multi-channel frame index on the time axis is reconstructed by inserting missing markers, repeating the metadata of the previous frame, or generating placeholder frames. This ensures that each unified sampling moment corresponds to a set of video acquisition data frames with device identifiers, channel identifiers, multi-level timestamps, and quality markers. The output is a sequence of city video acquisition data frames arranged in the unified sampling time order.

[0032] This implementation plan effectively mitigates timing misalignment issues caused by clock drift, inconsistent frame rates, and short-term frame drops across different devices through collaborative management of terminal resource information, sampling frame rates, and a unified time base. This ensures the logical consistency and playability of multi-channel images at the same time. It not only significantly improves the input quality and stability of video preprocessing, object detection, and multi-source correlation analysis, but also enhances the traceability and maintainability of video data in large-scale smart city scenarios.

[0033] Specifically, the process of constructing a spatiotemporal index view for urban video acquisition through field-of-view geometric modeling and regional coding is as follows: Inputting an urban basic geographic information dataset consisting of a base topographic map, road and building vector data, administrative division and functional zoning data, and a sequence of urban video acquisition data frames; geometrically modeling and regional coding of the installation location and field of view of each video terminal; mapping the urban video acquisition data frame sequence to a unified spatiotemporal coordinate system; and constructing a spatiotemporal index view for urban video acquisition. The urban basic geographic information dataset is obtained during the deployment phase by importing vector data, raster base maps, and attribute tables from the existing urban GIS platform, including road centerlines, building... The system includes building outlines, polygon area boundaries, and administrative division codes. During data import, coordinate systems from different sources are uniformly converted, and topological corrections are performed on duplicate or non-closed geometric objects to form basic geographic data with a unified coordinate system and topological consistency. The equipment basic information table is generated based on the video resource description table. During equipment installation and acceptance, maintenance personnel supplement the data with installation latitude and longitude, installation height, orientation, elevation angle, and the administrative division, road number, or park number. The system merges and verifies the automatically discovered equipment basic attributes with the manually corrected location information to ensure that each channel corresponds to a unique physical installation location and initial field of view. The regional coding adopts a hierarchical coding structure of administrative division code, functional zone code, and road and park number: the smallest administrative division unit where the camera installation point is located is used as the regional code prefix, and when functional zone data exists, the functional zone code is added as an intermediate layer. Then, a suffix is ​​generated based on the road number or park number that is closest to the installation point and has the smallest angle with the field of view. When the same camera intersects with multiple roads or multiple parks at the same time, the road or park identifier corresponding to the main direction is used as the main regional code, and the other intersecting objects are recorded as auxiliary regional codes in the auxiliary field.

[0034] Based on the equipment's installation latitude and longitude, installation height, orientation, and tilt angle parameters, the installation point of each fixed camera is mapped to the city's unified coordinate system through coordinate transformation. The city's basic geographic information dataset is queried to find the administrative division boundaries, road centerlines, and park boundaries that intersect with the installation point and surrounding area. The administrative division code and road or park identifier of the installation point are used as regional codes to establish a one-to-one or one-to-many mapping relationship. For vehicle-mounted, ship-mounted, and UAV-mounted cameras, they are marked as mobile field-of-view devices in the acquisition task configuration. Combined with the odometer, GNSS coordinates, speed, and attitude information attached to the video acquisition data frame sequence, the spatial position record is updated in real time according to a unified sampling time during the acquisition process. The position of the moving terminal at each sampling moment is projected onto the unified coordinate system to form a time-varying set of equipment trajectory points.

[0035] Based on the intrinsic and extrinsic parameters of the camera, a field-of-view geometric model is constructed for each camera in a unified coordinate system. The camera imaging plane is mapped to a polygonal region on the ground or in 3D space through geometric back-projection. In a fixed camera scene, the installation point position, installation height, horizontal field of view angle, and pitch angle are used as inputs to derive the field-of-view coverage polygon on the ground. The polygon is encoded as a field-of-view region identifier, and the inclusion or intersection relationship with the region code is recorded. For mobile field-of-view devices, the current device position and attitude parameters are taken according to discrete unified sampling time within each time window. The above back-projection calculation is repeated to obtain the field-of-view coverage region that changes over time, forming a time series trajectory. When constructing the field-of-view geometric model, a buffer zone is introduced for slight inconsistencies caused by installation errors or map data errors. Extended polygons or rasterization approximations are used for the edge regions of the field of view to avoid the problem of unstable subsequent region matching due to overly idealized geometric boundaries.

[0036] The region code, time series, and frame-level globally unique identifier, device unique identifier, channel identifier, and unified analysis time in the urban video acquisition data frame sequence are bound together to generate a frame-level spatiotemporal association record for each frame of urban video acquisition data sequence, which is then written into the video spatiotemporal index table. In the video spatiotemporal index table, a time-dimensional index is established using the unified analysis time as the primary sorting key, a spatial-dimensional auxiliary index is established using the region code and time series, and a channel-dimensional joint index is established using the device unique identifier, channel identifier, and frame-level globally unique identifier, enabling combined retrieval by time interval, spatial region, and channel identifier. During the construction process, the time distribution of the unified analysis of unique identifiers of different devices under the same regional encoding or time series is aligned and statistically analyzed to identify camera combinations with overlapping fields of view within the same time period, providing a set of candidate channels for cross-camera event reconstruction and comparison of multi-source detection results; potential correlations are established for frame pairs with overlapping fields of view and time, and their frame-level globally unique identifiers are recorded as pairs as field of view overlap index entries to support the reuse of the same video samples under the same scene and the same time window, and to perform parallel inference and effect evaluation for different algorithm model versions; and a spatiotemporal index view of urban video acquisition is output.

[0037] A frame-level globally unique identifier is used to uniquely mark a frame of video acquisition data globally. It is preferably concatenated and encoded in the order of device unique identifier, channel identifier, unified analysis time, and intra-frame sequence number, or a fixed-length hash value is calculated based on the above concatenated string. The time series is used to assist in the joint indexing of time and space. It is preferably discretized and encoded according to the time step based on the unified analysis time, mapping the unified analysis time to a monotonically increasing integer index. Within the same time step, multiple frames of data are further distinguished according to the intra-frame sequence number, so as to ensure time accuracy while taking into account indexing efficiency and storage overhead.

[0038] As shown in Table 1, the frame-level quality feature vector table has the following characteristics: Frame-level globally unique identifier F20240520080001: Sharpness 85.3, Luminosity 120, Noise Level 0.02, Occlusion Ratio 3.20%, and Freeze Mark "No". Based on these indicators, this frame's quality level is Grade A, indicating excellent image quality that meets the requirements for high-precision analysis. Frame-level globally unique identifier F20240520080002: Sharpness 62.1, Luminosity 85, Noise Level 0.05, Occlusion Ratio 5.70%, and Freeze Mark "No". Based on these indicators, this frame's quality level is Grade B, indicating good image quality that meets the requirements for routine analysis. Frame-level globally unique identifier F20240520080003: Sharpness 30.7, Luminosity 200, Noise Level 0.11, Occlusion Ratio 12.50%, Freeze Mark "No"; Based on all indicators, this frame's quality level is Grade C, indicating poor image quality, requiring caution or preprocessing for analysis. Frame-level globally unique identifier F20240520080004: Sharpness 78.9, Luminosity 110, Noise Level 0.03, Occlusion Ratio 1.80%, Freeze Mark "No"; Based on all indicators, this frame's quality level is Grade A, indicating excellent image quality, meeting the requirements for high-precision analysis. Frame-level globally unique identifier F20240520080005: Sharpness 45.2, Luminosity 50, Noise Level 0.08, Occlusion Ratio 8.30%, Freeze Mark "No"; Based on all indicators, this frame quality level is Grade B, indicating good image quality that meets routine analysis needs. Frame-level globally unique identifier F20240520080006: Sharpness 22.5, Luminosity 180, Noise Level 0.15, Occlusion Ratio 35.60%, Freeze Mark "Yes"; Based on all indicators, this frame quality level is Grade C, indicating extremely poor image quality, and is not recommended for direct analysis. In summary, frame-level quality level is significantly correlated with various indicators: Accurate frame-level quality grading is achieved through multi-dimensional quantitative features, providing a quantitative basis for image data screening and analysis priority determination.

[0039] Table 1 Frame-level Quality Feature Vector Table

[0040] This implementation scheme achieves fine binding between the spatiotemporal location, acquisition channel, and image quality level of each video frame, enabling orderly screening and graded use of images of different quality levels in the same scene. On the one hand, it provides a spatiotemporally consistent and quality-controllable data foundation for target detection, model evaluation, and multi-version comparison. On the other hand, by identifying and labeling low-quality frames, it improves the reliability and stability of the overall analysis results.

[0041] Specifically, the process of edge-side preprocessing and image quality measurement for urban video capture data frame sequences is as follows: Input the urban video capture data frame sequence and the urban video capture spatiotemporal index view; perform edge-side preprocessing and image quality measurement on the urban video capture data frame sequence to form a frame-level quality feature vector sequence; configure a preprocessing task configuration set in the edge computing node according to the channel type, scene category, and computing resources. The preprocessing task configuration set includes parameters such as whether to enable geometric distortion correction, whether to enable noise filtering, whether to enable image stabilization, target output resolution, and target output frame rate; bind a corresponding preprocessing task template to each channel; and generate a preprocessing task instance that can be loaded at runtime.

[0042] In the edge computing nodes, a circular frame buffer queue is configured for each path. The sequence of city video acquisition data frames is written into the buffer according to a unified analysis time order, ensuring that the buffer always retains a continuous sequence of frames from the most recent period. The preprocessing task scheduler selects frames to be processed from the circular buffer queue at time intervals based on the target output frame rate of the preprocessing task instance, driving the preprocessing pipeline to process the frame data. In the geometric distortion correction stage, radial and tangential distortion corrections are performed on the image based on the distortion parameters involved in the camera, eliminating geometric distortion caused by barrel or pincushion distortion. In the denoising and enhancement stage, spatial filtering, edge-preserving filtering, or temporal filtering are applied to the image. Domain smoothing and other methods are used to reduce random noise and enhance structural details. In the resolution and cropping stage, the image is scaled to a uniform target resolution, and occluded borders and overlapping masking areas that have no effective information for a long time are cropped. In the image stabilization stage, continuous frames of vehicle-mounted and drone-mounted moving devices are subjected to motion compensation through feature point matching and motion estimation to suppress image jitter and improve the stability of target contours and textures. The preprocessing pipeline records the key parameters and activation status used in each step of the process and encapsulates them as preprocessing metadata, which is bound to the corresponding frame-level globally unique identifier to provide preprocessing context information for the algorithm's running data frame recording.

[0043] For each frame of the image, frame-level image quality feature vector calculations are performed: multiple image quality evaluation operators are selected to analyze the image. In the sharpness dimension, Laplacian energy, gradient energy, or high-frequency component energy are used to measure image sharpness and blurriness. In the brightness and contrast dimension, the average brightness, contrast, and dynamic range occupancy of the image are statistically analyzed based on grayscale histograms or brightness distribution to detect overexposure, underexposure, and low contrast. In the noise and compression artifact dimension, the noise level and compression artifact intensity are estimated through flat area texture analysis and block effect detection. In the occlusion and image integrity dimension, the proportion of areas covered by fixed occlusions, stains, or occlusion masks in the image is estimated using edge density, texture distribution, or occlusion masks. In the frozen frame detection dimension, frozen frames with unchanged content for a long time are identified through pixel differences between consecutive frames, feature hashing, or motion vector statistics, and frozen or stuttering frames are marked. These are combined sequentially to form a frame-level quality feature vector, and the physical meaning, value range, and suggested threshold are recorded for each dimension of the feature in the quality feature dictionary. The quality feature dictionary is designed as a structured configuration table. For each quality feature dimension, it records the feature name, the corresponding physical dimension, the feature calculation formula, the typical range of feature values, the set of judgment thresholds, and the method of obtaining the thresholds. For example, for the sharpness feature, the variance based on the Laplacian operator can be used as the sharpness index. The image is filtered by Laplacian and the variance of the filtering result is used as the sharpness feature value. The corresponding fields in the feature dictionary record "Laplacian variance" as the feature name, "used to characterize edge strength and blur degree" as the physical meaning, and "statistical index based on spatial domain edge operator" as the formula source explanation. A set of recommended thresholds for classifying sharpness levels is also provided.

[0044] Based on frame-level quality feature vectors, a quality score and quality level label are obtained for each frame. This is achieved by weighted fusion of quality features such as sharpness, noise level, contrast, brightness stability, compression artifact degree, occlusion ratio, and jitter amplitude. Frames with a quality score not lower than the first quality threshold are labeled as Grade A, frames with a quality score between the first and second quality thresholds are labeled as Grade B, frames with a quality score between the second and third quality thresholds are labeled as Grade C, and frames with a quality score lower than the third quality threshold are labeled as Grade D. The values ​​of each quality threshold are obtained by offline statistical analysis and subjective evaluation labeling fitting on a large-scale sample set and can be adjusted according to the actual deployment scenario. Frame-level quality feature vectors, quality scores, and quality level labels are uniformly written into a frame-level quality feature table, forming a sequence of urban video quality feature vectors arranged in chronological order. During the writing process, the regional coding and field of view area identification fields from the spatiotemporal index view of urban video acquisition are retained, so that each quality feature record can be traced not only to the specific device and channel, but also to the specific spatial region and field of view coverage, providing a spatial dimension of quality context for statistical quality status by region and performance analysis model by field of view; the frame-level quality feature vector sequence and quality level label set are output.

[0045] This implementation scheme enables the screening, hierarchical scheduling, and differentiated processing of video samples based on quality thresholds. It achieves a quality-driven strategy that prioritizes high-quality images for high-precision analysis and uses degraded images for robustness assessment or quality alarms. This effectively reduces the interference of low-quality images on detection results and improves the reliability, interpretability, and resource utilization efficiency of the overall visual analysis chain.

[0046] Specifically, the process of performing channel quality time series evaluation and quality-driven frame scheduling to output the video detection input set is as follows: Input the frame-level quality feature vector sequence, the urban video acquisition data frame sequence, and the urban video acquisition spatiotemporal index view; perform channel quality time series evaluation on the image quality changes of multi-source video channels in the time dimension; and perform quality-driven frame filtering and scheduling accordingly to reconstruct a quality-aware video detection input set; the frame-level quality feature vector sequence is a set of quality records with a frame-level globally unique identifier as the primary key. Each record includes a frame-level globally unique identifier, a device unique identifier, a channel identifier, a unified analysis time, a frame-level quality feature vector, a quality score, a quality level, and a quality reason code field.

[0047] Using the unique device identifier and channel identifier as channel indexes, frame-level quality records belonging to the same channel are sorted by a unified analysis time to construct a channel quality time series. Quality characteristics are statistically analyzed within a sliding time window: in a short time window, the proportion of quality frames of levels A, B, and C, the average quality score, and the standard deviation are statistically analyzed to identify short-term quality fluctuations and transient anomalies; in a long time window, the longest duration of consecutive level C low-quality frames, the number of quality score mutations, and the number of frozen frame events are statistically analyzed to identify long-term channel degradation trends and intermittent fault manifestations. During the statistical process, different weights are assigned to different quality feature dimensions according to the application scenario, with higher weights given to indicators highly correlated with detection performance, such as sharpness, occlusion ratio, and freeze duration, making the channel quality evaluation more consistent with the actual dependence of visual inspection tasks on image quality.

[0048] A channel quality health index is constructed. Channels with a Class A frame ratio higher than the threshold, a Class C frame ratio lower than the threshold, and no abnormal events for a long period are rated as normally usable. Channels with short-term quality degradation or occlusion but still usable overall are rated as slightly degraded. Channels with a long-term Class C frame ratio higher than the threshold, frequent freezing or occlusion events, are rated as severely degraded or temporarily not recommended for use. Simultaneously with generating the channel quality health index, the channel quality status and its changes are written to the channel quality status table and the quality anomaly log table. The channel quality status table records: device unique identifier, channel identifier, latest quality health index, channel quality status enumeration value, last update time, and a summary of the quality problem type, used to indicate the current channel availability. The quality anomaly log table records the time interval for each degradation from high quality to degraded or severely degraded status, the corresponding area code and field of view set, the main quality anomaly characteristics, and recommended maintenance actions, such as "recommend cleaning the lens," "recommend adjusting the installation angle," and "recommend checking the power supply and encoding configuration."

[0049] Based on channel quality status and frame-level quality levels, a quality-driven video detection frame scheduling strategy is constructed: For channels in a normal and usable state, all Class A frames participate in detection as detection candidate frames, Class B frames enter the detection queue according to the sampling ratio, and Class C frames are generally not included in the main detection link, only reserved for quality diagnosis or robustness assessment; for channels in a slightly degraded state, the sampling ratio of Class B frames can be reduced, and a channel quality status field can be added to the detection task metadata, enabling the algorithm's running data records to perceive the channel's degraded state; for channels in a severely degraded state, Class C frames can be temporarily suspended from participating in the main detection, and only Class A and Class B frames are retained as monitoring samples, while the time period and reason for the channel's degraded use are recorded in the quality anomaly log. For each frame entering the detection process, when generating the detection task, the frame identifier, the globally unique frame-level identifier, and the corresponding quality score, quality level, channel quality health index, region coding, and field of view region identifier quality and spatiotemporal context fields are written into the detection task queue list, so that the visual detection module can directly reference this context information when constructing the algorithm's running data frames and model version view.

[0050] Output quality-aware video detection input set. The input set consists of a set of frame-level task entries with a globally unique frame identifier, a unique device identifier, a channel identifier, a unified analysis time, and quality and spatiotemporal context fields. Each task entry can be traced back to the original video frame content, acquisition link information, and channel quality evolution record in the storage system.

[0051] This implementation scheme achieves automatic screening and graded use of input samples for detection. On the one hand, it prioritizes high-quality frames to enter the main detection link, reducing the interference of low-quality images on the accuracy and stability of the algorithm. On the other hand, it explicitly marks and downgrades degraded or abnormal channels, providing an interpretable quality context for operation and maintenance decisions and model evaluation. Thus, without increasing the complexity of the front-end equipment, it significantly improves the utilization rate of effective samples, the reliability of detection results, and the level of precision in system operation and maintenance.

[0052] Specifically, the process of constructing the algorithm execution data frame set based on the video detection input set is as follows: Input the video detection input set, uniformly collect and structure-encapsulate the detection results, model configuration, and runtime environment characteristics generated during the execution of each detection task, and construct the algorithm execution data frame set; based on the task entries in the video detection input set, select the model version and runtime configuration that match the task type, scene category, and quality conditions: for example, for a pedestrian detection task under nighttime illumination, prioritize the night vision optimized version of the pedestrian detection model; for a vehicle detection task in a high-resolution road scene, select the wide field-of-view vehicle detection model version; when existing... When multiple candidate versions are used for comparative evaluation, a primary model version and a comparison model version can be configured simultaneously for the same frame task, and the roles of the primary and comparison versions are marked in the task metadata. Key information such as the selected model identifier, model version number, input resolution, confidence threshold, non-maximum suppression threshold, multi-scale switch, and batch size are bound to the corresponding detection task to form a detection task record. The detection scheduling component selects an inference node that meets the model hardware requirements and has a suitable load based on task priority and the current load of the inference node, and distributes the detection task to the corresponding inference node for execution, recording the inference node identifier and distribution time bound to the task. Frame-level globally unique identifiers are preferably generated according to unified rules during the acquisition phase, mapped to a fixed-length identifier value using a hash algorithm or UUID generation function, ensuring that frame-level globally unique identifiers corresponding to different frames do not conflict, and facilitating their use as primary or foreign keys across multiple tables.

[0053] On the inference node side, the corresponding preprocessed image frame is located and loaded into the city video acquisition data frame table based on the frame-level globally unique identifier. The image is then input into the specified model version for inference. The detection results output by the model are structurally encapsulated to obtain the basic record of the detection results. The basic record of the detection results includes: the internal identifier of the target or event, the corresponding frame-level globally unique identifier, the device unique identifier and channel identifier, the target category label or event type, the target bounding box or region contour coordinates, the detection confidence score, and the target quantity statistics. If it is a temporal event detection, it further includes the event duration estimate and start and end time estimate. In the scenario of parallel inference of multiple model versions, the detection results obtained by the same frame-level globally unique identifier under different model versions are assigned the same candidate target group identifier, so that the detection differences of the same target under different model versions can be compared one by one in the model version view. The structured encapsulation adopts a fixed set of fields and type constraints. Geometric information such as bounding box coordinates and contour point sets are stored in a unified coordinate system and a unified format. Category labels and event types are constrained by a standardized encoding table to ensure that the output results of different model versions and different inference nodes can be directly compared and aggregated under the same data structure.

[0054] During inference execution, runtime characteristic information of this inference call is collected in real time, and the inference environment and resource consumption are quantified into runtime characteristic vectors. The runtime characteristic information includes: the start and end time of inference in this frame or batch, the average inference latency per frame, the CPU utilization range during inference, the GPU utilization range, the memory and video memory usage range, and the number of concurrent inference tasks or queue length of the current node. When the inference latency is detected to be significantly higher than the historical statistical upper limit threshold of the model on the node, this inference is marked as a latency anomaly. When the video memory is close to full or multiple memory shortage alarms occur, this inference is marked as resource stressful. When model loading fails, operator execution errors occur, or the process is interrupted during inference, the corresponding error code and brief error information are recorded, and this inference is marked as a runtime error. The runtime characteristic vector and the anomaly markers together constitute the runtime environment descriptor, which is used to characterize the resource status and environmental conditions under which this detection is completed. The running feature vector is defined as a fixed-length multi-dimensional vector, with each dimension corresponding to a fixed meaning field. Each field registers its physical meaning, unit, and statistical method in the running feature dictionary, and provides a reference threshold for anomaly detection. The running environment descriptor is stored in the algorithm running data frame table in the form of key-value pairs or columnar fields during structured encapsulation, so that the running feature vector has comparability and interpretability at different nodes and at different time periods.

[0055] The algorithm execution data frame is constructed using a combined primary key of a frame-level globally unique identifier and the model version number. Acquisition information, model and configuration information, detection results information, preprocessing context, quality context, and runtime environment information are uniformly encapsulated into a single algorithm execution data frame record. Acquisition information includes: frame-level globally unique identifier, device unique identifier, channel identifier, unified analysis time, region encoding, and field of view region identifier. Model and configuration information includes: model identifier, model version number, task type, input resolution, threshold configuration, and multi-model parallelism marker. Detection results information includes: target category, bounding box or region outline, confidence score, and target quantity. Preprocessing context includes corresponding preprocessing metadata identifier and preprocessing parameter summary. Quality context includes: frame-level quality score, quality level, and channel quality health index. Runtime environment information includes: inference node identifier, hardware type, inference framework version, runtime feature vector, and anomaly marker. These are stored in the algorithm execution data frame table as fixed fields or expandable field groups. Different types of fields maintain consistency through primary keys and foreign keys, thus supporting the expansion of the field set without changing the primary key design.

[0056] The algorithm execution data frame records are written into the algorithm execution data frame table, and a relationship is established with the city video acquisition data frame table, video spatiotemporal index table, frame-level quality feature table and channel quality status table through the frame-level globally unique identifier. A relationship is established with the model identifier, model version number and model version registration information set, and the algorithm execution data frame set is output.

[0057] As shown in Table 2, the channel quality status is as follows: Device unique identifier DEV-001: Channel identifier CH-01, latest quality health index 92.5, channel quality status normal and usable, last updated 2024 / 5 / 20-9:30, quality problem type none, recommended maintenance action no maintenance required; Device unique identifier DEV-002: Channel identifier CH-03, latest quality health index 78.3, channel quality status slight degradation, last updated 2024 / 5 / 20-9:15, quality problem type slight obstruction, recommended maintenance action: clean lens; Device unique identifier DEV-003: Channel identifier CH-02, latest quality health index 65.7, channel quality status slight degradation, last updated 2024 / 5 / 20-8:45, quality problem type short-term jitter, recommended maintenance action: check power supply stability; Device unique identifier DEV-004: Channel identifier CH-01, latest quality health index 42.1, channel quality status severely degraded, last updated 2024 / 5 / 20-8:30, quality issue type is large-area obstruction + noise, recommended maintenance action is to adjust the installation angle + clean the lens; Device unique identifier DEV-005: Channel identifier CH-04, latest quality health index 38.6, channel quality status severely degraded, last updated 2024 / 5 / 20-8:10, quality issue type is persistent frame freezing, recommended maintenance action is to check encoding configuration + network link; Device unique identifier DEV-006: Channel identifier CH-02, latest quality health index 89.7, channel quality status normal and usable, last updated 2024 / 5 / 20-9:00, quality issue type is none, recommended maintenance action is no maintenance required.

[0058] Table 2 Channel Quality Status Table

[0059] like Figure 3The time-series trend chart of the channel quality health index shows a "decline followed by recovery" trend over time: Initially, the index is close to 90, in a relatively healthy range. It then declines daily, briefly falling below the mild degradation threshold of 80 in the middle, but remaining above the severe degradation threshold of 60, forming the "mild degradation risk zone" highlighted in orange in the chart. After this zone, the health index gradually recovers and crosses the 80 threshold again, indicating that the channel quality has recovered after a period of degradation. The two dashed lines represent the mild degradation threshold and the severe degradation threshold, respectively. Together with the health index curve, they visually indicate the channel's state changes between normal, mild degradation, and potentially severe degradation, providing a basis for channel quality monitoring and early warning.

[0060] In this implementation plan, a correlation is established between the inference node identifier and the inference node environment information set, so that each detection result has a complete traceable link at the data level, from the source of data collection to its spatiotemporal location, quality status, model version, operating environment, and detection output, providing basic data for the construction and comparative evaluation of the model version view.

[0061] Specifically, the process of constructing the model version view and traceable evaluation index set is as follows: Input the set of algorithm running data frames, the sequence of urban video acquisition data frames, the spatiotemporal index view of urban video acquisition, the sequence of frame-level quality feature vectors, and the channel quality status records. Aggregate and compare the detection results and running characteristics of different model versions on a unified data basis to construct a model version view and traceable evaluation index set for version management: Group and aggregate the set of algorithm running data frames with model identifier and model version number as the main dimensions to construct the model version view; establish a one-to-many relationship with the corresponding field in the algorithm running data frame table as a foreign key, so that any model version view can trace back to all running data frame records covered. In the time dimension, the runtime data is divided into time buckets according to a unified analysis time. The number of frames processed, the number of targets detected, and the number of runtime anomalies within each time bucket are statistically analyzed to form a time series statistics. The time bucketing rules are implemented by configuring an adjustable time window length. In the spatial dimension, the distribution of detection samples and output characteristics of the model version under different regions and different fields of view are statistically analyzed by combining region coding and field of view region identification to identify the applicability differences of the model in different scenarios. In the quality dimension, the detection results under different quality conditions are statistically analyzed by combining frame-level quality scores, quality levels, and channel quality health indices to distinguish the detection situation on high-quality frames, medium-quality frames, and low-quality frames, avoiding misjudging image quality issues as model issues. In the runtime characteristic dimension, the average inference latency, resource consumption, and anomaly rate of the model version at different nodes or different time periods are calculated based on the runtime feature vector to evaluate the performance and stability of the model.

[0062] In scenarios where reference results are available, the detection results in the algorithm's running data frames are compared with the reference results to obtain the model version's detection accuracy and robustness metrics: Accuracy metrics can include hit rate, false negative rate, and false positive rate, obtained by statistically analyzing the target matching relationship between the model output and the reference results under the same frame-level globally unique identifier; Robustness metrics can include the performance variation under different quality conditions such as complex lighting, strong occlusion, and low resolution, obtained by analyzing the difference and ratio between the baseline performance on high-quality frames and the performance on low-quality frames; Operational stability metrics can be obtained by statistically analyzing the proportion of operation anomaly markers and latency anomaly markers appearing in all running data frames of the model version, used to measure the engineering robustness of the model after deployment; Samples under different quality levels and different channel quality states are statistically analyzed separately so that the performance of the same model version under ideal, normal, and degraded image quality conditions can be distinguished.

[0063] For scenarios with parallel detection using multiple model versions, a cross-version comparison subview is constructed by using a globally unique identifier at the same frame level as the association key to compare the detection results of multiple model versions for the same frame and the same target region one-to-one. In the comparison subview, the target category judgment, confidence level, whether there are missed or false detections, corresponding runtime latency and resource consumption of different model versions are recorded, and these differences are saved in a structured form. By comparing the detection accuracy and runtime performance of multiple versions on a unified sample set, it is possible to objectively evaluate whether the version upgrade brings substantial improvements without changing the collected data and business environment. Versions with abnormal behavior are marked with a risk in the model version view to prompt maintenance personnel to roll back or retrain.

[0064] For each model version, a version evaluation record containing multi-dimensional indicators is generated. The evaluation results are written to the model version view table and the version evaluation indicator table. The version evaluation record includes: model identifier, model version number, statistical coverage time range and regional range, detection accuracy indicators under different quality levels and channel states, operational stability indicators, average inference latency and resource consumption indicators, multi-version comparison result summary, and risk and suggestion tags. It is associated with the original configuration record of the model version registration information set. The model version view and evaluation indicator set are output. The generation rule of the version evaluation record is as follows: when the number of algorithm running data frames accumulated by a certain model version within a set statistical time window reaches the lower limit threshold of the sample, the version evaluation task is triggered. All relevant running data within the time window are aggregated and calculated centrally to generate one or more version evaluation records and write them to the version evaluation indicator table.

[0065] like Figure 4The comparison chart of detection accuracy across multiple model versions shows the overall trend of detection hit rate improvement with version iterations across different image quality levels, from model version v1.0 to v2.1 on the horizontal axis: Level A images consistently have the highest hit rate, maintaining a small but stable increase across versions; Level B images have a slightly lower hit rate than Level A, but similarly, they gradually approach the high accuracy range with version upgrades; Level C and D images had relatively low detection hit rates in early versions, but as the model evolved from v1.x to v2.0 and v2.1, the curve slopes became steeper and the improvement more significant, indicating that the new versions are significantly more robust to low-quality and degraded images. The v2.0 version is marked with an "upgrade inflection point," reflecting a significant improvement in the overall detection performance of the model across all quality levels after this version.

[0066] This implementation plan enables the analysis of traceable detection effects of different model versions in urban multi-source video scenarios and the evaluation of quantifiable version comparisons from the perspective of structured data. It provides a unified data foundation for version selection, rolling upgrades and anomaly location without relying on any business strategies or policy rules.

[0067] Specifically, the process of constructing a unified detection result service view based on the algorithm execution data frame set is as follows: Inputting the algorithm execution data frame set, model version view, and evaluation index set, as well as the urban video acquisition data frame sequence, urban video acquisition spatiotemporal index view, frame-level quality feature vector sequence, and channel quality status record; performing field trimming, semantic merging, and service encapsulation on the multi-source algorithm execution data frames to construct the unified detection result service view; based on the algorithm execution data frame set, performing hierarchical extraction and semantic merging of fields, structurally decoupling the internally used execution process fields from the externally served detection result fields; the field set of the outer service view adopts a fixed... The field list is constrained and includes: detection result service record identifier, frame-level globally unique identifier, running data frame identifier, unified analysis time, device unique identifier, channel identifier, target category label, target bounding box or region contour coordinates, target confidence or score, target quantity statistics, region encoding, field of view region identifier, frame-level quality level, simplified quality label, channel quality status label, model identifier, model version number, version health score summary, version recommendation level, and context reference identifier field. Each field is registered with semantic definition and value constraints in the service view metadata table to ensure that different upper-layer systems obtain consistent technical meaning when accessing the service view. The outer service view retains only technical fields directly related to the upper-layer system, including target category, target location description, detection time, device and channel identifier, region code, field of view identifier, image quality level and simplified quality label, model version number and version health summary, as well as the runtime data frame identifier and frame-level globally unique identifier for traceability. Fields such as preprocessing metadata identifier, complete quality feature vector, and detailed runtime feature vector inside the algorithm runtime data frame are not directly expanded in the service view, but are retained in the form of context reference identifiers: that is, a context reference field is recorded for each detection result in the service view, and the field points to the complete record in the algorithm runtime data frame table. This allows the external system to consume only simplified service fields by default, but can still recover all technical details when needed through context references.

[0068] The model version view and version health information and recommendation levels from the evaluation metric set are injected into each detection result record. The algorithm execution data frame is associated with the corresponding model version evaluation record through the model identifier and model version number. Two service fields, a version health score summary and a version recommendation level, are added to each detection result to indicate to the upper-level system the model version status from which the result originates, without exposing complex evaluation details. Version health is optimally defined as a comprehensive score obtained through normalized aggregation of multi-dimensional evaluation metrics. The evaluation metrics used include detection accuracy metrics (a function of hit rate, false negative rate, and false positive rate), operational stability metrics (a function of the proportion of operational anomaly markers and latency anomaly markers), and quality robustness metrics (stable performance on samples of different quality levels). The health score is calculated by weighting and fusing the health score with a function representing the fluctuation range of the performance and a cross-scenario consistency index (a function representing the performance dispersion in different regions and fields of view). The version recommendation level is determined segmented based on the version health score and thresholds in the evaluation indicators. When the version health score is not lower than the first health threshold and both operational stability and quality robustness indicators meet the stability threshold, the model version is marked as the first recommendation level. When the version health score is between the first and second health thresholds, or some indicators are close to the boundary, the model version is marked as the second recommendation level. When the version health score is lower than the second health threshold, or the false positive rate or anomaly rate significantly exceeds the safety threshold, the model version is marked as the third recommendation level or risk level. To avoid confusion caused by differences in evaluation results across different time periods and regions, the version health score is recorded together with the evaluation coverage time range when injecting the version health score. When the detection time exceeds the version evaluation coverage range, an "outside evaluation range" marker is added to the service view to remind the analysis team to pay attention to the uncertainty brought about by time extrapolation when interpreting the results.

[0069] A multi-level index structure is constructed for service fields: a time index is built with the unified analysis time as the main sorting key; a spatial index is built with the region code and field of view region identifier; a quality index is built with the quality level and channel quality health index; a version index is built with the model identifier and model version number; and a traceability index is built with the running data frame identifier. All of the above indexes are built on the service view layer, rather than on the original algorithm running data frame table. By maintaining an index metadata table in the service view to record the key value range, sharding strategy and refresh time of each type of index, the decoupling between external service queries and internal storage is achieved: external queries only access the service view and index structure, and internal deep backtracking is required before jumping to the algorithm running data frame table and other basic data tables through the traceability index; a unified detection result service view is output.

[0070] like Figure 5The ER diagram showing the relationships between algorithm execution data frame tables places the algorithm execution data frame table at the center of the relationships. It is associated with the city video acquisition frame table through a globally unique frame-level identifier, which is used to accurately trace each algorithm detection result back to the corresponding original video frame. At the same time, the algorithm execution data frame table is also associated with the model version table through the model identifier and model version number, and with the inference node environment table through the inference node identifier. Thus, a single execution record is bound with complete information of "acquisition frame - model version - inference node", realizing the traceability and reproducibility of detection results in three dimensions: acquisition data, algorithm version, and execution environment.

[0071] In this implementation plan, by leveraging a multi-dimensional index structure of time, space, quality, and version, the upper-layer system can efficiently query massive amounts of detection results according to business needs. Furthermore, when false alarms, missed alarms, or version anomalies occur, it can accurately locate specific frames, specific model versions, and specific operating environments. This effectively supports version comparison and evaluation, root cause analysis of problems, and strategy optimization decisions, thereby improving the availability, maintainability, and long-term evolution capabilities of the smart city intelligent video system as a whole.

[0072] Specifically, the process of constructing the traceable backtracking and original scene reconstruction interface is as follows: Input a unified detection result service view and an algorithm running data frame set, organize and index the relationship between detection results, model version, image quality, running environment and original video acquisition process, and construct a traceable backtracking interface that supports combined query and original scene reconstruction: The unified detection result service view provides simplified detection result records for the upper-layer system. Each record contains at least the target category, target location, detection time, device and channel identifier, area code and field of view identifier, quality level, version health summary and running data frame identifier and frame identifier traceability key fields.

[0073] The system employs a combined condition query mechanism: users can specify model version, time interval, region range, image quality conditions, channel quality status characteristics, and operational anomaly markers. With the support of multi-level indexes in the service view, matching detection result records are quickly filtered to obtain a candidate result set. In implementation, a time index quickly limits the time range of candidate records, a spatial index limits the range of region encoding and field of view identifiers, a quality index limits the range of quality levels or quality labels, a version index limits specific model versions or version sets, and a traceability index binds the filtered results to the operational data frame identifier, forming a set of results suitable for centralized backtracking. Result sets under the same query conditions can be further grouped by model version, quality level, or channel status to construct sample subsets for specific scenarios, such as "all false alarm records of a certain version on low-quality images" or "all detection records of a certain channel in a severely degraded state."

[0074] After obtaining the candidate detection result records and corresponding running data frame identifiers, the original technical context is reconstructed layer by layer through the source index: using the running data frame identifier as the key, the corresponding record is loaded from the algorithm running data frame table to restore the model identifier and model version number used in this detection; using the frame identifier as the frame-level globally unique identifier as the key, the corresponding original frame or adjacent frame sequence is located from the city video acquisition data frame table, and the unified analysis time, spatial coordinate position and field of view coverage area of ​​the frame are obtained by combining the city video acquisition spatiotemporal index view; through the frame-level quality feature vector sequence and channel quality status record, the image quality change and channel health index time curve within the time window before and after the frame are restored, and the image quality and channel status background of a detection result are fully unfolded in the time dimension; during the reconstruction process, preprocessing steps and parameters can be optionally loaded from the preprocessing configuration table corresponding to the preprocessing metadata identifier to determine whether distortion correction, denoising, image stabilization and cropping operations were performed on the frame image before inputting it into the detection model, so as to distinguish "acquisition problem, preprocessing problem, model problem or running environment problem" during false alarm or false alarm analysis.

[0075] The system provides a unified, traceable backtracking interface to the upper-layer system, exposing backtracking capabilities externally in standardized technical interface formats (such as request-response interfaces, subscription / push alarm backtracking interfaces, or offline batch export interfaces). Interface inputs include multi-dimensional query conditions or specific detection result service record identifiers, while the output is a structured backtracking result package. This package includes: the corresponding detection result service record, a summary of the associated algorithm execution data frame record, a reference to the corresponding original video frame or frame sequence, a summary of the image quality and channel status time curves, and model version information and runtime environment summary fields. The interface design does not introduce any specific business strategies, pricing strategies, or policy rules; it only provides data access and reconstruction capabilities. After obtaining the backtracking result package, the upper-layer business system can make business-level judgments and decisions based on its own logic. Simultaneously, log records of the input conditions and returned results are retained in each backtracking call, constructing a backtracking call log table for analyzing the upper-layer system's usage patterns of the backtracking capabilities and further optimizing the index structure and storage layout.

[0076] In this implementation plan, a traceable backtracking mechanism enables any detection result in the smart city intelligent video system to be technically retrieved back to the scene and linked to the acquisition environment, image quality, model version, and operating resource status, providing basic support for model iteration, system operation and maintenance, and anomaly diagnosis.

[0077] Specifically, the second aspect of this invention provides an intelligent video system application system applicable to smart cities, and an application method for an intelligent video system applicable to smart cities, comprising: a video acquisition module, used to construct urban video acquisition data frames from multi-source video frame data, perform field-of-view geometric modeling and region coding on the installation location and field of view of the acquisition terminal, and construct a spatiotemporal index view of urban video acquisition; a quality perception module, used to perform edge-side preprocessing and image quality measurement on the urban video acquisition data frame sequence, and perform channel quality time series evaluation and quality-driven frame scheduling, filter and output a video detection input set that meets the detection quality requirements, and output the video detection input set; a visual detection module, used to construct an algorithm running data frame set based on the video detection input set, record detection output, model identifier and model version number, running node and running feature information at the frame level, and construct a model version view and a traceable evaluation index set; and a view and traceability evaluation module, used to construct a unified detection result service view based on the algorithm running data frame set, provide a structured detection result access interface to the outside world, and construct a traceable backtracking and original scene reconstruction interface.

[0078] This implementation plan not only significantly improves the automation and precision of large-scale smart city video systems in terms of data management, quality control, and algorithm evaluation, but also enhances the interpretability and traceability of detection results in multiple dimensions, including time, space, quality, and model version, providing a stable and reliable technical foundation for the iterative optimization of visual algorithms and system operation and maintenance.

[0079] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0080] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A method for applying an intelligent video system suitable for smart cities, characterized in that, Includes the following steps: S1, collect multi-source video frame data to construct urban video acquisition data frames, perform field geometry modeling and regional coding to construct urban video acquisition spatiotemporal index view; S2 performs edge-side preprocessing and image quality measurement on the urban video acquisition data frame sequence, and performs channel quality time series evaluation and quality-driven frame scheduling to output the video detection input set; S3, based on the video detection input set, constructs a set of algorithm running data frames, and builds a model version view and a set of traceable evaluation indicators; S4 constructs a unified detection result service view based on the algorithm's data frame set, and builds an interface for traceable backtracking and original scene reconstruction.

2. The application method of an intelligent video system suitable for smart cities according to claim 1, characterized in that: The specific process of constructing urban video acquisition data frames from the collected multi-source video frame data is as follows: Input multi-source video frame data collected by fixed cameras, PTZ cameras, vehicle-mounted terminals, ship-mounted terminals, and cameras mounted on drones; construct a unified time base and schedule channel sampling for the multi-source video frame data to form a sequence of urban video acquisition data frames. During the acquisition process, a unified time source calibrated by a unified time service is used as the benchmark to assign a unified sampling timestamp to each video sample. Sampling records from different video terminals are sorted by sampling timestamp and assigned to the corresponding sampling period. The sampling timestamps are normalized using a unified time unit and starting reference point, and the local time of different devices is mapped to a relative time index under a unified time benchmark. When generating a unified acquisition frame sequence, the sampling timestamps assigned by the unified time source are used as the main time axis. The video sampling records of each channel on the same time axis are time-aligned, and a sequence of urban video acquisition data frames arranged in a unified sampling time order is output.

3. The application method of an intelligent video system suitable for smart cities according to claim 1, characterized in that: The specific process of constructing a spatiotemporal index view of urban video acquisition through field-of-view geometric modeling and region coding is as follows: The input consists of a city basic geographic information dataset composed of a base topographic map, road and building vector data, administrative division and functional zoning data, and a sequence of city video acquisition data frames. Geometric modeling and regional coding are performed on the installation location and field of view of each video terminal. The city video acquisition data frame sequence is mapped to a unified spatiotemporal coordinate system to construct a spatiotemporal index view of city video acquisition. The installation point of each fixed camera is mapped to the city unified coordinate system through coordinate transformation. The administrative division boundaries, road centerlines and park boundaries that intersect with the installation point and surrounding area in the city basic geographic information dataset are queried. The administrative division code and road or park identifier to which the installation point belongs are used as regional codes to establish a one-to-one or one-to-many mapping relationship, forming a time-varying set of device trajectory points. The region code, time series, and frame-level globally unique identifier, device unique identifier, channel identifier, and unified analysis time in the urban video acquisition data frame sequence are bound together to generate a frame-level spatiotemporal association record for each frame of urban video acquisition data frame sequence and write it into the video spatiotemporal index table; output the urban video acquisition spatiotemporal index view.

4. The application method of an intelligent video system suitable for smart cities according to claim 1, characterized in that: The specific process of performing edge-side preprocessing and image quality measurement on the urban video acquisition data frame sequence is as follows: Input the sequence of urban video capture data frames and the spatiotemporal index view of urban video capture, perform edge-side preprocessing and image quality measurement on the sequence of urban video capture data frames, and form a sequence of frame-level quality feature vectors. In the edge computing nodes, a circular frame buffer queue is configured for each path. The sequence of urban video acquisition data frames is written into the buffer according to a unified analysis time order, so that the buffer always retains the continuous frame sequence of the most recent period. Frame-level image quality feature vector operation is performed on each frame: multiple image quality evaluation operators are selected to analyze the image, and they are combined in order to form a frame-level quality feature vector. The physical meaning, value range and suggested threshold are recorded for each feature dimension in the quality feature dictionary. Based on frame-level quality feature vectors, the quality score and quality level label of each frame are obtained; the frame-level quality feature vectors, quality scores and quality level labels are written into the frame-level quality feature table to form a sequence of urban video quality feature vectors arranged in chronological order; the frame-level quality feature vector sequence and the set of quality level labels are output.

5. The application method of an intelligent video system suitable for smart cities according to claim 1, characterized in that: The specific process of performing channel quality time series evaluation and quality-driven frame scheduling to output the video detection input set is as follows: The system inputs a sequence of frame-level quality feature vectors, a sequence of urban video capture data frames, and a spatiotemporal index view of urban video capture data. It then performs a time-series evaluation of the image quality changes across multiple video channels over time, and executes quality-driven frame filtering and scheduling accordingly, reconstructing a quality-aware video detection input set. Using the unique device identifier and channel identifier as channel indices, it sorts frame-level quality records belonging to the same channel according to a unified analysis time, constructing a channel quality time series, and statistically analyzes the quality features within a sliding time window. Construct a channel quality health index, and while generating the channel quality health index, write the channel quality status and change process into the channel quality status table and the quality anomaly log table; Based on channel quality status and frame-level quality level, a quality-driven video detection frame scheduling strategy is constructed: outputting a quality-aware video detection input set.

6. The application method of an intelligent video system suitable for smart cities according to claim 1, characterized in that: The specific process of constructing the data frame set based on the video detection input set is as follows: The input video detection input set is used to uniformly collect and structurally encapsulate the detection results, model configuration, and runtime environment characteristics generated during the execution of each detection task, and construct an algorithm runtime data frame set. Based on the task entries in the video detection input set, a model version and runtime configuration matching the task type, scene category, and quality conditions are selected. On the inference node side, the corresponding preprocessed image frame is located and loaded into the urban video acquisition data frame table based on the frame-level globally unique identifier. The image is input into the specified model version for inference, and the detection results output by the model are structurally encapsulated to obtain the basic record of the detection results. During the inference execution process, the runtime characteristic information of this inference call is collected in real time, and the inference environment and resource consumption are quantified into runtime characteristic vectors; Construct algorithm execution data frames, record the algorithm execution data frame records into the algorithm execution data frame table, and output the set of algorithm execution data frames.

7. The application method of an intelligent video system suitable for smart cities according to claim 1, characterized in that: The specific process for constructing the model version view and traceable evaluation indicator set is as follows: The algorithm runs a set of data frames, as well as a sequence of urban video capture data frames, a spatiotemporal index view of urban video capture, a sequence of frame-level quality feature vectors, and channel quality status records. The detection results and running characteristics of different model versions are aggregated and compared on a unified data basis. A model version view and a traceable evaluation index set for version management are constructed: the algorithm runs a set of data frames are grouped and aggregated based on model identifier and model version number as the main dimensions to construct a model version view. Samples at different quality levels and under different channel quality conditions are statistically analyzed separately, so that the performance of the same model version under ideal, normal and degraded image quality conditions can be distinguished. For scenarios with multiple model versions configured for parallel detection, the detection results of multiple model versions for the same frame and the same target area are compared one-to-one using the same frame-level globally unique identifier as the association key, and a cross-version comparison sub-view is constructed; the model version view and evaluation index set are output.

8. The application method of an intelligent video system suitable for smart cities according to claim 1, characterized in that: The specific process of constructing a unified detection result service view based on the algorithm-driven data frame set is as follows: The system takes as input a set of algorithm execution data frames, a model version view, and an evaluation index set, as well as a sequence of urban video acquisition data frames, a spatiotemporal index view of urban video acquisition, a sequence of frame-level quality feature vectors, and channel quality status records. It then performs field trimming, semantic merging, and service-oriented encapsulation on the multi-source algorithm execution data frames to construct a unified detection result service view. Based on the set of algorithm execution data frames, it performs hierarchical extraction and semantic merging of fields, structurally decoupling the internally used execution process fields from the externally used detection result fields. In the outer service view, it retains only the technical fields directly related to the upper-layer system, injects the version health information and recommendation level from the model version view and evaluation index set into each detection result record, and outputs a unified detection result service view.

9. The application method of an intelligent video system suitable for smart cities according to claim 1, characterized in that: The specific process of constructing the traceable backtracking and original scene reconstruction interface is as follows: The system takes a unified detection result service view and a set of algorithm execution data frames as input, organizes and indexes the relationships between detection results, model version, image quality, operating environment and the original video acquisition process, and builds a traceable backtracking interface that supports combined queries and original scene reconstruction. It designs a combined condition query mechanism: users can specify model version, time interval, area range, image quality conditions, channel quality status characteristics and operation anomaly marker conditions, and quickly filter matching detection result records under the support of multi-level indexing in the service view to obtain a set of candidate results. After obtaining the candidate detection result record and the corresponding running data frame identifier, the original technical context is reconstructed layer by layer through the source index: using the running data frame identifier as the key, the model identifier and model version number used in this detection are restored; using the frame identifier as the frame-level globally unique identifier as the key, the unified analysis time, spatial coordinate position and field of view coverage area of ​​the frame are obtained; the image quality change and channel health index time curve within the time window before and after the frame are restored, providing a unified traceable backtracking interface for the upper layer system.

10. An intelligent video system application system suitable for smart cities, employing the intelligent video system application method for smart cities as described in any one of claims 1-9, characterized in that, include: The video acquisition module is used to collect multi-source video frame data to construct urban video acquisition data frames, perform field-of-view geometric modeling and regional coding to construct a spatiotemporal index view of urban video acquisition; The quality perception module is used to perform edge-side preprocessing and image quality measurement on the urban video acquisition data frame sequence, and to perform channel quality time series evaluation and quality-driven frame scheduling to output the video detection input set. The visual inspection module is used to construct a set of data frames for algorithm operation based on the video inspection input set, and to build a model version view and a set of traceable evaluation indicators. The View and Traceability Evaluation Module is used to build a unified detection result service view based on the algorithm-run data frame set, and to build traceable backtracking and original scene reconstruction interfaces.

Citation Information

Patent Citations

  • A Distributed and Extensible Intelligent Video Surveillance System

    CN103475870B

  • An intelligent video recognition system and its application

    CN109271554B