Big data rapid processing operation method based on smart city
By adopting multimodal deep learning model and adaptive compression technology in the smart city video surveillance system, intelligent partitioning and dynamic storage of video data are realized, solving the problems of resource waste and data loss in traditional methods, and improving processing efficiency and storage efficiency.
Patent Information
- Application Number
- CN202510836087.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-06-20
AI Technical Summary
In the existing smart city video surveillance system, the traditional full storage method leads to wasting of storage resources, and the frame extraction process at fixed time intervals is prone to loss of key information. The general compression algorithm is inefficient, and the data processing process and storage strategies lack a dynamic coordination mechanism, resulting in significant system response delays.
Adaptive compression method based on multimodal deep learning model is adopted, and video acquisition intervals are dynamically adjusted through edge computing nodes, combined with spatiotemporal similarity calculation and multi-dimensional feature vectors to realize intelligent partitioning and dynamic storage, and the hybrid compression engine and dual-layer redundancy elimination mechanism are used to optimize data processing and storage strategies.
On the premise of ensuring information integrity, it significantly improves processing efficiency, reduces storage requirements, improves video quality and identification accuracy, especially in low-light scenarios to ensure the integrity of key information, solving the problems of resource waste and data loss in traditional methods.
Smart Images

Figure CN120434362A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer processing technology, and in particular to a method for rapidly processing and operating big data based on a smart city. Background Art
[0002] As the construction of smart cities accelerates, the scale of urban road video surveillance networks is growing exponentially. Statistics show that a single camera can generate hundreds of GB of video data daily. The massive amount of data generated by a cluster of millions of cameras poses a severe challenge to existing processing systems.
[0003] Traditional video data processing methods have the following main defects:
[0004] 1) Using full storage results in a storage resource waste rate exceeding 60%;
[0005] 2) Frame extraction based on fixed time intervals is prone to losing key event information;
[0006] 3) General compression algorithms are difficult to take into account the characteristics of traffic scenes, resulting in low compression efficiency;
[0007] 4) There is a lack of dynamic coordination between data processing procedures and storage strategies, resulting in significant system response delays.
[0008] Therefore, there is an urgent need to develop a big data processing method that integrates intelligent partitioning, adaptive compression and dynamic storage collaboration to achieve coordinated optimization of processing efficiency and storage efficiency while ensuring data integrity. Summary of the Invention
[0009] In response to the above technical problems, the technical solution adopted by the present invention is a method for rapid processing and operation of big data based on smart cities, which includes the following steps:
[0010] S01. Obtaining video data packets collected by urban road cameras within a preset time period, and extracting picture content features reflecting traffic dynamics from the video data packets, wherein the picture content features divide a coverage area into a plurality of data blocks;
[0011] S02, extracting key frames from the video data packet for each data block, calculating the spatiotemporal similarity of adjacent key frames, and combining the time series analysis to generate a data redundancy evaluation result;
[0012] S03, performing adaptive compression processing on the video data packet according to the data redundancy evaluation result to obtain a compressed video data packet including a first key data segment including vehicle trajectory, pedestrian behavior, and traffic signs;
[0013] S04. Using a multimodal deep learning model to perform parallel processing on the compressed video data packet, extracting structured traffic element information and generating a visual analysis result as the second key data segment;
[0014] S05. Based on real-time data traffic and processing throughput, dynamically configure storage paths and resource allocation strategies of distributed storage nodes to store the compressed video data packet and the second key data segment corresponding to the timestamp of the first key data segment.
[0015] Preferably, the acquisition of the video data packet in step S01 includes:
[0016] S11, controlling the camera through the edge computing node to collect video streams at variable time intervals, where the time intervals are dynamically adjusted according to light intensity;
[0017] S12, performing motion region detection on the original video stream using a frame difference method to segment video segments containing valid traffic events into data packets;
[0018] S13, forming a data packet input from the video clip to build an invalid frame filtering model based on YOLOv5, identifying and eliminating completely black frames, still frames and repeated frames, and generating a pre-processed data packet index.
[0019] Preferably, the step S01 of dividing the coverage area into a plurality of data blocks includes:
[0020] S14, extracting a multi-dimensional feature vector from the video data packet, wherein the multi-dimensional feature vector includes a vehicle density gradient, a pedestrian movement heat map, and a traffic sign distribution matrix;
[0021] S15. Apply the DBSCAN clustering algorithm with spatiotemporal constraints and integrate GIS coordinate data and time period weight factors , where t is the length of traffic collection time, T m is the midpoint of the peak traffic period, N e is the number of historical accidents in the current period, and L is the road segment grade coefficient, which is used to generate characteristic blocks with similar traffic patterns;
[0022] S16. Establish a dynamic block update mechanism to trigger block re-division when the feature vectors of three consecutive frames deviate from the cluster center threshold.
[0023] Preferably, the data redundancy assessment in step S02 includes:
[0024] S21, using an improved ORB feature point extraction algorithm to perform multi-scale spatial pyramid feature matching on the key frame;
[0025] S22. Construct a spatiotemporal similarity calculation model, integrating HSV histogram intersection-over-union ratio, SIFT descriptor cosine similarity, and optical flow motion vector similarity;
[0026] S23. Design a time series analysis module based on LSTM to predict the redundancy change trend of the next five frames and generate a dynamic redundancy evaluation coefficient matrix.
[0027] Preferably, the adaptive compression processing in step S3 includes:
[0028] S31, creating a priority queue of the key frames, and setting a retention weight based on the vehicle license plate recognition confidence and the pedestrian posture abnormality index;
[0029] S32: Based on the deployment of a configurable hybrid compression engine, H.265 inter-frame prediction coding is enabled for high-redundancy blocks, and AV1 intra-frame coding is used for low-redundancy blocks;
[0030] S33, based on embedding a SSIM-based quality feedback loop, automatically switches the encoding mode when the structural similarity after compression is lower than 0.95.
[0031] Preferably, the step S04 of using a multimodal deep learning model to perform parallel processing on the compressed video data packets includes:
[0032] S41, build a multi-task deep learning network, the backbone network uses , the parallel output branches include:
[0033] Vehicle detection branch: Vehicle type classification and trajectory prediction based on the improved CenterNet;
[0034] Pedestrian analysis branch: Identify abnormal behaviors through OpenPose skeleton extraction combined with spatiotemporal graph convolutional networks;
[0035] Traffic sign recognition branch: Applying attention mechanism to enhance small target detection capabilities;
[0036] S42. Design a feature fusion module to perform cross-modal alignment of visual features with radar point cloud data and meteorological sensor data;
[0037] S43. Generate a structured traffic situation map, including a dynamic lane-level flow heat map and an event spatiotemporal distribution matrix.
[0038] Preferably, the cross-modal alignment in step S42 includes:
[0039] S421, performing spatiotemporal synchronization based on Kalman filtering on the radar point cloud data;
[0040] S422. Matching the meteorological data acquired by the meteorological sensor data with the video data through a sliding window correlation coefficient;
[0041] S423. Establish a cross-modal attention map, the calculation formula is , where Q v is the visual feature query vector, K r is the radar feature key vector, and d is the dimension scaling factor.
[0042] Preferably, the dynamic adjustment of the storage strategy in step S05 includes:
[0043] S51. Real-time calculation of storage cost factors Energy consumption, where: α is the access delay, β is the storage density, and γ is the energy consumption penalty coefficient;
[0044] S5, based on The storage path optimization algorithm automatically selects local SSD, distributed object storage, or edge cache based on the hot and cold characteristics of data;
[0045] S53. Implement a two-layer redundancy elimination mechanism, perform BloomFilter-based deduplication at the edge node, and implement global deduplication based on content signatures at the central cloud.
[0046] The present invention has at least the following beneficial effects:
[0047] 1. By integrating multi-dimensional feature vectors such as vehicle density gradients and pedestrian heat maps, combined with the spatiotemporal-constrained DBSCAN clustering algorithm, this method achieves intelligent dynamic segmentation of traffic scenarios. Compared to traditional fixed segmentation methods, this method automatically triggers segmentation when traffic flow changes suddenly, precisely matching computing resources to the changing characteristics of each region. This effectively addresses the resource waste associated with traditional methods due to their poor scenario adaptability.
[0048] 2. A multidimensional redundancy assessment system is constructed using an improved ORB feature point extraction and spatiotemporal similarity calculation model, integrated with cross-dimensional analysis of HSV histograms, SIFT descriptors, and optical flow motion vectors, and combined with an LSTM prediction module. Compared to single-pixel comparison methods, this improves the accuracy of redundant data identification. Combined with a hybrid compression engine and a two-layer deduplication mechanism, storage requirements are reduced while retaining key information, such as license plates and abnormal behavior, with greater integrity.
[0049] 3. Dynamic optimization of compression parameters is achieved through the coordinated control of the SSIM quality feedback loop and the priority queue. At the same compression rate, this solution improves the structural similarity index by 0.12 compared to traditional H.265 encoding, significantly improving vehicle license plate recognition rates. Especially for low-light scenarios, edge nodes can adaptively switch encoding modes to ensure the integrity of keyframe information in nighttime videos, resolving the nighttime data distortion issue caused by fixed encoding in existing technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0051] Figure 1 A flowchart of a method for rapid processing of big data based on smart cities provided in the first embodiment of the present invention;
[0052] Figure 2 This is a flowchart of S01 provided in the first embodiment of the present invention;
[0053] Figure 3 This is a flowchart of S02 provided in the first embodiment of the present invention;
[0054] Figure 4 This is a flowchart of S03 provided in the first embodiment of the present invention;
[0055] Figure 5 This is a flowchart of S04 provided in the first embodiment of the present invention. DETAILED DESCRIPTION
[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0057] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0058] Example 1
[0059] This embodiment provides a method for rapid processing of big data based on smart city, which includes the following steps: Figure 1 As shown:
[0060] S01. Obtaining video data packets collected by urban road cameras within a preset time period, and extracting image content features reflecting traffic dynamics from the video data packets, wherein the image content features divide the coverage area into multiple data blocks;
[0061] Specific, combined Figure 2 As shown, the acquisition of the above-mentioned video data packet includes:
[0062] S11, controlling the camera through the edge computing node to collect video streams at variable time intervals, where the time intervals are dynamically adjusted according to light intensity;
[0063] S12, using a frame difference method to perform motion region detection on the original video stream to segment video segments containing valid traffic events to form data packets;
[0064] S13, forming a data packet input from the video clip to build an invalid frame filtering model based on YOLOv5, identifying and eliminating completely black frames, still frames and repeated frames, and generating a preprocessing data packet index.
[0065] As mentioned above, through edge computing nodes, the system can intelligently control the camera's video collection interval. This interval is no longer fixed, but is dynamically adjusted based on actual scene conditions such as light intensity. For example, during daytime when there is ample light, the system can appropriately extend the collection interval to reduce data volume; while at night or in low light conditions, the collection interval can be shortened to ensure image clarity, thereby optimizing resource utilization. Furthermore, the use of "variable interval video stream collection" effectively avoids the resource waste (e.g., collecting large amounts of useless data) and the loss of critical information (e.g., excessively long collection intervals during important events) that can occur with traditional fixed-interval collection.
[0066] Secondly, the system uses frame differencing to detect motion regions in the raw video stream. By comparing the differences between consecutive frames, it can accurately identify areas of motion within the image. Based on the results of motion region detection, it can automatically segment video clips containing valid traffic events. These clips typically contain important traffic information, such as vehicles in motion and pedestrians crossing the road.
[0067] Furthermore, the YOLOv5 model intelligently analyzes video frames, accurately identifying and removing invalid frames such as completely black frames, still frames, and repeated frames. These invalid frames typically contain no useful traffic information, and removing them significantly reduces the amount of data required for subsequent processing. After removing invalid frames, the system generates a pre-processed data packet index. This index records information such as the location and length of valid video segments, facilitating subsequent processing.
[0068] Furthermore, the method for dividing the coverage area into a plurality of data blocks in the above embodiment includes:
[0069] S14, extracting a multi-dimensional feature vector from the video data packet, the multi-dimensional feature vector including a vehicle density gradient, a pedestrian movement heat map, and a traffic sign distribution matrix;
[0070] S15. Apply the DBSCAN clustering algorithm with spatiotemporal constraints and integrate GIS coordinate data and time period weight factors , where t is the length of traffic collection time, T m is the midpoint of the peak traffic period, N e is the number of historical accidents in the current period, and L is the road segment grade coefficient, which is used to generate characteristic blocks with similar traffic patterns;
[0071] S16. Establish a dynamic block update mechanism to trigger block re-division when the feature vectors of three consecutive frames deviate from the cluster center threshold.
[0072] The above-mentioned multi-dimensional feature vectors are extracted from the video data packets. These vectors contain rich information about the video content, such as vehicle density and pedestrian movement trajectories. Using the extracted feature vectors, the system applies the DBSCAN clustering algorithm with spatiotemporal constraints to divide the video image into multiple feature blocks with similar traffic patterns. This helps the system perform targeted processing based on the characteristics of different areas and improve processing efficiency. For example, one area may be dominated by straight-moving vehicles, while another area may be dominated by turning vehicles. The system can divide these areas into different feature blocks through feature vector extraction and clustering algorithms, and optimize processing based on the characteristics of each block.
[0073] Secondly, when the feature vectors of three consecutive frames deviate from the cluster center threshold, a block re-division is triggered, ensuring that the system can adapt to changes in the traffic scene in real time and maintain efficient processing. By dynamically updating the block division, the system can continue to maintain efficient and accurate processing of traffic scenes, even when traffic patterns change significantly.
[0074] S02. Extract key frames from the video data packet for each data block, calculate the spatiotemporal similarity of adjacent key frames, and generate a data redundancy evaluation result by combining time series analysis;
[0075] Specific, combined Figure 3 As shown, the above-mentioned data redundancy evaluation method includes:
[0076] S21. Use the improved ORB feature point extraction algorithm to perform multi-scale spatial pyramid feature matching on the key frames;
[0077] S22. Construct a spatiotemporal similarity calculation model, integrating HSV histogram intersection-over-union ratio, SIFT descriptor cosine similarity, and optical flow motion vector similarity;
[0078] S23. Design a time series analysis module based on LSTM to predict the redundancy change trend of the next five frames and generate a dynamic redundancy evaluation coefficient matrix.
[0079] The above-mentioned key frames are accurately extracted from the video data packets. These key frames usually contain important information about the video content. By calculating the spatiotemporal similarity of adjacent key frames, the system can accurately assess the redundancy of the video data. This assessment provides a scientific basis for subsequent compression processing, helping to reduce the amount of data while ensuring video quality. For example, in a video, if the image content of multiple consecutive frames is extremely similar, the system can identify this feature through spatiotemporal similarity calculation. In this way, in the subsequent compression processing, the system can take corresponding optimization measures, such as increasing the compression ratio or skipping the processing of certain frames, thereby further improving processing efficiency.
[0080] Furthermore, the generation of the above-mentioned LSTM-based time series analysis module includes:
[0081] Data input:
[0082] The LSTM-based time series analysis module receives feature vectors of video data as input. These feature vectors may include visual features such as brightness, color, and texture of the video frame, as well as processing features such as the compression ratio and encoding method of the video data.
[0083] Network training:
[0084] During the training phase, the LSTM network learns the changing patterns of redundancy from a large amount of historical video data. The network adjusts its internal parameters based on the input feature vector and the corresponding redundancy label (e.g., high redundancy, low redundancy) to minimize prediction error.
[0085] Trend Forecast:
[0086] During the prediction phase, the LSTM network receives the feature vector of the current video frame and, based on its internal state and learned patterns, predicts the redundancy trend for the next five frames. The prediction results are output as a redundancy trend curve, visually demonstrating how redundancy changes over time.
[0087] Dynamic redundancy assessment coefficient matrix generation
[0088] Redundancy quantification:
[0089] Based on the redundancy trend predicted by the LSTM network, the system quantifies the redundancy into specific numerical values. These values reflect the degree of redundancy of the video data at different points in time, with higher values indicating greater redundancy.
[0090] Matrix construction:
[0091] The system arranges the quantized redundancy values in chronological order to construct a dynamic redundancy evaluation coefficient matrix. The rows of this matrix represent different time points (frames), and the columns represent the dimensions of redundancy evaluation (such as visual feature redundancy and processing feature redundancy).
[0092] Dynamic Adjustment:
[0093] As video data is continuously input and processed, the dynamic redundancy evaluation coefficient matrix is updated in real time. The system adjusts the video data processing strategy, such as compression ratio and encoding method, based on the latest redundancy evaluation results to optimize processing efficiency and quality.
[0094] Therefore, when processing a traffic surveillance video, the LSTM-based time series analysis module predicts that the redundancy of the next five frames will initially increase and then decrease. Based on this prediction, the system dynamically adjusts the compression strategy: during periods of increasing redundancy, a more efficient compression algorithm is used to reduce the data volume; during periods of decreasing redundancy, the compression ratio is appropriately reduced to preserve more video details. The system also outputs the redundancy assessment results in matrix form for subsequent processing and analysis.
[0095] S03, performing adaptive compression processing on the video data packet according to the data redundancy evaluation result to obtain a compressed video data packet containing a first key data segment including vehicle trajectory, pedestrian behavior and traffic signs;
[0096] Specific, combined Figure 4 As shown, the adaptive compression process includes:
[0097] S31, creating a priority queue of key frames, and setting retention weights based on the vehicle license plate recognition confidence and pedestrian posture abnormality index;
[0098] S32: Based on the deployment of a configurable hybrid compression engine, H.265 inter-frame prediction coding is enabled for high-redundancy blocks, and AV1 intra-frame coding is used for low-redundancy blocks;
[0099] S33, based on embedding a SSIM-based quality feedback loop, automatically switches the encoding mode when the structural similarity after compression is lower than 0.95.
[0100] During operation, the system creates a priority queue for key frames, assigning retention weights based on factors such as license plate recognition confidence and pedestrian posture anomaly index. This mechanism ensures that critical information is not lost during the compression process. A hybrid compression engine then employs adaptive compression of video data packets based on data redundancy assessment. For high-redundancy blocks, the engine uses efficient compression algorithms (such as H.265 inter-frame prediction coding); for low-redundancy blocks, it uses detail-preserving compression algorithms (such as AV1 intra-frame coding). Finally, an embedded SSIM quality feedback loop automatically switches encoding modes when the post-compression structural similarity falls below a threshold, ensuring that the compressed video quality meets application requirements. This means that when processing a video containing a moving vehicle and a pedestrian crossing the road, the system can identify key information such as the license plate and pedestrian posture, and ensure that this information is not lost during the compression process through the priority queue. Furthermore, the system uses efficient compression algorithms for high-redundancy areas (such as static backgrounds) and detail-preserving compression algorithms for low-redundancy areas containing critical information.
[0101] S04 uses a multimodal deep learning model to process compressed video data packets in parallel, extracting structured traffic element information and generating visual analysis results as the second key data segment;
[0102] Further, combined Figure 5 As shown, the parallel processing of compressed video data packets using a multimodal deep learning model includes:
[0103] S41, build a multi-task deep learning network, the backbone network uses , the parallel output branches include:
[0104] Vehicle detection branch: Vehicle type classification and trajectory prediction based on the improved CenterNet;
[0105] Pedestrian analysis branch: Identify abnormal behaviors through OpenPose skeleton extraction combined with spatiotemporal graph convolutional networks;
[0106] Traffic sign recognition branch: Applying attention mechanism to enhance small target detection capabilities;
[0107] S42. Design a feature fusion module to perform cross-modal alignment of visual features with radar point cloud data and meteorological sensor data;
[0108] S43. Generate a structured traffic situation map, including a dynamic lane-level flow heat map and an event spatiotemporal distribution matrix.
[0109] Secondly, the above cross-modal alignment includes:
[0110] S421, implementing spatiotemporal synchronization of radar point cloud data based on Kalman filtering;
[0111] S422, matching the meteorological data acquired by the meteorological sensor data with the video data through a sliding window correlation coefficient;
[0112] S423. Establish a cross-modal attention map, the calculation formula is , where Q v is the visual feature query vector, K r is the radar feature key vector, and d is the dimension scaling factor.
[0113] Specifically, a multi-task deep learning network was built to The system serves as the backbone network, concurrently processing multiple tasks such as vehicle detection, pedestrian analysis, and traffic sign recognition. A feature fusion module aligns visual features with radar point cloud data and meteorological sensor data across modalities, improving detection accuracy and robustness. This ultimately generates a structured traffic situation map, including visualizations such as dynamic lane-level traffic flow heat maps and event spatiotemporal distribution matrices, providing powerful support for traffic management.
[0114] Secondly, Kalman filtering synchronizes radar point clouds and video data, combined with sliding window correlation coefficients to match meteorological information, reducing spatiotemporal alignment errors in multi-source data. Trajectory prediction errors were also reduced in rainy and snowy weather tests. This approach addresses the "spatiotemporal asynchrony" issue caused by differences in sampling frequency and coordinate systems in multi-sensor data. For example, the high-precision ranging of radar point clouds and the low-latency visual information of videos are fused through attention maps to improve the robustness of target recognition in complex scenarios.
[0115] S05. Based on real-time data traffic and processing throughput, dynamically configure storage paths and resource allocation strategies of distributed storage nodes to store compressed video data packets and second key data segments corresponding to the timestamps of the first key data segments.
[0116] The above-mentioned dynamic adjustment of storage policies includes:
[0117] S51. Real-time calculation of storage cost factors Energy consumption, where α is the access latency (range: 0.8-1.2), β is the storage density (range: 0.5-0.9), and γ is the energy penalty coefficient (range: 0.3-0.7);
[0118] S5, based on The storage path optimization algorithm automatically selects local SSD, distributed object storage, or edge cache based on the hot and cold characteristics of data;
[0119] S53. Implement a two-layer redundancy elimination mechanism, perform BloomFilter-based deduplication at the edge node, and implement global deduplication based on content signatures at the central cloud.
[0120] Specifically, the storage cost factor is calculated in real time to guide resource allocation. A storage path optimization algorithm automatically selects storage options such as local SSD, distributed object storage, or edge caching based on data hot and cold characteristics. Furthermore, a two-tier redundancy elimination mechanism is implemented: BloomFilter-based deduplication at edge nodes and global deduplication based on content signatures in the central cloud, further reducing storage costs and improving access performance.
[0121] This first embodiment achieves intelligent dynamic segmentation of traffic scenes by integrating multidimensional feature vectors such as vehicle density gradients and pedestrian heatmaps, combined with the spatiotemporal-constrained DBSCAN clustering algorithm. Compared to traditional fixed segmentation methods, this method automatically triggers block repartitioning when traffic flow changes suddenly, allowing computing resources to precisely match the changing characteristics of each region. This effectively addresses the resource waste caused by traditional methods' poor scene adaptability. Furthermore, it employs an improved ORB feature point extraction and spatiotemporal similarity calculation model, integrating cross-dimensional analysis of HSV histograms, SIFT descriptors, and optical flow motion vectors, and combining it with an LSTM prediction module to construct a multidimensional redundancy assessment system. Compared to single-pixel comparison methods, this method improves the accuracy of redundant data recognition. Combined with a hybrid compression engine and a two-layer deduplication mechanism, it reduces storage requirements while retaining more integrity of key information, such as license plates and abnormal behavior. Furthermore, through the coordinated control of the SSIM quality feedback loop and priority queue, dynamic optimization of compression parameters is achieved. At the same compression rate, this solution improves the structural similarity index by 0.12 compared to traditional H.265 encoding, significantly improving the vehicle license plate recognition rate. Especially for low-light scenes, edge nodes can adaptively switch encoding modes to ensure the integrity of nighttime video key frame information, solving the problem of nighttime data distortion caused by fixed encoding in existing technologies.
[0122] Example 2
[0123] An embodiment of the present invention provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by a processor to implement the steps:
[0124] Obtaining video data packets collected by urban road cameras within a preset time period, and extracting image content features reflecting traffic dynamics from the video data packets, wherein the image content features divide the coverage area into multiple data blocks;
[0125] Extract key frames from the video data packet for each data block, calculate the spatiotemporal similarity of adjacent key frames, and combine with time series analysis to generate data redundancy assessment results;
[0126] Adaptively compressing the video data packet according to the data redundancy evaluation result to obtain a compressed video data packet containing a first key data segment including vehicle trajectory, pedestrian behavior and traffic signs;
[0127] A multimodal deep learning model is used to process compressed video data packets in parallel, extracting structured traffic element information and generating visual analysis results as the second key data segment;
[0128] Based on real-time data traffic and processing throughput, storage paths and resource allocation strategies of distributed storage nodes are dynamically configured to store compressed video data packets and second key data segments corresponding to the timestamps of the first key data segments.
[0129] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware using a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media used in the various embodiments provided herein may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).
[0130] Those skilled in the art will clearly understand that for the sake of convenience and brevity in description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0131] Example 3
[0132] An embodiment of the present invention provides an electronic device, including a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the following steps:
[0133] Obtaining video data packets collected by urban road cameras within a preset time period, and extracting image content features reflecting traffic dynamics from the video data packets, wherein the image content features divide the coverage area into multiple data blocks;
[0134] Extract key frames from the video data packet for each data block, calculate the spatiotemporal similarity of adjacent key frames, and combine with time series analysis to generate data redundancy assessment results;
[0135] Adaptively compressing the video data packet according to the data redundancy evaluation result to obtain a compressed video data packet containing a first key data segment including vehicle trajectory, pedestrian behavior and traffic signs;
[0136] A multimodal deep learning model is used to process compressed video data packets in parallel, extracting structured traffic element information and generating visual analysis results as the second key data segment;
[0137] Based on real-time data traffic and processing throughput, storage paths and resource allocation strategies of distributed storage nodes are dynamically configured to store compressed video data packets and second key data segments corresponding to the timestamps of the first key data segments.
[0138] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as a preferred embodiment as above, it is not intended to limit the present invention. Any technician familiar with the present profession can make some changes or modifications to equivalent embodiments of the technical contents disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A method for rapid processing of big data based on smart cities, characterized in that: The method comprises the following steps: S01. Obtaining video data packets collected by urban road cameras within a preset time period, and extracting picture content features reflecting traffic dynamics from the video data packets, wherein the picture content features divide a coverage area into a plurality of data blocks; S02, extracting key frames from the video data packet for each data block, calculating the spatiotemporal similarity of adjacent key frames, and combining the time series analysis to generate a data redundancy evaluation result; S03, performing adaptive compression processing on the video data packet according to the data redundancy evaluation result to obtain a compressed video data packet including a first key data segment including vehicle trajectory, pedestrian behavior, and traffic signs; S04. Using a multimodal deep learning model to perform parallel processing on the compressed video data packet, extracting structured traffic element information and generating a visual analysis result as the second key data segment; S05. Based on real-time data traffic and processing throughput, dynamically configure storage paths and resource allocation strategies of distributed storage nodes to store the compressed video data packet and the second key data segment corresponding to the timestamp of the first key data segment.
2. The method for rapid processing of big data based on smart city according to claim 1 is characterized in that: The acquisition of the video data packet in step S01 includes: S11, controlling the camera through the edge computing node to collect video streams at variable time intervals, where the time intervals are dynamically adjusted according to light intensity; S12, performing motion region detection on the original video stream using a frame difference method to segment video segments containing valid traffic events into data packets; S13, forming a data packet input from the video clip to build an invalid frame filtering model based on YOLOv5, identifying and eliminating completely black frames, still frames and repeated frames, and generating a pre-processed data packet index.
3. The method for rapid processing of big data based on smart city according to claim 1, characterized in that: The step S01 of dividing the coverage area into a plurality of data blocks includes: S14, extracting a multi-dimensional feature vector from the video data packet, wherein the multi-dimensional feature vector includes a vehicle density gradient, a pedestrian movement heat map, and a traffic sign distribution matrix; S15. Apply the DBSCAN clustering algorithm with spatiotemporal constraints and integrate GIS coordinate data and time period weight factors , where t is the length of traffic collection time, T m is the midpoint of the peak traffic period, N e is the number of historical accidents in the current period, and L is the road segment grade coefficient, which is used to generate characteristic blocks with similar traffic patterns; S16. Establish a dynamic block update mechanism to trigger block re-division when the feature vectors of three consecutive frames deviate from the cluster center threshold.
4. The method for rapid processing of big data based on smart city according to claim 1, characterized in that: The data redundancy assessment in step S02 includes: S21, using an improved ORB feature point extraction algorithm to perform multi-scale spatial pyramid feature matching on the key frame; S22. Construct a spatiotemporal similarity calculation model, integrating HSV histogram intersection-over-union ratio, SIFT descriptor cosine similarity, and optical flow motion vector similarity; S23. Design a time series analysis module based on LSTM to predict the redundancy change trend of the next five frames and generate a dynamic redundancy evaluation coefficient matrix.
5. The method for rapid processing of big data based on smart city according to claim 1 is characterized in that: The adaptive compression processing in step S3 includes: S31, creating a priority queue of the key frames, and setting a retention weight based on the vehicle license plate recognition confidence and the pedestrian posture abnormality index; S32: Based on the deployment of a configurable hybrid compression engine, H.265 inter-frame prediction coding is enabled for high-redundancy blocks, and AV1 intra-frame coding is used for low-redundancy blocks; S33, based on embedding a SSIM-based quality feedback loop, automatically switches the encoding mode when the structural similarity after compression is lower than 0.
95.
6. The method for rapid processing of big data based on smart city according to claim 1 is characterized in that: The parallel processing of the compressed video data packets using a multimodal deep learning model in step S04 includes: S41, build a multi-task deep learning network, the backbone network uses , the parallel output branches include: Vehicle detection branch: Vehicle type classification and trajectory prediction based on the improved CenterNet; Pedestrian analysis branch: Identify abnormal behaviors through OpenPose skeleton extraction combined with spatiotemporal graph convolutional networks; Traffic sign recognition branch: Applying attention mechanism to enhance small target detection capabilities; S42. Design a feature fusion module to perform cross-modal alignment of visual features with radar point cloud data and meteorological sensor data; S43. Generate a structured traffic situation map, including a dynamic lane-level flow heat map and an event spatiotemporal distribution matrix.
7. The method for rapid processing of big data based on smart city according to claim 6 is characterized in that: The cross-modal alignment in step S42 includes: S421, performing spatiotemporal synchronization based on Kalman filtering on the radar point cloud data; S422. Matching the meteorological data acquired by the meteorological sensor data with the video data through a sliding window correlation coefficient; S423. Establish a cross-modal attention map, the calculation formula is , where Q v is the visual feature query vector, K r is the radar feature key vector, and d is the dimension scaling factor.
8. The method for rapid processing of big data based on smart city according to claim 1 is characterized in that: The dynamic adjustment of the storage strategy in step S05 includes: S51. Real-time calculation of storage cost factors Energy consumption, where: α is the access delay, β is the storage density, and γ is the energy consumption penalty coefficient; S5, based on The storage path optimization algorithm automatically selects local SSD, distributed object storage, or edge cache based on the hot and cold characteristics of data; S53. Implement a two-layer redundancy elimination mechanism, perform BloomFilter-based deduplication at the edge node, and implement global deduplication based on content signatures at the central cloud.
9. A non-transitory computer-readable storage medium, wherein at least one instruction or at least one program is stored in the non-transitory computer-readable storage medium, characterized in that: The at least one instruction or the at least one program is loaded and executed by the processor to implement the steps of the method for rapid processing and operation of big data based on a smart city as described in any one of claims 1-8.
10. An electronic device, characterized in that: It includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the steps of the method for rapid processing and operation of big data based on smart cities as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Method for identifying traffic jam evolution law based on gridding and space-time clustering
CN110750573A
Vehicle moving track data compression method and system and electronic equipment
CN119135185A
Road traffic signal lamp management system based on big data
CN119694146A
Urban road traffic flow detection method and system
CN120148248A
System and method for edge based multi-modal homomorphic compression
US20250190339A1
Cited By
Deep learning-based computing power center edge computing task scheduling method and system
CN121233252A