Multi-protocol communication-based smart city multi-scene monitoring method and system
By deploying heterogeneous terminals and edge servers across the entire smart city for lightweight data processing, combined with LoRa and high-bandwidth selective transmission, the problems of data transmission latency and low analysis efficiency in traditional smart city monitoring systems have been solved, enabling efficient and secure multi-scenario monitoring.
Patent Information
- Application Number
- CN202510858461.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Traditional smart city monitoring systems suffer from problems such as large data transmission latency, high server storage pressure, and low data analysis efficiency.
Heterogeneous terminals are deployed throughout the city. The monitoring data is preprocessed in real time by the edge server to generate lightweight summary data, which is then transmitted to the central server using LoRa communication units. The central server performs preliminary event screening and triggers high-bandwidth data transmission only when preset events are identified. Event verification is performed by combining a multi-level collaborative processing architecture and a weighted probability model.
It reduces network load and storage pressure, improves event analysis efficiency, reduces false alarm rate, and ensures privacy compliance during transmission, achieving cost-effective multi-scenario monitoring support.
Smart Images

Figure CN120529272B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of communication, in particular to a smart city multi-scene monitoring method and system based on multi-protocol communication. BACKGROUND
[0002] With the continuous advancement of smart city construction, various aspects of the city are gradually included in the management category of intelligent monitoring systems. These systems are not only used for security monitoring of the city, but also involve environmental protection, traffic management, infrastructure maintenance and other fields. Traditional monitoring systems usually rely on a single type of terminal device for data collection, and data transmission mainly relies on centralized servers. This architecture has problems such as large data transmission delay and high server storage pressure.
[0003] A similar prior art is Chinese patent application CN114518960A, which relates to a data preprocessing method of an Internet of Things edge gateway. The method includes: configuring a data preprocessing module through a remote server; the data preprocessing module includes an abnormal data processing unit, a stable data processing unit and a local data processing unit; the remote server issues the data preprocessing module to the edge gateway; the edge gateway stores the data preprocessing module to the local storage; the edge gateway preprocesses the data reported by the Internet of Things terminal device according to the data preprocessing module; the edge gateway stores the preprocessed data to the local storage and simultaneously reports to the remote server. However, this invention only considers saving transmission bandwidth and server resources, and does not consider the analysis efficiency of data after transmission.
[0004] A similar prior art is Chinese patent application CN119603650A, which discloses a key area intelligent monitoring system based on communication towers, relating to the fields of communication technology, IoT and intelligent monitoring. It includes: each communication tower is provided with a front-end sensing layer; the front-end sensing layer is used to obtain sensor data of the key area and preprocess the sensor data to generate preprocessed data; the channel transmission layer is used to transmit the preprocessed data to the data processing and analysis layer; the data processing and analysis layer is used to generate an adaptive abnormal behavior recognition model based on a machine learning model library, and to process the preprocessed data according to the adaptive abnormal behavior recognition model, analyze abnormal behavior, and generate data analysis results; the application service layer is used to transmit the data analysis results to the mobile terminal and formulate an early warning strategy according to the abnormal behavior in different scenarios; the early warning strategy includes a warning threshold and a response rule. However, this application does not consider the transmission efficiency and analysis efficiency of data when the data volume is large. SUMMARY
[0005] To solve the above technical problems, the application provides a smart city multi-scene monitoring method and system based on multi-protocol communication, which can reduce network load and storage pressure and improve event analysis efficiency.
[0006] In a first aspect, the application provides a smart city multi-scene monitoring method based on multi-protocol communication, which comprises:
[0007] Deploying heterogeneous terminals in the whole city and collecting monitoring data of the terminals in real time and sending them to an edge server;
[0008] The edge server processes the monitoring data from multiple terminals to generate summary data and transmits it to a central server through a LoRa communication unit, wherein the summary data includes video summary data, audio summary data and text summary data, and comes from different terminals respectively;
[0009] After the central server receives the summary data from multiple edge servers, it determines whether there is a preset event in the summary data, and if so, requests the edge server to send target monitoring data of the corresponding time period of the summary data;
[0010] The edge server transmits the target monitoring data to the central server through a 4G communication unit or a 5G communication unit;
[0011] The central server verifies the event in combination with the target monitoring data, and after determining the event, generates alarm information of the event and feeds it back to the relevant management system.
[0012] In combination with the first aspect, in a first implementation manner of the first aspect of the application, the summary data is generated, comprising:
[0013] Obtaining video data from the monitoring data, and extracting multiple object contours and background images from each frame of image of the video data, wherein the video data is any one or several segments of the monitoring data;
[0014] Comparing multiple object contours in the video data, identifying the same object, generating the same object in the same image as an object image, and generating a unique identifier for the object image, wherein the object image includes the contour, position, size and direction of the object;
[0015] Analyzing the characteristics and behaviors of the object image to generate video summary data of the video data, wherein the video summary data includes terminal ID, time, position, size, speed, direction and motion trajectory of the object in the video data;
[0016] The video summary data is stored together with the corresponding video data and is associated with the background image and the object image.
[0017] In a second implementation form of the first aspect, the extracting the plurality of object contours and the background image comprises:
[0018] An i-th frame image is obtained from the video data, and the i-th frame image is processed based on a background modeling algorithm to obtain an i-th background image, where i represents a positive integer from 1 to A, and A represents a total number of frames of the video data;
[0019] When a background image similarity between the i-th background image and an (i+1)-th frame image following the i-th frame image is greater than or equal to a first threshold value, pixels of the i-th frame image are compared with pixels of the (i+1)-th frame image to detect a changed object contour;
[0020] When the background image similarity between the i-th background image and the (i+1)-th frame image is less than the first threshold value, an (i+1)-th background image is obtained from the (i+1)-th frame image, and when a background image similarity between the (i+1)-th background image and an (i+2)-th frame image following the (i+1)-th frame image is greater than or equal to the first threshold value, pixels of the (i+1)-th frame image are compared with pixels of the (i+2)-th frame image to detect a changed object contour;
[0021] Similarly, until all changed object contours and background images are detected, the object contours include shapes, sizes, positions and motion directions of objects.
[0022] In a third implementation form of the first aspect, the identifying the same object comprises:
[0023] An object feature is extracted based on the object contour, and the object feature includes a geometric shape, a color, a texture, an edge and a corner point of the object;
[0024] The (i+1)-th object contour in the (i+1)-th frame image is compared with the i-th object contour in the i-th frame image, and when a similarity between the (i+1)-th object contour and the i-th object contour is greater than or equal to a second threshold value, it is determined that the (i+1)-th object contour and the i-th object contour belong to the same object, otherwise, it is determined that the (i+1)-th object contour and the i-th object contour belong to different objects.
[0025] In a fourth implementation form of the first aspect, the identifying the same object further comprises:
[0026] acquire a first location where a first terminal of the video data source is located, acquire other video collection terminals closest to the first location, and acquire video data of the other video collection terminals; when a similarity between an object contour in the video data of the other video collection terminals and an object contour in the video data is greater than or equal to the first threshold, merge the video data of the other video collection terminals and the video data into one video data, and regenerate video summary data of the video data.
[0027] With reference to the first aspect, in a fifth implementation form of the first aspect of the application, the determining whether the summary data contains a preset event comprises:
[0028] Based on the summary data and a pre-trained event prediction model, an event corresponding to the summary data is predicted, and an event label is added to the summary data, wherein the event refers to a specific behavior model, and at least includes an abnormal person entering a restricted area, a vehicle reversing, a rapid gathering, a fire, an explosion and a fight.
[0029] With reference to the first aspect, in a sixth implementation form of the first aspect of the application, the generating the summary data further comprises:
[0030] Acquiring audio data from the monitoring data, analyzing the audio data, generating audio summary data, the audio summary data including a terminal ID, a time of sound occurrence and a content description of sound prediction;
[0031] Acquiring sensor data from the monitoring data, pre-processing the sensor data, generating text summary data, the text summary data including a terminal ID, a recording time and a content description.
[0032] With reference to the first aspect, in a seventh implementation form of the first aspect of the application, the verifying the event comprises:
[0033] Acquiring first monitoring data and second monitoring data of other two terminals closest to a location where a terminal of a source of the target monitoring data is located, if the target monitoring data is video data, the first monitoring data represents audio data, and the second monitoring data represents text data, if the target monitoring data is audio data, the first monitoring data represents video data, and the second monitoring data represents text data, if the target monitoring data is text data, the first monitoring data represents video data, and the second monitoring data represents audio data;
[0034] Supposing that contribution weights of video data, audio data and text data to an event occurrence probability are w1, w2 and w3 respectively, and the event occurrence probability is calculated by a formula Calculate a probability P of an event occurring, and determine that the event occurs when the probability P of the event occurring is greater than a preset probability, or ignore the event otherwise, wherein T1 represents video data, T2 represents audio data, T3 represents text data, w1 is greater than w2, and w2 is greater than w3.
[0035] With reference to the first aspect, in an eighth implementation form of the first aspect of the present application, generating the summary data further includes:
[0036] When the summary data is generated, if sensitive information exists in the summary data, the sensitive information is covered by encoding or text.
[0037] In a second aspect, the present application provides a smart city multi-scene monitoring system based on multi-protocol communication, which comprises:
[0038] A data acquisition unit is configured to deploy heterogeneous terminals in the whole city and collect monitoring data of the terminals in real time and send the monitoring data to an edge server;
[0039] An edge data processing unit is configured to process the monitoring data from multiple terminals to generate summary data and transmit the summary data to a central server through a LoRa communication unit, wherein the summary data comprises video summary data, audio summary data and text summary data, and the summary data is from different terminals respectively;
[0040] A service data processing unit is configured to determine whether a preset event exists in the summary data after the central server receives the summary data from multiple edge servers, and if the preset event exists, request the edge server to send target monitoring data corresponding to a time period of the summary data;
[0041] A data transmission unit is configured to transmit the target monitoring data to the central server through a 4G communication unit or a 5G communication unit.
[0042] An event verification unit is configured to verify the event in combination with the target monitoring data, and generate alarm information of the event and feed back the alarm information to a related management system after determining the event.
[0043] Compared with the prior art, the present application has at least the following advantages:
[0044] The technical scheme provided in the application is characterized in that: multi-dimensional data acquisition is realized by deploying heterogeneous terminals in the whole city, original monitoring data is preprocessed in real time by an edge server to generate lightweight summary data, the transmission bandwidth requirement is reduced by using the transmission characteristics of a LoRa low-power wide-area network, and the network congestion problem is effectively alleviated; secondly, the central server performs preliminary event screening based on the summary data, and only when a preset event is identified, a 4G / 5G channel with high bandwidth is triggered to request target monitoring data, the selective transmission mechanism can significantly reduce the core network traffic, and can effectively compress the server storage load; in terms of analysis efficiency, a multi-level cooperative processing architecture is adopted, the edge side completes object contour extraction, cross-terminal target association and behavior feature coding, the central side verifies the multi-modal data by using a weighted probability model, the event analysis response speed is improved, and the false alarm rate is reduced; in addition, a dynamic privacy protection mechanism is introduced, sensitive information such as faces and license plates is automatically covered in the summary data generation stage, and secondary desensitization processing is performed during target data transmission, so that the data usability is ensured, and the privacy compliance requirement is met. Finally, the transmission efficiency, analysis accuracy, resource consumption and safety compliance of the monitoring system are comprehensively improved, and high-performance multi-scene monitoring support is provided for smart cities. BRIEF DESCRIPTION OF DRAWINGS
[0045] In order to more clearly illustrate the technical solutions of the embodiments of the application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor based on these drawings.
[0046] Figure 1 An embodiment schematic diagram of the smart city multi-scene monitoring method based on multi-protocol communication in the embodiments of the application;
[0047] Figure 2 A structural schematic diagram of the terminal, edge server and central server in the embodiments of the application;
[0048] Figure 3 A schematic diagram of the method for extracting multiple object contours and background images in the embodiments of the application;
[0049] Figure 4 An embodiment schematic diagram of the smart city multi-scene monitoring system based on multi-protocol communication in the embodiments of the application. DETAILED DESCRIPTION
[0050] The embodiments of the present application provide a smart city multi-scene monitoring method and system based on multi-protocol communication. The terms "first", "second", "third", "fourth" and the like (if any) in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" or "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0051] For ease of understanding, the specific process of the embodiments of the present application is described below. Please refer to Figure 1 One embodiment of the smart city multi-scene monitoring method based on multi-protocol communication in the embodiments of the present application includes:
[0052] Step S1: Deploying heterogeneous terminals in the whole city, and collecting the monitoring data of the terminals in real time and sending them to the edge server.
[0053] Specifically, a plurality of types of monitoring terminals are arranged in the whole city, such as Figure 2 The first terminal, the second terminal and the third terminal shown in the figure can be a camera, an audio sensor, an environmental sensor, etc., respectively, to collect multi-modal original monitoring data in real time, including video, audio and text sensor data, and uniformly transmit them to the edge server. The core is to break the data limitation of traditional single terminal, realize the whole coverage and stereoscopic data collection of the city in multiple scenes, and provide rich data basis for subsequent analysis.
[0054] Step S2: The edge server processes the monitoring data from the plurality of terminals to generate summary data, and transmits the summary data to the central server through the LoRa communication unit. The summary data includes video summary data, audio summary data and text summary data, and is respectively from different terminals.
[0055] Specifically, the edge server performs lightweight processing on the original data, including for video data, using an AI algorithm to extract a background image and a moving object contour frame by frame, and the specific extraction of the background image and the object contour will be described below; cross-frame object tracking is also achieved through feature matching technology to generate video summary data containing object trajectory, speed and other key information; for audio and sensor data, corresponding summary data is generated through voiceprint analysis and text summary, respectively, which is described in detail above; this method can significantly compress the data volume, and at the same time, through a dynamic desensitization mechanism, privacy security can also be ensured during data compression.
[0056] Step S3: After the central server receives the summary data from the multiple edge servers, it determines whether there is a preset event in the summary data. If there is, it requests the edge server to send target monitoring data corresponding to the time period of the summary data.
[0057] Specifically, after the central server receives the summary data of the multi-edge node, it identifies a preset event label based on a machine learning model, such as vehicle reverse, fire, etc. If a potential event is detected, a target monitoring data request is immediately initiated to the edge server. The core advantage of this method is that it uses low-bandwidth LoRa to transmit summary data and only requests high-precision raw data when a high-probability event is triggered, which can effectively reduce network transmission load and also avoid the central server processing a large amount of invalid data.
[0058] Step S4: The edge server transmits the target monitoring data to the central server through a 4G communication unit or a 5G communication unit.
[0059] Specifically, after receiving the request, the edge server automatically selects a 4G / 5G high-speed channel to transmit the raw monitoring data in the target time period, and synchronously activates a secondary privacy protection process to cover sensitive information in the target data. This method ensures data integrity while optimizing network resource allocation through differentiated protocol selection, such as using low-power LoRa for transmitting summary data and high-speed cellular networks for transmitting large amounts of data.
[0060] Step S5: The central server verifies the event in combination with the target monitoring data and generates an alarm information of the event and feeds back to the related management system after determining the event.
[0061] Specifically, the central server combines the target monitoring data and auxiliary data of adjacent terminals, such as video events supplemented by audio or text verification, and calculates the event occurrence probability through a weighted probability model with video weight w1> audio w2> text w3. If the probability exceeds the threshold, an alarm information is generated and pushed to the management system. This method can effectively reduce the event false positive rate and improve the event judgment accuracy through multi-modal data cross verification.
[0062] The present application can reduce network load and storage pressure while improving event analysis efficiency through three technical linkages of data hierarchical compression, protocol on-demand calling, and multi-modal collaborative verification, achieving the core goal of "low load, high precision, and fast response" of the smart city monitoring system.
[0063] It can be understood that the execution subject of the present application can be a smart city multi-scene monitoring system based on multi-protocol communication, and can also be a terminal or a server, which is not limited here. The server is taken as an example for illustration in the embodiments of the present application.
[0064] In a specific embodiment, the process of generating summary data can specifically include the following steps:
[0065] Obtaining video data from the monitoring data, and extracting a plurality of object contours and background images from each frame of image of the video data, the video data being any one or several segments in the monitoring data; comparing the plurality of object contours in the video data, identifying the same objects, generating the identified same objects in the same image as object images, and generating unique identifiers for the object images, the object images including the contours, positions, sizes and directions of the objects; analyzing the features and behaviors of the object images, and generating video summary data of the video data, the video summary data including the terminal ID, the time, position, size, speed, direction and motion trajectory of the objects displayed in the video data; wherein the video summary data is stored together with the corresponding video data, and is associated with the background images and the object images.
[0066] Specifically, first, an AI algorithm is used to separate the static background and dynamic object contours frame by frame, detect moving targets by comparing the pixel changes of adjacent frames, and extract contour information containing object shape, position and motion direction; then, cross-frame object tracking is performed based on geometric features such as color, texture, edge, etc., to fuse the multiple frame contours of the same object into a single object image with a unique identifier, and record the motion trajectory completely; finally, the object behavior characteristics such as speed and direction change are analyzed, and structured video summary data is generated, including terminal ID, timestamp, position coordinates, size change, motion vector and trajectory path, etc. Key metadata, and the summary data is associated with the original video, background image and object image to establish an index; for non-video data, audio is analyzed synchronously to generate a voiceprint summary, and sensor data is converted into a text summary to form a multi-modal lightweight data set. By converting the original video into summary data, the data volume is reduced to effectively reduce the transmission load; at the same time, the edge server completes data distillation and feature purification, so that the central server only needs to process a small amount of high-value data, which can effectively improve the initial screening response speed of events.
[0067] In a specific embodiment, referring to Figure 3 The process of extracting a plurality of object contours and background images can specifically include the following steps:
[0068] An i-th image is acquired from the video data, and the i-th image is processed based on a background modeling algorithm to acquire an i-th background image, i represents a positive integer from 1 to A, and A represents a total frame number of the video data; when a background image similarity between the i-th background image and an i+1-th image after the i-th image is greater than or equal to a first threshold value, pixels of the i-th image are compared with pixels of the i+1-th image to detect a changed object contour; when the background image similarity between the i-th background image and the i+1-th image after the i-th image is less than the first threshold value, an i+1-th background image is acquired from the i+1-th image, and when a background image similarity between the i+1-th background image and an i+2-th image after the i+1-th image is greater than or equal to the first threshold value, pixels of the i+1-th image are compared with pixels of the i+2-th image to detect a changed object contour; similarly, until all changed object contours and background images are detected, the object contour includes a shape, a size, a position and a moving direction of the object.
[0069] Specifically, first, the IA algorithm such as Gaussian mixture model is applied to the i-th frame of the video to generate a reference background image as a reference template of the static elements of the scene; then, the similarity between the reference background and the i+1-th background image is calculated, and the similarity can be measured by using the structural similarity index SSIM and the like, and if the similarity reaches a preset threshold value, for example, the preset threshold value is greater than or equal to 0.9, it indicates that the scene background is stable, at this time, the pixel difference between the i-th frame and the i+1-th frame is directly compared, and the changed region caused by the object movement is extracted as the contour data; when the similarity is lower than the threshold value, for example, sudden light change or lens switching, the i+1-th frame is set as a new reference background, and the second verification is delayed to the i+2-th frame: only when the similarity between the new reference and the background of the subsequent frame meets the standard, the pixel comparison is performed to detect the contour, which can effectively avoid the false extraction caused by the instantaneous interference; the double background verification method can reduce the contour false detection rate in complex environments such as sudden change of sunshine and lens shaking, and improve the extraction accuracy; finally, the detected region is subjected to morphological optimization and edge fitting to generate a structured data set containing the geometric shape, pixel-level size, center coordinates and motion direction vector of the object, and a quantifiable object motion fingerprint library is constructed; the lightweight contour data can perfectly adapt to LoRa narrowband transmission, and reduce the occupancy rate of network bandwidth; the method can be long-term used for infrastructure state monitoring such as road crack identification through the separated background image, and the structured contour data directly drives the cross-camera target association to realize the reconstruction of the global object motion trajectory; at the same time, the central server is provided with pre-screened high-value features, so that the event preliminary screening response time is effectively compressed to the target range, and the event capture rate can be guaranteed in complex scenes such as traffic hubs and dense human flow areas, and finally the system goal of low bandwidth, high precision and fast response is achieved.
[0070] In a specific embodiment, the process of identifying the same object can specifically include the following steps:
[0071] extracting object features based on the object contours, the object features including geometry, color, texture, edge and corner point of the object; comparing the i+1th object contour in the i+1th frame image with the ith object contour in the ith frame image, and determining that the i+1th object contour and the ith object contour belong to the same object when the similarity between the i+1th object contour and the ith object contour is greater than or equal to a second threshold, otherwise determining that the i+1th object contour and the ith object contour belong to different objects.
[0072] Specifically, by extracting the object contour of each frame image from the video data, the features such as geometry, color, texture, edge and corner point of the object are extracted; the features can comprehensively describe the appearance and behavior of the object; then the features are used to compare the objects, specifically, by comparing the object contour in the i+1th frame image with the object contour in the ith frame image, the similarity is calculated; when the similarity is greater than or equal to the set second threshold, it can be determined that the two object contours are the same object, otherwise it is determined to be different objects; the method can effectively ensure the consistency of the object in the continuous frames, accurately track the motion trajectory of the same object, and realize accurate identification of the same object, and at the same time, it can maintain high identification accuracy in complex environment and dynamic change.
[0073] In a specific embodiment, the process of identifying the same object can further include the following steps:
[0074] obtaining a first position of a first terminal where the video data is from, obtaining other video acquisition terminals closest to the first position, and obtaining video data of the other video acquisition terminals; when the similarity between the object contour in the video data of the other video acquisition terminals and the object contour in the video data is greater than or equal to a first threshold, merging the video data of the other video acquisition terminals and the video data into one video data, and regenerating the video summary data of the video data.
[0075] Specifically, by acquiring the position information of the first terminal of the video data source, the position information of the other video acquisition terminals closest to the position is further acquired. The terminals can respectively acquire video data of different perspectives. Next, by comparing the object contours in the video data of each terminal, if the similarity of the object contours extracted from the other video acquisition terminals and the object contours in the first terminal is greater than or equal to a preset first threshold, it is considered that the objects belong to the same object. At this time, the video data from different terminals is merged to form a unified video data, and the summary data of the video data is regenerated; the summary data includes key information such as the position information, size, and motion trajectory of the object, thereby improving the object recognition capability across terminals and perspectives. The method can effectively realize multi-perspective and all-around recognition of the same object by combining the similarity analysis of the object contours and the geographical positions of the video acquisition terminals. By merging the data from multiple terminals, the accuracy of the recognition is enhanced, the monitoring range is expanded, and the utilization efficiency of the video data is improved. In a complex monitoring environment, especially in the case of multiple perspectives and different geographical positions, the continuity and accuracy of object tracking can be ensured.
[0076] In a specific embodiment, the process of determining whether the preset event exists in the summary data can specifically include the following steps:
[0077] Based on the summary data and the pre-trained event prediction model, the event corresponding to the summary data is predicted, and an event label is added to the summary data, wherein the event refers to a specific behavior model, at least including an abnormal person entering a restricted area, a vehicle reversing, a rapid gathering, a fire, an explosion, and a fight.
[0078] Specifically, first, the video, audio, and sensor data extracted from the summary data are analyzed, and the event prediction model is used for processing. The event prediction model is pre-trained and can identify different types of behavior patterns and predict whether a preset event occurs according to the characteristics of the summary data. The preset event includes behavior patterns such as an abnormal person entering a restricted area, a vehicle reversing, a rapid gathering, a fire, an explosion, and a fight. For each piece of summary data, the model will evaluate the matching degree with a specific behavior model, and add a corresponding event label to the data to mark the occurrence of the event; if the prediction model identifies that there is a behavior in the summary data that matches the preset event, it is determined that the event is a real event, and an alarm mechanism can be further triggered. The method realizes automatic event detection by deeply analyzing the key information in the summary data and combining the pre-trained behavior model. This process can not only quickly and accurately extract potential security events from a large amount of monitoring data, but also can respond in the early stage of the event, improving the timeliness of event warning.
[0079] In a specific embodiment, the process of generating summary data can specifically include the following steps:
[0080] Audio data is obtained from the monitoring data, the audio data is parsed, audio summary data is generated, the audio summary data includes terminal ID, time of sound occurrence and predicted content description of the sound; sensor data is obtained from the monitoring data, the sensor data is preprocessed, and text summary data is generated, the text summary data includes terminal ID, recording time and content description.
[0081] Specifically, audio data is obtained from the monitoring data and parsed, key information in the audio data such as the time of sound occurrence and the type of sound will be extracted; based on these information, audio summary data is generated, the content includes terminal ID, time of sound occurrence, and predicted content description of the sound; through the analysis of the sound, the possible occurrence of the event can be identified, such as abnormal sound, shouting, vehicle brake sound, etc., thereby providing important clues for subsequent event prediction and analysis; secondly, the sensor data in the monitoring data will also be obtained and preprocessed, the sensor data may contain temperature, humidity, pressure and other information, after analysis and processing of the data, text summary data is generated, the content includes terminal ID, recording time and corresponding content description, to reflect the changes of the environment, physical state or potential safety risks. By parsing the audio and sensor data, a simplified representation of the environmental changes and events is obtained from multiple dimensions, thereby reducing the complexity of the data and improving the data processing efficiency, and at the same time, by generating the summary data of the audio and the sensor, the system can reduce the data volume while maintaining the data accuracy, and improve the speed of subsequent processing and event identification.
[0082] In a specific embodiment, the process of verifying the event can specifically include the following steps:
[0083] The first monitoring data and the second monitoring data of the other two terminals closest to the location of the terminal of the target monitoring data source are obtained, if the target monitoring data is video data, the first monitoring data represents audio data, and the second monitoring data represents text data, if the target monitoring data is audio data, the first monitoring data represents video data, and the second monitoring data represents text data, if the target monitoring data is text data, the first monitoring data represents video data, and the second monitoring data represents audio data;
[0084] The contribution weights of the video data, the audio data and the text data to the event occurrence probability are w1, w2 and w3 respectively, and the probability P of the event occurrence is calculated by the formula, when the probability P of the event occurrence is greater than a preset probability, it is determined that the event occurs, otherwise the event is ignored, wherein T1 represents video data, T2 represents audio data, T3 represents text data, w1 is greater than w2, and w2 is greater than w3.
[0085] Specifically, by introducing multi-modal data: video, audio, text and their respective weights, the contribution of different types of data to the occurrence of the event is considered comprehensively, so as to verify the event from multiple angles and all aspects; through the weight weighting model, the authenticity of the event can be judged more accurately, the possibility of false positives or false negatives can be reduced, and the accuracy and reliability of event verification can be improved, especially in complex monitoring environments, to ensure that the system can make more accurate responses, reduce invalid alarms and optimize the use of resources.
[0086] In a specific embodiment, the process of generating summary data can specifically include the following steps:
[0087] When generating summary data, if there is sensitive information in the summary data, the sensitive information is covered by encoding or text, including face images, identity information, license plate numbers, etc.
[0088] When transmitting target monitoring data, the target monitoring data can also be processed to cover sensitive information in the target monitoring data, and if necessary, the original data can be restored through the relationship between the encoding or text of the covering data and the sensitive information.
[0089] Specifically, by automatically identifying and covering sensitive information during generation and transmission, and using encoding and text substitution, data privacy protection is achieved; this method not only effectively prevents the leakage of sensitive information, but also ensures compliance with privacy protection requirements while maintaining data integrity and usability, improving data security, which is particularly important for smart city systems that involve a large amount of monitoring data, and ensures that data is always in a safe and controlled state during processing and transmission.
[0090] The above describes the smart city multi-scene monitoring method based on multi-protocol communication in the embodiments of the present application, and the following describes the smart city multi-scene monitoring system based on multi-protocol communication in the embodiments of the present application. Please refer to Figure 4 An embodiment of the smart city multi-scene monitoring system based on multi-protocol communication in the embodiments of the present application includes:
[0091] The data acquisition unit is configured to deploy heterogeneous terminals in the whole city and collect monitoring data of the terminals in real time and send the monitoring data to the edge server;
[0092] The edge data processing unit is configured to process the monitoring data from multiple terminals to generate summary data, and transmit the summary data to the central server through the LoRa communication unit, wherein the summary data includes video summary data, audio summary data and text summary data, and each of the summary data comes from a different terminal.
[0093] The service data processing unit is configured to determine whether a preset event exists in the summary data after the central server receives the summary data from the plurality of edge servers, and request the edge server to send target monitoring data corresponding to a time period of the summary data if the preset event exists.
[0094] The data transmission unit is configured to transmit the target monitoring data to the central server by the 4G communication unit or the 5G communication unit.
[0095] The event verification unit is configured to verify the event in combination with the target monitoring data, generate alarm information of the event and feed back to a related management system after determining the event.
[0096] Through the cooperation of the above-mentioned components, the application can further reduce the network load and storage pressure while improving the event analysis efficiency.
[0097] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-mentioned system, system and unit can refer to the corresponding process in the foregoing method embodiments, which will not be described here.
[0098] The integrated unit, if realized in the form of a software function unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the present application essentially or say the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0099] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A smart city multi-scene monitoring method based on multi-protocol communication, characterized in that, The method comprises: Deploying heterogeneous terminals in the whole city and collecting monitoring data of the terminals in real time and sending the monitoring data to an edge server; The edge server processes the monitoring data from multiple terminals to generate summary data, and transmits the summary data to a central server through a LoRa communication unit, wherein the summary data includes video summary data, audio summary data and text summary data, and is from different terminals respectively; After the central server receives the summary data from multiple edge servers, it determines whether there is a preset event in the summary data, and if so, requests the edge server to send target monitoring data of a time period corresponding to the summary data; The edge server transmits the target monitoring data to the central server through a 4G communication unit or a 5G communication unit; The central server verifies the event in combination with the target monitoring data, and after determining the event, generates alarm information of the event and feeds back to the related management system; Verifying the event includes: obtaining first monitoring data and second monitoring data of other two terminals closest to the location of the terminal where the target monitoring data comes from, if the target monitoring data is video data, the first monitoring data represents audio data, and the second monitoring data represents text data, if the target monitoring data is audio data, the first monitoring data represents video data, and the second monitoring data represents text data, if the target monitoring data is text data, the first monitoring data represents video data, and the second monitoring data represents audio data; The contribution weights of the video data, the audio data and the text data to the event occurrence probability are w1, w2 and w3 respectively, and the event occurrence probability P is calculated by the formula The probability P of the event occurrence is calculated, and when the probability P of the event occurrence is greater than a preset probability, it is determined that the event occurs, otherwise the event is ignored, wherein T1 represents video data, T2 represents audio data, T3 represents text data, w1 is greater than w2, w2 is greater than w3, the central server combines target monitoring data and auxiliary data of adjacent terminals, and calculates the event occurrence probability through a weighted probability model, video weight w1> audio w2> text w3; if the probability exceeds a threshold value, an alarm information is generated and pushed to a management system.
2. The method of claim 1, wherein, Generating summary data includes: Obtaining video data from the monitoring data, and extracting multiple object contours and background images from each frame of image of the video data; Comparing multiple object contours in the video data to identify the same object, generating the same object identified in the same image as an object image, and generating a unique identifier for the object image; Analyzing the characteristics and behaviors of the object image to generate video summary data of the video data, wherein the video summary data includes terminal ID, time, position, size, speed, direction and motion trajectory of the object displayed in the video data.
3. The method of claim 2, wherein, Extracting multiple object contours and background images includes: Obtaining the i-th frame image from the video data, and processing the i-th frame image based on an AI algorithm to obtain the i-th background image; When the similarity between the i-th background image and the background image of the i+1-th frame image after the i-th frame image is greater than or equal to a first threshold, comparing the pixels of the i-th frame image with the pixels of the i+1-th frame image to detect the changed object contour; When the background image similarity between the i-th background image and the background image of the i+1-th frame image after the i-th frame image is less than a first threshold, an i+1-th background image is acquired from the i+1-th frame image, when the background image similarity between the i+1-th background image and the background image of the i+2-th frame image after the i+1-th frame image is greater than or equal to the first threshold, pixels of the i+1-th frame image are compared with pixels of the i+2-th frame image, and a changed object contour is detected, i represents a positive integer from 1 to A, and A represents a total frame number of the video data; Similarly, until all changed object contours and background images are detected.
4. The method of claim 2, wherein, Identifying the same object comprises: Based on the object contour, an object feature is extracted, an i+1-th object contour in the i+1-th frame image is compared with an i-th object contour in the i-th frame image, when the similarity between the i+1-th object contour and the i-th object contour is greater than or equal to a second threshold, it is determined that the i+1-th object contour and the i-th object contour belong to the same object, otherwise, it is determined that the i+1-th object contour and the i-th object contour belong to different objects, i represents a positive integer from 1 to A, and A represents a total frame number of the video data.
5. The method of claim 3, wherein, Identifying the same object further comprises: A first position where a first terminal of the video data source is located is acquired, other video acquisition terminals closest to the first position are acquired, and video data of the other video acquisition terminals is acquired; when the similarity between an object contour in the video data of the other video acquisition terminals and an object contour in the video data is greater than or equal to the first threshold, the video data of the other video acquisition terminals and the video data are merged into one video data, and video summary data of the video data is regenerated.
6. The method of claim 1, wherein, Determining whether a preset event exists in the summary data comprises: Based on the summary data and a pre-trained event prediction model, an event corresponding to the summary data is predicted, and an event label is added to the summary data.
7. The method of claim 1, wherein, Generating the summary data further comprises: Audio data is acquired from the monitoring data, the audio data is analyzed, and audio summary data is generated, the audio summary data comprising a terminal ID, a sound occurrence time and a content description of a sound prediction; Sensor data is acquired from the monitoring data, the sensor data is preprocessed, and text summary data is generated, the text summary data comprising a terminal ID, a recording time and a content description.
8. The method of claim 1, wherein, Generating the summary data further comprises: When the summary data contains sensitive information, the sensitive information is covered by encoding or text when the summary data is generated. 9.A smart city multi-scene monitoring system based on multi-protocol communication, used to implement the smart city multi-scene monitoring method based on multi-protocol communication according to any one of claims 1-8. The system comprises: A data acquisition unit is configured to deploy heterogeneous terminals in the whole city and collect monitoring data of the terminals in real time and send the monitoring data to an edge server; An edge data processing unit is configured to process the monitoring data from multiple terminals to generate summary data by the edge server, and transmit the summary data to a central server through a LoRa communication unit, the summary data comprising video summary data, audio summary data and text summary data, and being from different terminals respectively; A service data processing unit is configured to determine whether a preset event exists in the summary data after the central server receives the summary data from the plurality of edge servers, and request the edge server to send target monitoring data corresponding to a time period of the summary data if the preset event exists; A data transmission unit is configured to transmit the target monitoring data to the central server by a 4G communication unit or a 5G communication unit of the edge server; An event verification unit is configured to verify the event in combination with the target monitoring data by the central server, generate alarm information of the event and feed back to a related management system after determining the event.
Citation Information
Patent Citations
Data preprocessing method and system for edge gateway of Internet of Things
CN114518960A
Intelligent monitoring system for key area based on communication tower
CN119603650A
Smart city management system
CN112153464A
Smart city video monitoring system and method based on 5G edge computing
CN112437259A