A monitoring system for urban public safety with intelligent question answering and decision support.
The monitoring system, which integrates intelligent question answering and decision support, enables closed-loop management of the entire urban monitoring system. This solves the problems of data isolation and insufficient emergency response efficiency in existing systems, improves emergency response speed and public interaction capabilities, and provides accurate information and decision support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XUZHOU HIGH TECH ZONE SAFETY EMERGENCY EQUIPMENT INDUSTRIAL TECHNOLOGY RESEARCH INSTITUTE
- Filing Date
- 2026-01-28
- Publication Date
- 2026-06-02
AI Technical Summary
Existing urban monitoring systems are limited in function, data is isolated, and they lack cross-scenario correlation analysis and comprehensive decision-making capabilities. This limits their emergency response effectiveness, public interaction capabilities, and command centers lack intelligent decision support.
The system employs a monitoring system with intelligent question-and-answer capabilities and decision support, including front-end sensing and data acquisition, intelligent video analysis, AI big data model analysis and decision-making, public safety dashboard interaction, and proactive intervention units. This enables closed-loop management throughout the entire process. Through AI big data models, it conducts multimodal in-depth analysis and risk assessment, providing precise question-and-answer services and targeted voice alarms.
It has achieved closed-loop management of the entire process from safety risk perception to public interaction, improved emergency response speed and processing efficiency, provided scenario-based and actionable information and decision support, surpassed the limitations of traditional systems, and achieved precise intervention and proactive service.
Smart Images

Figure CN122134530A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of smart city and public safety monitoring technology, specifically relating to a monitoring system for urban public safety with intelligent question answering and decision support. Background Technology
[0002] With the acceleration of urbanization, public safety management faces increasing complexity and challenges. Existing urban monitoring systems rely heavily on manual inspection of video footage, which has limitations such as blind spots, personnel fatigue, and response delays. With the development of technology, video analysis technology based on computer vision has been gradually applied, which can realize the automatic detection of specific events such as flames, smoke, and crowd gatherings. However, such systems are often single-function, and the data between subsystems is isolated, making it difficult to support cross-scenario correlation analysis and comprehensive decision-making.
[0003] Currently, the commonly used methods for disseminating public safety information mostly rely on one-way broadcasting or static signage, lacking the ability to interact with the public in real time and intelligently. When emergencies occur, the public finds it difficult to obtain emergency guidance that is specific to the situation and highly operable. At the same time, command centers also lack intelligent decision support tools that can integrate real-time situational awareness and historical experience, thus restricting the overall effectiveness of emergency response.
[0004] Therefore, in response to the aforementioned technical problems, it is necessary to provide a monitoring system for urban public safety that features intelligent question answering and decision support.
[0005] The information disclosed in this background section is intended only to enhance the understanding of the overall background of the invention and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this invention is to provide a monitoring system for urban public safety with intelligent question-and-answer capabilities and decision support. This system enables closed-loop management of the entire process, from automatic risk perception, intelligent analysis and assessment, tiered proactive intervention to intelligent interactive services for the public.
[0007] To achieve the above objectives, a specific embodiment of the present invention provides the following technical solution: A monitoring system for urban public safety with intelligent question answering and decision support includes: The front-end sensing and acquisition unit is used to collect real-time video data of the monitored area; The video intelligent analysis unit is used to identify and analyze the real-time video data to obtain structured security event information; The AI large model analysis and decision-making unit is used to receive and fuse the structured security event information and associated context data, perform multimodal deep analysis and risk assessment, and generate alarm information and response strategies. The public safety dashboard interactive unit includes a display interface and an intelligent question-and-answer window integrated thereon. The display interface is used to display monitoring screens, analysis results and alarm information. The intelligent question-and-answer window is used to receive natural language queries and, based on the output of the AI big model analysis and decision-making unit, to provide question-and-answer services and auxiliary decision-making solutions related to the current security situation. An active intervention unit is used to execute voice alarms in the monitored area according to the instructions of the AI big model analysis and decision-making unit. The system management platform is used to coordinate and control the data flow and instruction flow between the various units.
[0008] In one or more embodiments of the present invention, the video intelligent analysis unit includes multiple dedicated target detection models for identifying at least one of fire safety incidents and road traffic safety incidents.
[0009] In one or more embodiments of the present invention, when the AI large model analysis and decision-making unit performs multimodal deep analysis, the fused contextual data includes at least one of time information, weather information, historical event data, and geographic information system data.
[0010] In one or more embodiments of the present invention, when the AI large model analysis and decision-making unit generates a response strategy, the output content includes at least one of the following: risk level, suggested handling steps, and emergency resource information to be mobilized.
[0011] In one or more embodiments of the present invention, the question-and-answer service and decision-making assistance provided by the intelligent question-and-answer window are dynamically linked to the current security situation obtained by the AI big model analysis and decision-making unit from the analysis of real-time video data.
[0012] In one or more embodiments of the present invention, the active intervention unit includes a directional speaker column bound to a specific camera or monitoring point, used to play customized voice alarm content for specific event types and risk levels according to the instructions of the AI big model analysis and decision-making unit.
[0013] In one or more embodiments of the present invention, the system management platform is configured with a rule engine for setting the mapping relationship between different risk levels, event types and the display strategy of the public safety dashboard interaction unit and the response strategy of the proactive intervention unit.
[0014] In one or more embodiments of the present invention, the intelligent question-and-answer window of the public safety dashboard interactive unit has a background engine connected to a public safety knowledge graph to enhance the accuracy of routine safety questions and answers.
[0015] In one or more embodiments of the present invention, the following are included: Real-time video streams are acquired through a front-end sensing and acquisition unit; The video stream is analyzed in real time by the video intelligent analysis unit to identify potential security incidents; The identified events and related data are input into the AI big model analysis and decision-making unit for comprehensive evaluation and generation of decision instructions; Based on the decision instruction, at least one of the following operations shall be executed synchronously: a. Update and display alarm information on the public safety dashboard interactive unit; b. Provide targeted voice alarms through active intervention units; c. Update the context of the smart question-and-answer window to support interactive question-and-answer and decision assistance based on the current event.
[0016] In one or more embodiments of the present invention, the program, when executed by a processor, implements the steps of the method as described in claim 9.
[0017] Compared with existing technologies, the urban public safety monitoring system of the present invention, which features intelligent question answering and decision support, realizes a complete closed loop of "perception-analysis-decision-intervention-interaction," significantly improving emergency response speed and processing efficiency. Through multimodal fusion analysis of AI big data models, it can understand complex scenarios, perform correlation reasoning and risk assessment, and surpass the limitations of traditional single algorithms that can only identify specific targets. The intelligent question answering window is deeply integrated with the real-time monitoring situation, providing the public and managers with scenario-based, actionable, and accurate information and decision support, transforming passive queries into proactive intelligent services. Through a tiered intervention strategy, it achieves precise intervention from information prompts to targeted voice alarms, avoiding disturbance to residents and improving intervention effectiveness. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of the overall architecture of the system of the present invention; Figure 2 This is a flowchart of the AI large model analysis and decision-making unit in the system of this invention; Figure 3 This is a sequence diagram of how the system of the present invention processes a specific security event. Detailed Implementation
[0020] To enable those skilled in the art to better understand the technical solutions in this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this disclosure.
[0021] like Figure 1 As shown, the urban public safety monitoring system with intelligent question answering and decision support in one embodiment of the present invention adopts a cloud-edge-device collaborative architecture. The edge consists of network high-definition cameras widely deployed in urban roads, squares, key buildings, and other areas. These cameras continuously collect real-time video streams via wired or wireless networks. Environmental sensors can be integrated at some key locations, and the data is uploaded along with them.
[0022] The edge / cloud side includes a video intelligent analysis unit and an AI large model analysis and decision-making unit; The video intelligent analysis unit can be deployed on an edge computing gateway or cloud server. This unit has a variety of trained dedicated deep learning models built in. AI large-scale model analysis and decision-making units are typically deployed in cloud computing centers and have powerful multimodal processing capabilities.
[0023] The application layer includes a public safety dashboard interaction unit and a proactive intervention unit; Public safety dashboard interactive units are deployed on command center screens, community service center displays, or information kiosks in public places; Active intervention units are paired with outdoor speaker columns or network cameras with voice functionality, which are compatible with specific cameras.
[0024] The management and coordination layer is a centralized software platform used for system configuration, monitoring, operation and maintenance, and log management.
[0025] Implementation methods for each unit of the system: The video intelligent analysis unit receives video streams from the front-end camera. To achieve the function in claim 2, this embodiment deploys multiple lightweight dedicated target detection models, which can be called in parallel or on demand. These models include a fire safety detection model, a road traffic safety detection model, and a general safety detection model. The fire safety detection model is trained based on YOLOv8 or a similar architecture and can detect targets such as flames, smoke, blocked fire exits, and electric vehicles entering elevators in video frames in real time. The road traffic safety detection model is also based on deep learning and can identify events such as vehicles driving in the wrong direction, pedestrians running red lights, non-motorized vehicles occupying motor vehicle lanes, and traffic accidents. A general safety detection model is used to identify abnormal gatherings of people, falls, and items left behind. This unit performs real-time analysis of the video stream and immediately generates a structured alarm message once a preset security event is identified. This message includes at least: event type, confidence level, time of occurrence, associated camera ID, and bounding box coordinates in the frame. This information is encapsulated in JSON format and sent to the AI big data model analysis and decision-making unit via a message queue.
[0026] AI large-scale model analysis and decision-making units are the intelligent brain of the system, such as Figure 2 As shown, its workflow is as follows: Data reception and fusion: The unit subscribes to the message queue to receive structured alarm information from the video analysis unit. At the same time, it obtains rich contextual data from external systems through the API interface, including real-time weather data, historical alarm records of the location, real-time traffic flow data, and geographic information system data of the camera location, such as the street, surrounding building types, and the location of fire-fighting facilities.
[0027] Multimodal deep analysis and risk assessment involves inputting the aforementioned multi-source data into a pre-trained and fine-tuned multimodal large-scale model. This model, based on the Transformer architecture, integrates visual understanding and text reasoning capabilities. The large-scale model performs the following tasks: For example, if smoke is detected, combined with GIS tags of the kitchen area and midday time information, it can be preliminarily judged as possibly being restaurant fumes or a fire; if people are detected gathering, combined with large-scale event registration data and real-time crowd heat map, it can be judged whether the gathering is within the permitted range and whether there is a risk of stampede; based on the nature of the event, the surrounding environment, and historical data, a quantitative risk level is output, which can be divided into low risk, medium risk, high risk, emergency, etc.
[0028] Strategy generation is based on the analysis results. The large model generates specific response strategies. For example, for a high-risk initial fire event, the strategy may include alarm content, handling suggestions, and intervention instructions. When a suspected fire is detected near the designated location, nearby personnel should take precautions and staff should proceed to investigate; notify the nearest patrol officer A to verify; prepare to call upon resources from the mini fire station within 200 meters; continuously monitor the fire's development via cameras; trigger a voice alarm and specify the content to be played and the target camera / speaker.
[0029] The interactive interface of the public safety signboard is divided into several areas: The main monitoring area displays real-time video of key areas in a rotating or split-screen manner; the event list area displays various security events identified by the system in a scrolling manner, highlighted with different colors according to risk level; the electronic map area marks the location of the events on the GIS map; The intelligent question-answering window is an interactive dialog box embedded on one side of the interface. The backend service of the intelligent question-answering window maintains a long WebSocket connection with the AI big model analysis and decision-making unit to obtain the overall security status of the current system and the context of all active events in real time. When the user enters a natural language question, the question is sent to the big model. The big model combines the user's question with the current real-time data. For general knowledge, it calls the integrated public security knowledge graph to answer; for specific events, it directly generates a response based on the analysis results of that event. For emergency response questions, the large model combines emergency plans from the knowledge graph with historical successful handling cases to generate step-by-step auxiliary decision-making solutions. For example, if a user selects a traffic accident icon on the map and asks how to handle it, the window might reply: "A rear-end collision of two vehicles has been detected, but no one is on the ground." It might also suggest: using a camera to alert drivers to turn on their hazard lights and place warning signs behind them; automatically dispatching the incident to the nearest traffic police patrol post; and synchronizing information with the traffic signal control platform, suggesting an appropriate extension of the green light for that direction.
[0030] The proactive intervention unit consists of network speaker columns and speakers integrated into IP cameras. Each speaker column / speaker is geographically bound to one or more cameras within the system management platform. When it receives an intervention command from the AI big data model analysis and decision-making unit, the unit converts the text alert content into speech using a TTS engine, or directly plays pre-recorded audio clips. The broadcast is then directed and looped through the sound devices at designated locations to avoid noise interference in unrelated areas. The broadcast content is customized based on the event type and risk level.
[0031] The system management platform is a comprehensive web-based management backend. Its core is a rules engine, allowing administrators to configure various rules through a graphical interface, including display rules, intervention rules, and linkage rules. When a high-risk event occurs, a pop-up window automatically displays the camera's image on the command center's large screen, accompanied by a flashing alarm. When "Fire Safety - Open Flame" is detected and the risk level is emergency, a Level 1 voice alarm is automatically triggered, and an alarm SMS is sent to the pre-set fire safety supervisor's mobile phone. When a major traffic accident is detected, the system can automatically link with the traffic signal system to implement temporary traffic control at relevant intersections. The platform is also responsible for user permission management, device status monitoring, and the storage and auditing of all alarms and Q&A records.
[0032] like Figure 3 As shown, taking a street trash can fire incident as an example, the closed-loop workflow of the system is illustrated: The front-end camera captures a video stream containing flames. The video intelligent analysis unit identifies the flame target, generates a structured alarm, and reports it. The AI big model unit receives the alarm, combines weather, geographical location, and historical data, assesses the risk as high, generates a strategy, issues an alarm on the dashboard, makes targeted voice announcements, and prompts the location of nearby fire extinguishing equipment.
[0033] Afterwards, the system management platform pushes the event to the public safety dashboard according to the rules, the main screen switches to the monitoring screen, and sends a command to the speaker attached to the monitoring screen to broadcast: "Attention citizens, there is a fire in the trash can near you. Please do not approach. Staff are handling the situation. The nearest fire extinguisher is located in the red box 20 meters to the southeast." At this point, personnel in the command center can enter the question "Will the fire spread to the nearby shops?" through the intelligent question-and-answer window. The large model, based on real-time camera footage and knowledge of the shop's materials, will provide analytical conclusions and suggestions for strengthening monitoring.
[0034] After on-site personnel have completed the handling, they can mark the event as "handled" on the platform. This handling result can be used as empirical data to optimize future decisions of the large model.
[0035] The above embodiments fully disclose the technical solutions claimed in claims 1-8. The method described in claim 9 is fully embodied in the workflow of the above system. When the program stored on the storage medium described in claim 10 is executed by the processor, it realizes the functions and method flow of each unit of the above system. Those skilled in the art can fully implement the present invention based on the above detailed description.
[0036] It will be apparent to those skilled in the art that this disclosure is not limited to the details of the exemplary embodiments described above, and that this disclosure can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of this disclosure is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within this disclosure. No reference numerals in the claims should be construed as limiting the scope of the claims.
[0037] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A monitoring system for urban public safety with intelligent question-and-answer and decision-making support, characterized in that, include: The front-end sensing and acquisition unit is used to collect real-time video data of the monitored area; The video intelligent analysis unit is used to identify and analyze the real-time video data to obtain structured security event information; The AI large model analysis and decision-making unit is used to receive and fuse the structured security event information and associated context data, perform multimodal deep analysis and risk assessment, and generate alarm information and response strategies. The public safety dashboard interactive unit includes a display interface and an intelligent question-and-answer window integrated thereon. The display interface is used to display monitoring screens, analysis results and alarm information. The intelligent question-and-answer window is used to receive natural language queries and, based on the output of the AI big model analysis and decision-making unit, to provide question-and-answer services and auxiliary decision-making solutions related to the current security situation. An active intervention unit is used to execute voice alarms in the monitored area according to the instructions of the AI big model analysis and decision-making unit. The system management platform is used to coordinate and control the data flow and instruction flow between the various units.
2. The system according to claim 1, characterized in that, The video intelligent analysis unit includes multiple dedicated target detection models for identifying at least one of fire safety incidents and road traffic safety incidents.
3. The system according to claim 1, characterized in that, When the AI large model analysis and decision-making unit performs multimodal deep analysis, the contextual data it integrates includes at least one of time information, weather information, historical event data, and geographic information system data.
4. The system according to claim 1, characterized in that, When the AI large-scale model analysis and decision-making unit generates response strategies, the output includes at least one of the following: risk level, suggested handling steps, and information on emergency resources that need to be mobilized.
5. The system according to claim 1, characterized in that, The question-and-answer service and decision-making support provided by the intelligent question-and-answer window are dynamically linked to the current security situation obtained by the AI big model analysis and decision-making unit from the analysis of real-time video data.
6. The system according to claim 1, characterized in that, The active intervention unit includes directional sound columns that are bound to specific cameras or monitoring points, and are used to play customized voice alarm content based on the instructions of the AI big data model analysis and decision-making unit for specific event types and risk levels.
7. The system according to claim 1, characterized in that, The system management platform is equipped with a rule engine, which is used to set the mapping relationship between different risk levels, event types and the display strategy of the public safety dashboard interaction unit and the response strategy of the active intervention unit.
8. The system according to claim 1, characterized in that, The intelligent question-and-answer window of the public safety dashboard interactive unit has a backend engine connected to a public safety knowledge graph to enhance the accuracy of routine safety questions and answers.
9. A method for urban public safety monitoring based on the system described in any one of claims 1-8, characterized in that, include: Real-time video streams are acquired through a front-end sensing and acquisition unit; The video stream is analyzed in real time by the video intelligent analysis unit to identify potential security incidents; The identified events and related data are input into the AI big model analysis and decision-making unit for comprehensive evaluation and generation of decision instructions; Based on the decision instruction, at least one of the following operations shall be executed synchronously: a. Update and display alarm information on the public safety dashboard interactive unit; b. Provide targeted voice alarms through active intervention units; c. Update the context of the smart question-and-answer window to support interactive question-and-answer and decision assistance based on the current event.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method as described in claim 9.