Safety management camera module and method based on visual model

By adopting visual model-based technology in the security management camera, combining high-definition optical lenses, image sensors and powerful visual big models, the problem of insufficient recognition and analysis capabilities of traditional cameras in complex scenes is solved, high-precision recognition and in-depth data analysis are achieved, and the effectiveness and intelligence of security monitoring are improved.

CN120075566APending Publication Date: 2025-05-30CHENGDU BESTNAC TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510148783.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When facing complex surveillance scenarios, traditional security management cameras are difficult to accurately identify and analyze images, and are prone to false alarms and missed alarms, and lack in-depth data mining capabilities, so they cannot extract valuable security information in a timely manner.

Method used

The security management camera module based on vision models is adopted, including high-definition optical lens, image sensor, Doubao vision big model embedding module, data processing and analysis module, communication module, power management module, storage module and alarm module. Through the coordinated work of these modules, high-precision identification and in-depth data analysis of monitoring scenarios are achieved.

Benefits of technology

It greatly reduces misjudgment, improves the effectiveness of security monitoring, accurately captures key security-related information, discovers potential safety hazards in advance, and generates detailed security risk assessment reports, which improves the intelligence level and actual effect of security management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075566A_ABST
    Figure CN120075566A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of safety monitoring, and particularly relates to a safety management camera module and method based on a visual model.The safety management camera module comprises an image collecting module, a bean bag visual large model embedding module, a Qwen2-VL, a data processing and analyzing module and a communication module.The image collecting module continuously collects images; after deep analysis and processing of the bean bun visual large model embedding module and the data processing and analyzing module, information is transmitted outwards through the communication module, data is stored through the storage module, power supply is guaranteed through the power management module, and an alarm is given out in time when necessary, so that a set of complete intelligent safety management monitoring system is formed. Through combination with an advanced visual large model, more accurate, intelligent and efficient safety monitoring and management functions are realized, the false alarm rate and the missing report rate are effectively reduced, safety management personnel are assisted to better fulfill responsibilities, and the safety of a monitored area is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of security monitoring, and specifically relates to a security management camera module and method based on a vision model. Background Art

[0002] With the development of society and the continuous improvement of people's safety awareness, security management cameras have been widely used in many places such as residential communities, commercial office areas, industrial parks, etc. Traditional security management cameras can often only achieve simple video image acquisition functions, and their ability to analyze and judge the collected images is relatively limited. They usually rely on preset fixed rules to identify basic situations such as moving targets and objects with specific shapes, and it is difficult to accurately respond to complex and changeable actual monitoring scenarios, prone to false alarms, missed alarms and other situations.

[0003] In the face of a large amount of monitoring video data, traditional cameras lack effective data integration and in-depth mining capabilities, and are unable to extract valuable security-related information in a timely manner to assist managers in making efficient decisions, which greatly limits the intelligent level and actual effect of security management. Therefore, there is an urgent need for a new camera technology with stronger image analysis capabilities, capable of deeply mining data value and accurately making security management judgments.

[0004] For this reason, the present invention provides a security management camera module and method based on a vision model. Summary of the Invention

[0005] In order to make up for the deficiencies of the prior art and solve at least one of the technical problems proposed in the background art.

[0006] The technical solution adopted by the present invention to solve its technical problems is: A security management camera module based on a vision model according to the present invention is characterized in that: the security management camera module includes:

[0007] An image acquisition module, composed of a high-definition optical lens and an image sensor component, responsible for collecting real-time video image information in the monitoring area;

[0008] A Doubao Vision Large Model Embedding Module, after receiving the video image data from the image acquisition module, the vision large model can accurately identify and classify the objects, people, and behavioral actions in the image;

[0009] Qwen2-VL, which can identify complex scene object relationships, handwritten and multi-language image texts, has excellent visual reasoning ability to solve problems through chart analysis, can understand long videos and support real-time conversations for multiple applications, supports multiple languages for global users, and with an innovative architecture can process images of any resolution, integrate multi-dimensional information, and also has the ability to expand application such as answering questions by taking pictures;

[0010] The data processing and analysis module, in collaboration with the Doubao visual large model embedding module, not only further organizes, judges and integrates the output results of the module, but also compares the real-time and historical monitoring data through algorithms, mines security trends and anomalies, generates evaluation reports, and assists in security management decision-making;

[0011] The communication module supports multiple communication protocols, such as Wi-Fi, Ethernet, and 4G / 5G. It can transmit the video image data collected by the camera, the analysis results of the Doubao visual model, and the security risk assessment report generated by the data processing and analysis module to the remote monitoring management platform or the mobile terminal of the security personnel in real time, ensuring the timely transmission of security information and facilitating the management personnel to remotely control the status of the monitoring area and respond quickly;

[0012] The power management module provides a stable and reliable power supply for the entire camera system. It can be connected to the mains power supply and has a built-in backup power supply such as a lithium battery. It has power monitoring, intelligent switching and energy-saving control functions.

[0013] The storage module is equipped with a large-capacity storage medium, such as a solid-state hard disk, which is mainly used to store the original video image data collected by the camera, as well as the key information and analysis results obtained after processing in various links;

[0014] The alarm module works with the data processing and analysis module to monitor the safety status of the surveillance area; triggers multiple alarm modes, deters potentially dangerous persons by emitting sound and light signals, and pushes messages and voice prompts to the manager's mobile terminal.

[0015] Preferably, the image acquisition module has functions such as adjustable focal length and aperture, and can adapt to the needs of acquiring clear images under different distances and lighting conditions. The acquired image data is transmitted in real time in the form of digital signals to subsequent modules for processing.

[0016] Preferably, the Doubao visual big model embedding module has a pre-trained Doubao visual big model built in, which is deeply trained based on a large amount of different scenes, various objects, and human behavior image data, and has powerful image feature extraction, semantic understanding and pattern recognition capabilities; after receiving the video image data from the image acquisition module, the visual big model can accurately identify and classify objects, people, and behavioral actions in the image, for example, accurately distinguish whether they are normal pedestrians, staff or suspicious persons, identify specific dangerous items, abnormal scene arrangements, such as blocked fire escape routes, etc., and at the same time, it can also analyze and judge the behavior trajectory and movement posture of the characters to determine whether there is abnormal behavior, such as climbing and fighting.

[0017] Preferably, Qwen2-VL comprises:

[0018] Powerful recognition capabilities, capable of accurately identifying multiple objects and their relationships in complex scenes, as well as handwritten text and multilingual image text including most European languages, Japanese, Korean, and Arabic;

[0019] Excellent visual reasoning skills, able to solve complex math problems through graphical analysis, able to extract information from real-world images and diagrams, and better follow instructions to solve real-world problems;

[0020] Long video comprehension and real-time conversation: It can understand video content of more than 20 minutes and continuously provide information and support in real-time conversation. It can be applied to video Q&A, conversation and content creation, etc.

[0021] Multi-language support: In addition to the common English and Chinese, it also supports understanding image text in multiple languages, making it convenient for users around the world to use;

[0022] The architecture is innovative, adopting the series structure of VIT and QWEN2, supporting native dynamic resolution and multi-modal rotation position embedding technology, and can process image input of any resolution. It can simultaneously capture and integrate multi-dimensional location information, take photos to answer questions, generate AI images, make phone calls, and send messages.

[0023] Preferably, the data processing and analysis module cooperates with the Doubao visual large model embedding module to, on the one hand, further organize the data and make logical judgments on the recognition and analysis results output by the large model, and correlate and integrate the relevant information collected at different times and angles; on the other hand, it uses the built-in algorithm to compare and analyze the real-time monitoring data with the historical monitoring data, and explore potential security trends and abnormal changes, such as the recent changes in the frequency of abnormal gatherings of people in a certain area, and generate a security risk assessment report based on this, providing detailed data support for subsequent security management decisions.

[0024] Preferably, the power management module is responsible for providing a stable and reliable power supply for the entire camera system. It can be connected to the mains power supply and have a built-in backup power supply, such as a lithium battery. It has power monitoring, intelligent switching and energy-saving control functions. When the mains power is cut off, it can automatically switch to the backup power supply to continue to maintain the normal operation of the camera. At the same time, the power supply power of each module can be dynamically adjusted according to the actual monitoring needs.

[0025] Preferably, the storage module is equipped with a large-capacity storage medium, such as a solid-state hard drive, for storing the collected original video image data and processed key information and analysis results. It supports data classification and storage in multiple ways according to time and event type, which is convenient for subsequent query, playback and data tracing operations, and the stored data can be backed up regularly to prevent loss.

[0026] A method for using a security management camera based on a vision model, characterized in that the method comprises the following steps:

[0027] S1. Installation and initialization: Install the camera at a suitable location for security monitoring, such as the entrance and exit of a building, the corridor, or the key passage area of a park, ensuring that the lens field of view of the image acquisition module covers the target monitoring area; Connect the mains power supply and turn on the camera. At this time, the power management module performs self-check and supplies power to each module normally, and the camera system starts the initialization program. The communication module automatically connects to a preset network, such as connecting to a local area network or accessing the Internet through a 4G / 5G network. The storage module completes the formatting preparation work and waits for data storage. The Doubao vision large model embedding module loads the pre-trained model parameters into the ready state;

[0028] S2. Image acquisition and analysis: The image acquisition module continuously acquires video images of the monitoring area at a set frame rate, such as 25 frames per second, and transmits the real-time image data to the Doubao vision large model embedding module; After receiving the image, the vision large model embedding module performs operations such as feature extraction, object recognition, and behavior analysis on each frame of the image. For example, it identifies the identity of the personnel in the picture, judges whether the walking direction and behavior of the personnel are compliant by comparing with the pre-stored authorized personnel image database, and at the same time, it also identifies various objects and their states in the scene, such as fire-fighting facilities and vehicles, and outputs the corresponding analysis results to the data processing and analysis module;

[0029] S3. Data processing and comprehensive judgment: After the data processing and analysis module collects the analysis results of the vision large model, on the one hand, it integrates the relevant data collected by cameras at different angles at the same time to construct a complete monitoring scene view; on the other hand, it combines historical monitoring data and judges the safety status of the current monitoring area through the built-in risk assessment algorithm, such as calculating the probability of abnormal behavior of current personnel and the level of safety hazards in the environment, and generating a security risk assessment report; For example, if it is found that a large number of strange people gather and behave abnormally in a certain area in a short period of time, and compared with the normal traffic flow data in this area in history, it is determined that this situation is a high-risk event and marked as a key concern situation;

[0030] S4. Information Transmission and Remote Monitoring: The communication module sends the collected original video image data, the analysis results of the vision large model, and the generated security risk assessment report to the remote monitoring and management platform and the mobile terminals of security personnel at a set cycle, such as every 1 minute or in real time for urgent high-risk events; security personnel can view the monitoring images of each camera, the analysis results, and the risk assessment situation in real time through the mobile APP or the interface of the monitoring and management platform, and remotely supervise the monitored area. If suspicious situations are found, they can also remotely control the camera to perform zooming and turning operations to obtain clearer and more accurate image information;

[0031] S5. Alarm Triggering and Response: When the data processing and analysis module determines that a security event reaching the preset alarm threshold occurs in the monitored area, such as detecting unauthorized personnel attempting to break into a restricted area or a fire and smoke situation, the alarm module is immediately activated; the audible and visual alarm emits strong audible and visual signals to deter potential dangerous personnel, and at the same time sends an alarm notification containing detailed event information, such as the location of the event and the type of the event, to the mobile terminals of security personnel and relevant responsible persons through the communication module. After receiving the notification, relevant personnel can quickly rush to the scene for handling or remotely command and dispatch to take corresponding countermeasures;

[0032] S6. Data Storage and Management: The storage module continuously stores the original video image data collected by the image acquisition module, classifies and archives it according to time sequence and event tags, such as using each alarm event and daily inspection period as tags, for convenient subsequent query and playback; at the same time, it also stores the processed analysis results and key information of the security risk assessment report, facilitating data statistics, trend analysis by management personnel, and serving as a basis for security management decisions. The stored data is regularly backed up to external storage devices or the cloud to ensure the security and integrity of the data.

[0033] The beneficial effects of the present invention are as follows:

[0034] 1. For the security management camera module and method based on a vision model described in the present invention, with the powerful image understanding ability of the Doubao vision large model, it can accurately identify various elements in the monitoring scene, greatly reducing misjudgment situations caused by factors such as environmental interference and object similarity. Whether it is day or night, in complex indoor and outdoor scenes and other different conditions, it can accurately capture key information related to security, improving the effectiveness of security monitoring.

[0035] 2. The security management camera module and method based on a vision model according to the present invention, through the comprehensive utilization of the output results of the vision large model and historical data by the data processing and analysis module, can not only know the current security situation, but also discover potential security hazard trends. The generated security risk assessment report can assist managers in formulating countermeasures in advance, changing passive response to active prevention, and enhancing the foresight and scientific nature of overall security management.

[0036] 3. The security management camera module and method based on a vision model according to the present invention, through the communication module, ensures that monitoring data and analysis results can be transmitted to the remote terminal in real time. Managers can master the situation and issue instructions anytime and anywhere without being at the monitoring site, greatly improving the convenience and timeliness of security management work, and is especially suitable for large-area and multi-region centralized security management scenarios.

[0037] 4. The security management camera module and method based on a vision model according to the present invention, through the power management module, ensures that the camera works stably in various power supply environments, avoiding monitoring blanks caused by power outages; while the convenient data storage and query functions of the storage module are conducive to subsequent event review, evidence search, etc., enhancing the practicality and reliability of the entire security management system.

[0038] 5. The security management camera module and method based on a vision model according to the present invention, the alarm module can quickly activate the corresponding alarm method according to the accurate judgment of security events, notify relevant personnel in the first time, helps to control security risks within the minimum range, avoid the further expansion of security accidents, and ensure the safety of personnel and property in the monitored area. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The present invention will be further described below with reference to the drawings.

[0040] Figure 1 is the camera module diagram in the present invention;

[0041] Figure 2 is the flowchart of the camera usage method in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] In order to make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the present invention will be further described below in conjunction with specific embodiments.

[0043] As Figure 1 and Figure 2 shown, a security management camera module based on a vision model according to an embodiment of the present invention is characterized in that: the security management camera module includes:

[0044] The image acquisition module, which consists of high-definition optical lenses and image sensor components, is responsible for collecting real-time video image information within the monitoring area;

[0045] The Doubao visual model embedding module receives the video image data from the image acquisition module. The visual model can accurately identify and classify the objects, people, and behaviors in the image.

[0046] Qwen2-VL, which can recognize complex scene object relationships, handwriting, and multi-language image text, has excellent visual reasoning ability and can solve problems through chart analysis. It can understand long videos and support real-time conversations for multiple applications. It supports multiple languages ​​to facilitate global users. With its innovative architecture, it can process images of any resolution, integrate multi-dimensional information, and expand its application capabilities by taking pictures to answer questions.

[0047] The data processing and analysis module, in collaboration with the Doubao visual large model embedding module, not only further organizes, judges and integrates the output results of the module, but also compares the real-time and historical monitoring data through algorithms, mines security trends and anomalies, generates evaluation reports, and assists in security management decision-making;

[0048] The communication module supports multiple communication protocols, such as Wi-Fi, Ethernet, and 4G / 5G. It can transmit the video image data collected by the camera, the analysis results of the Doubao visual model, and the security risk assessment report generated by the data processing and analysis module to the remote monitoring management platform or the mobile terminal of the security personnel in real time, ensuring the timely transmission of security information and facilitating the management personnel to remotely control the status of the monitoring area and respond quickly;

[0049] The power management module provides a stable and reliable power supply for the entire camera system. It can be connected to the mains power supply and has a built-in backup power supply such as a lithium battery. It has power monitoring, intelligent switching and energy-saving control functions.

[0050] The storage module is equipped with a large-capacity storage medium, such as a solid-state hard disk, which is mainly used to store the original video image data collected by the camera, as well as the key information and analysis results obtained after processing in various links;

[0051] The alarm module works with the data processing and analysis module to monitor the safety status of the surveillance area; triggers multiple alarm modes, deters potentially dangerous persons by emitting sound and light signals, and pushes messages and voice prompts to the manager's mobile terminal.

[0052] The image acquisition module has functions such as adjustable focal length and aperture, and can adapt to the needs of collecting clear images under different distances and lighting conditions. The collected image data is transmitted in real time in the form of digital signals to subsequent modules for processing.

[0053] The Doubao visual big model embedding module has a pre-trained Doubao visual big model built in. The model is deeply trained based on a large amount of image data of different scenes, various objects, and human behaviors, and has powerful image feature extraction, semantic understanding, and pattern recognition capabilities. After receiving video image data from the image acquisition module, the visual big model can accurately identify and classify objects, people, and behavioral actions in the image. For example, it can accurately distinguish whether they are normal pedestrians, staff or suspicious persons, identify specific dangerous items, abnormal scene arrangements, such as blocked fire escape routes, etc. At the same time, it can also analyze and judge the behavioral trajectory and posture of the characters to determine whether there are abnormal behaviors, such as climbing and fighting.

[0054] Qwen2-VL includes:

[0055] Powerful recognition capabilities, capable of accurately identifying multiple objects and their relationships in complex scenes, as well as handwritten text and multilingual image text including most European languages, Japanese, Korean, and Arabic;

[0056] Excellent visual reasoning skills, able to solve complex math problems through graphical analysis, able to extract information from real-world images and diagrams, and better follow instructions to solve real-world problems;

[0057] Long video comprehension and real-time conversation: It can understand video content of more than 20 minutes and continuously provide information and support in real-time conversation. It can be applied to video Q&A, conversation and content creation, etc.

[0058] Multi-language support: In addition to the common English and Chinese, it also supports understanding image text in multiple languages, making it convenient for users around the world to use;

[0059] The architecture is innovative, adopting the series structure of VIT and QWEN2, supporting native dynamic resolution and multi-modal rotation position embedding technology, and can process image input of any resolution. It can simultaneously capture and integrate multi-dimensional location information, take photos to answer questions, generate AI images, make phone calls, and send messages.

[0060] The data processing and analysis module works in conjunction with the Doubao visual large model embedding module. On the one hand, it further organizes data and makes logical judgments on the recognition and analysis results output by the large model, and correlates and integrates relevant information collected at different times and angles; on the other hand, it uses built-in algorithms to compare and analyze real-time monitoring data with historical monitoring data, and explore potential security trends and abnormal changes, such as the recent changes in the frequency of abnormal gatherings of people in a certain area. Based on this, a security risk assessment report is generated to provide detailed data support for subsequent security management decisions.

[0061] The power management module is responsible for providing a stable and reliable power supply for the entire camera system. It can be connected to the mains power supply and has a built-in backup power supply, such as a lithium battery. It has functions of power monitoring, intelligent switching, and energy-saving control. When the mains power fails, it can automatically switch to the backup power supply to continue maintaining the normal operation of the camera. At the same time, it can dynamically adjust the power supply power of each module according to the actual monitoring requirements.

[0062] The storage module is equipped with a large-capacity storage medium, such as a solid-state drive, for storing the collected original video image data, as well as the processed key information and analysis result content. It supports data classification storage in multiple ways according to time, event type, etc., which is convenient for subsequent query, playback, and data traceability operations. And the stored data can be backed up regularly to prevent loss.

[0063] All modules of the entire camera system work together. The image acquisition module continuously acquires images. After in-depth analysis and processing by the Doubao vision large model embedding module and the data processing and analysis module, it transmits information outward through the communication module. The storage module saves data, the power management module ensures power supply, and the alarm module issues an alarm in a timely manner when necessary, thus forming a complete intelligent security management and monitoring system.

[0064] A method for using a security management camera based on a vision model, characterized in that the method comprises the following steps:

[0065] S1. Installation and initialization: Install the camera in a suitable location for security monitoring, such as the entrance and exit of a building, the corridor, or the key passage area of a park, to ensure that the lens field of view of the image acquisition module covers the target monitoring area; Connect the mains power supply and turn on the camera. At this time, the power management module performs self-check and supplies power to each module normally, and the camera system starts the initialization program. The communication module automatically connects to a preset network, such as connecting to a local area network or accessing the Internet through a 4G / 5G network. The storage module completes the formatting preparation work and waits for data storage. The Doubao vision large model embedding module loads the pre-trained model parameters and enters the ready state;

[0066] S2. Image acquisition and analysis: The image acquisition module continuously acquires video images of the monitoring area at a set frame rate, such as 25 frames per second, and transmits the real-time image data to the Doubao vision large model embedding module; After receiving the image, the vision large model embedding module performs operations such as feature extraction, object recognition, and behavior analysis on each frame of the image. For example, it identifies the identity of the personnel in the picture, judges whether the walking direction and behavior actions of the personnel are compliant by comparing with the pre-stored authorized personnel image database, and at the same time, it also identifies various objects and their states in the scene, such as fire-fighting facilities and vehicles, and outputs the corresponding analysis results to the data processing and analysis module;

[0067] S3. Data Processing and Comprehensive Judgment: After the data processing and analysis module collects the analysis results of the visual large model, on the one hand, it integrates the relevant data collected by cameras from different angles at the same moment to construct a complete view of the monitoring scenario; on the other hand, it combines historical monitoring data and uses the built-in risk assessment algorithm to judge the safety status of the current monitoring area, such as calculating the probability of abnormal behavior of current personnel and the level of potential safety hazards in the environment, and generating a safety risk assessment report. For example, if it is found that a large number of strange people gather and behave abnormally in a certain area in a short period of time, and compared with the normal pedestrian flow data in this area in history, it is determined that this situation is a high-risk event and marked as a key concern situation.

[0068] S4. Information Transmission and Remote Monitoring: The communication module sends the collected original video image data, the analysis results of the visual large model, and the generated safety risk assessment report to the remote monitoring management platform and the mobile terminals of security personnel at a set cycle, such as every 1 minute or in real time for urgent high-risk events; security personnel can view the monitoring images, analysis results, and risk assessment situations of each camera in real time through the mobile APP or the interface of the monitoring management platform, and remotely supervise the monitoring area. If suspicious situations are found, they can also remotely control the camera to zoom in and turn to obtain clearer and more accurate image information.

[0069] S5. Alarm Triggering and Response: When the data processing and analysis module determines that a safety event reaching the preset alarm threshold occurs in the monitoring area, such as detecting unauthorized personnel attempting to break into a restricted area or a fire and smoke situation, the alarm module is immediately activated; the audible and visual alarm emits strong audible and visual signals to deter potential dangerous personnel, and at the same time sends an alarm notification containing detailed event information, such as the location and type of the event, to the mobile terminals of security personnel and relevant responsible persons through the communication module. After receiving the notification, relevant personnel can quickly rush to the scene for handling or remotely command and dispatch to take corresponding countermeasures.

[0070] S6. Data Storage and Management: The storage module continuously stores the original video image data collected by the image acquisition module, classifying and archiving it according to time sequence and event tags, such as using each alarm event and daily inspection period as tags, for convenient subsequent query and playback; at the same time, it also stores the processed analysis results and key information of the safety risk assessment report, facilitating data statistics and trend analysis by management personnel and serving as a basis for safety management decisions. The stored data is regularly backed up to external storage devices or the cloud to ensure data security and integrity.

[0071] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.

Claims

1. A security management camera module based on a visual model, characterized in that: The security management camera module includes: The image acquisition module, which consists of high-definition optical lenses and image sensor components, is responsible for collecting real-time video image information within the monitoring area; The Doubao visual model embedding module receives the video image data from the image acquisition module. The visual model can accurately identify and classify the objects, people, and behaviors in the image. Qwen2-VL, which can recognize complex scene object relationships, handwriting, and multi-language image text, has excellent visual reasoning ability and can solve problems through chart analysis. It can understand long videos and support real-time conversations for multiple applications. It supports multiple languages ​​to facilitate global users. With its innovative architecture, it can process images of any resolution, integrate multi-dimensional information, and expand its application capabilities by taking pictures to answer questions. The data processing and analysis module, in collaboration with the Doubao visual large model embedding module, not only further organizes, judges and integrates the output results of the module, but also compares the real-time and historical monitoring data through algorithms, mines security trends and anomalies, generates evaluation reports, and assists in security management decision-making; The communication module supports multiple communication protocols, such as Wi-Fi, Ethernet, and 4G / 5G. It can transmit the video image data collected by the camera, the analysis results of the Doubao visual model, and the security risk assessment report generated by the data processing and analysis module to the remote monitoring management platform or the mobile terminal of the security personnel in real time, ensuring the timely transmission of security information and facilitating the management personnel to remotely control the status of the monitoring area and respond quickly; The power management module provides a stable and reliable power supply for the entire camera system. It can be connected to the mains power supply and has a built-in backup power supply such as a lithium battery. It has power monitoring, intelligent switching and energy-saving control functions. The storage module is equipped with a large-capacity storage medium, such as a solid-state hard disk, which is mainly used to store the original video image data collected by the camera, as well as the key information and analysis results obtained after processing in various links; The alarm module works with the data processing and analysis module to monitor the safety status of the surveillance area; triggers multiple alarm modes, deters potentially dangerous persons by emitting sound and light signals, and pushes messages and voice prompts to the manager's mobile terminal.

2. The security management camera module based on visual model according to claim 1, characterized in that: The image acquisition module has functions such as adjustable focal length and aperture, and can adapt to the needs of collecting clear images under different distances and lighting conditions. The collected image data is transmitted in real time in the form of digital signals to subsequent modules for processing.

3. The security management camera module based on visual model according to claim 2, characterized in that: The Doubao visual big model embedding module has a pre-trained Doubao visual big model built in. The model is deeply trained based on a large amount of image data of different scenes, various objects, and human behaviors, and has powerful image feature extraction, semantic understanding, and pattern recognition capabilities. After receiving video image data from the image acquisition module, the visual big model can accurately identify and classify objects, people, and behavioral actions in the image. For example, it can accurately distinguish whether they are normal pedestrians, staff or suspicious persons, identify specific dangerous items, abnormal scene arrangements, such as blocked fire escape routes, etc. At the same time, it can also analyze and judge the behavioral trajectory and posture of the characters to determine whether there are abnormal behaviors, such as climbing and fighting.

4. The security management camera module based on visual model according to claim 3, characterized in that: Qwen2-VL includes: Powerful recognition capabilities, capable of accurately identifying multiple objects and their relationships in complex scenes, as well as handwritten text and multilingual image text including most European languages, Japanese, Korean, and Arabic; Excellent visual reasoning skills, able to solve complex math problems through graphical analysis, able to extract information from real-world images and diagrams, and better follow instructions to solve real-world problems; Long video comprehension and real-time conversation: It can understand video content of more than 20 minutes and continuously provide information and support in real-time conversation. It can be applied to video Q&A, conversation and content creation, etc. Multi-language support: In addition to the common English and Chinese, it also supports understanding image text in multiple languages, making it convenient for users around the world to use; The architecture is innovative, adopting the series structure of VIT and QWEN2, supporting native dynamic resolution and multi-modal rotation position embedding technology, and can process image input of any resolution. It can simultaneously capture and integrate multi-dimensional location information, take photos to answer questions, generate AI images, make phone calls, and send messages.

5. The security management camera module based on visual model according to claim 4, characterized in that: The data processing and analysis module works in conjunction with the Doubao visual large model embedding module. On the one hand, it further organizes data and makes logical judgments on the recognition and analysis results output by the large model, and correlates and integrates relevant information collected at different times and angles; on the other hand, it uses built-in algorithms to compare and analyze real-time monitoring data with historical monitoring data, and explore potential security trends and abnormal changes, such as the recent changes in the frequency of abnormal gatherings of people in a certain area. Based on this, a security risk assessment report is generated to provide detailed data support for subsequent security management decisions.

6. The visual model-based security management camera module according to claim 5, characterized in that: The power management module is responsible for providing a stable and reliable power supply for the entire camera system. It can be connected to the AC power supply and have a built-in backup power supply, such as a lithium battery. It has power monitoring, intelligent switching and energy-saving control functions. When the AC power is cut off, it can automatically switch to the backup power supply to continue to maintain the normal operation of the camera. At the same time, the power supply of each module can be dynamically adjusted according to actual monitoring needs.

7. The visual model-based security management camera module according to claim 6, characterized in that: The storage module is equipped with a large-capacity storage medium, such as a solid-state hard drive, which is used to store the collected original video image data and the processed key information and analysis results. It supports data classification and storage in multiple ways according to time and event type, which is convenient for subsequent query, playback and data tracing operations. The stored data can be backed up regularly to prevent loss.

8. The visual model-based security management camera module according to claim 7, characterized in that: The alarm module is connected to the data processing and analysis module. When it is determined that a security incident reaching the preset danger level occurs in the monitoring area based on the analysis and judgment of the Doubao visual large model and the comprehensive data processing results, such as illegal intrusion or fire hazard, the module can trigger multiple alarm methods, including sending sound and light alarm signals to deter potentially dangerous persons and sending alarm notifications to the mobile terminals of relevant managers, such as push messages and voice prompts, to ensure that security incidents can be handled in a timely manner.

9. A method for using a visual model-based security management camera, using a visual model-based security management camera module according to any one of claims 1 to 8, characterized in that: The method of use includes the following steps: S1. Installation and initialization: Install the camera at a suitable location where security monitoring is required, such as the entrance and exit of a building, corridors, and key channel areas of a park, to ensure that the lens field of view of the image acquisition module covers the target monitoring area; connect the AC power supply and turn on the camera. At this time, the power management module performs a self-check and supplies power to each module normally. The camera system starts the initialization program, and the communication module automatically connects to the preset network, such as connecting to the local area network or accessing the Internet through a 4G / 5G network. The storage module completes formatting preparations and waits for data storage. The Doubao visual large model embedding module loads the pre-trained model parameters and enters the ready state; S2, image acquisition and analysis. The image acquisition module continuously acquires video images of the monitored area at a set frame rate, such as 25 frames per second, and transmits real-time image data to the Doubao visual large model embedding module. After receiving the image, the visual large model embedding module performs feature extraction, object recognition, and behavior analysis on each frame of the image. For example, it identifies the identity of the person in the picture, and determines whether the walking direction and behavior of the person are compliant by comparing with the pre-stored authorized personnel image database. At the same time, it also identifies various objects in the scene and their states, such as fire-fighting facilities and vehicles, and outputs the corresponding analysis results to the data processing and analysis module. S3. Data processing and comprehensive judgment. After the data processing and analysis module collects the analysis results of the large visual model, it integrates the relevant data collected by cameras at different angles at the same time to build a complete monitoring scene view. On the other hand, it combines historical monitoring data and uses the built-in risk assessment algorithm to judge the safety status of the current monitoring area. For example, it calculates the probability of abnormal behavior of the current personnel and the level of safety hazards in the environment, and generates a safety risk assessment report. For example, if a large number of unfamiliar people gather in a certain area in a short period of time and behave abnormally, combined with the historical data of normal traffic in the area, it is judged as a high-risk event and marked as a key concern. S4, information transmission and remote monitoring. The communication module sends the collected original video image data, the analysis results of the visual large model and the generated security risk assessment report to the remote monitoring management platform and the mobile terminal of the security personnel in real time according to the set cycle, such as every 1 minute or for urgent high-risk events; the security personnel can view the monitoring screen, analysis results and risk assessment of each camera in real time through the mobile phone APP or the interface of the monitoring management platform, remotely supervise the monitoring area, and remotely control the camera to zoom and turn in order to obtain clearer and more accurate image information if suspicious situations are found; S5, alarm triggering and response, when the data processing and analysis module determines that a security event that reaches the preset alarm threshold occurs in the monitoring area, such as detecting an unauthorized person trying to break into a restricted area or a fire and smoke situation, the alarm module is immediately activated; the sound and light alarm sends out a strong sound and light signal to deter potentially dangerous persons, and at the same time, an alarm notification containing detailed information of the event, such as the location of the event and the type of event, is sent to the mobile terminal of the security personnel and the relevant person in charge through the communication module. After receiving the notification, the relevant personnel can quickly rush to the scene to deal with it, or remotely command and dispatch to take corresponding response measures; S6. Data storage and management. The storage module continuously stores the original video image data collected by the image acquisition module, and classifies and archives them according to chronological order and event labels, such as each alarm event and daily patrol period, to facilitate subsequent query and playback; it also stores processed analysis results and key information of security risk assessment reports to facilitate management personnel to conduct data statistics, trend analysis, and as a basis for security management decisions. The stored data is regularly backed up to external storage devices or the cloud to ensure data security and integrity.

Citation Information

Cited By

  • Environmental protection risk AI visual early warning terminal

    CN120957010A

  • Environmental protection risk ai visual early warning terminal

    CN120957010B

  • Multi-dimensional vehicle picture quality inspection auditing method and system

    CN121259511A