Scene monitoring methods, devices, equipment, and media based on multimodal large models

By combining a lightweight multimodal large model on the edge with a high-precision model on the cloud platform, the problems of high false alarm rate, high resource consumption and difficulty in identifying complex scenes in existing scene monitoring technologies are solved, achieving efficient and accurate scene monitoring and alarm decision-making.

CN119743570BActive Publication Date: 2025-10-28E SURFING IOT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411679418.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-22
Publication Date
2025-10-28
Estimated Expiration
2044-11-22

AI Technical Summary

Technical Problem

Existing scene monitoring solutions suffer from problems such as high false alarm rates, low efficiency of manual verification, inability of traditional AI algorithms to understand complex scenes, and high resource consumption of large-scale multimodal models that can only be deployed in the cloud.

Method used

A lightweight multimodal large model is deployed on the edge for initial identification, and a high-precision video scene recognition model is deployed on the cloud platform. Through the collaborative work of the multimodal large model, accurate understanding of video scenes and alarm decisions are achieved.

Benefits of technology

It improves the efficiency and accuracy of scene monitoring, reduces resource consumption, and enables real-time identification and efficient alarm decision-making for complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119743570B_ABST
    Figure CN119743570B_ABST
Patent Text Reader

Abstract

This invention discloses a scene monitoring method, apparatus, device, and medium based on a multimodal large model. The method includes acquiring a first video stream of a target scene; determining a corresponding set of video frame images based on the first video stream; inputting the video frame image set into a multimodal large model deployed on the edge to obtain corresponding video scene description information; and then uploading the first video stream and the video scene description information to a cloud platform; matching the corresponding video scene recognition model on the cloud platform based on the video scene description information; performing scene recognition on the first video stream based on the video scene recognition model to obtain a scene recognition result; adjusting the video scene description information based on the scene recognition result; and determining whether to issue an alarm based on the adjusted video scene description information. This invention improves the efficiency and accuracy of scene monitoring and can be widely applied in the field of video surveillance technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video surveillance technology, and in particular to a scene monitoring method, device, equipment and medium based on a multimodal large model. Background Technology

[0002] With the further advancement of smart city construction, the number of IoT devices such as cameras is growing exponentially. Efficiently and accurately identifying the information collected by these cameras is crucial for urban governance. Currently, traditional visual monitoring solutions based on cameras are widely used in various fields of production and daily life. For example, license plate recognition cameras installed at highway checkpoints can accurately identify the license plates of passing vehicles; high-altitude cameras installed in mountainous areas capture and monitor wildfires; and garbage monitoring cameras are installed in streets and alleys. These cameras, equipped with traditional AI algorithms, can achieve monitoring and classification in traditional single-scene scenarios, but they fundamentally cannot understand the video. For example, garbage monitoring cameras installed in streets and alleys cannot understand whether a serious fight has occurred on the street; license plate recognition cameras installed at traffic checkpoints cannot identify whether a traffic accident has occurred. With the development of multimodal large-scale model technology, it has become possible to integrate various information such as surveillance video and audio in communities, industrial parks, factories, and other places to achieve more accurate intelligent monitoring and early warning.

[0003] Existing scene monitoring solutions mainly include the following:

[0004] 1) AI Camera Monitoring. This primarily involves installing traditional CV models on camera devices for target object recognition in security, fire protection, and other fields. In actual scene monitoring, AI cameras are used to monitor violations and abnormal phenomena. When an anomaly occurs, event verification is performed by retrieving relevant camera videos to confirm the reported event. When a violation or event occurs, notifications and broadcasts can be linked via email, voice calls, SMS, WeChat, LED screens, and smart speakers. Alarm events are tiered and issued with initial alerts to the responsible person. If the issue is not addressed within the specified time, it is pushed to relevant leaders. Weekly and monthly production safety reports are also provided. After the event is pushed out, manual verification and processing are performed, followed by photo and video verification for confirmation.

[0005] 2) AI Box-Linked Camera Monitoring. This primarily utilizes AI boxes to AI-enable traditional cameras, giving them video image monitoring capabilities. It's often used for upgrading traditional cameras. The event monitoring, verification, alarm linkage, alarm push notifications, and case closure verification processes remain consistent with the above. Taking security scenarios as an example, traditional security addresses the issues of "seeing" and "seeing clearly," while intelligent security addresses the issue of "understanding." Previously, video was reviewed manually; now, intelligent security records useful information like "traffic flow" while filtering out useless information like "the slightest movement."

[0006] 3) Cloud-based multimodal model monitoring. This method extracts keyframes from videos and transmits them to a large multimodal model in the cloud for recognition. While this method can understand image information and identify some events to a certain extent, the information utilization rate is very low because it relies on keyframes. Secondly, the model scale is large, generally above 30 bits, resulting in significant hardware consumption and requiring cloud deployment. Furthermore, the real-time performance of image recognition is low, making it impossible to generate timely alarms.

[0007] In summary, existing scene monitoring solutions have the following drawbacks:

[0008] 1) Numerous false alarms and inefficient manual verification. Alarms generated by monitoring require manual review of video streams, making automated and efficient processing impossible. Therefore, in practical applications, insufficient manpower for verification results in a large amount of video alarm data not being fully utilized.

[0009] 2) Traditional AI algorithms with camera-mounted CV models can only "see" and "see clearly," but cannot understand or solve the problem of "understanding." A single CV model is difficult to meet the requirements for dealing with complex multi-scene discrimination.

[0010] 3) Large-scale multimodal large models commonly used in the industry consume a lot of resources and can only be deployed in the cloud. At the same time, in order to support a large number of monitoring devices, the real-time performance of local alarms is a huge challenge.

[0011] Terminology Explanation:

[0012] MultiModal Large Language Models (MM-LLMs): Multimodal large language models.

[0013] Computer Vision (CV)

[0014] Prompt Engineering (PE): Prompt Engineering. Summary of the Invention

[0015] The purpose of this invention is to at least partially solve one of the technical problems existing in the prior art.

[0016] Therefore, one objective of this invention is to provide a scene monitoring method based on a multimodal large model, which improves the efficiency and accuracy of scene monitoring.

[0017] Another objective of this invention is to provide a scene monitoring device based on a multimodal large model.

[0018] To achieve the above-mentioned technical objectives, the technical solutions adopted in the embodiments of the present invention include:

[0019] On one hand, embodiments of the present invention provide a scene monitoring method based on a multimodal large model, comprising the following steps:

[0020] Acquire the first video stream of the target scene, and determine the corresponding video frame image set based on the first video stream;

[0021] The video frame image set is input into a multimodal large model deployed on the edge to obtain the corresponding video scene description information, and then the first video stream and the video scene description information are uploaded to the cloud platform.

[0022] Based on the video scene description information, a corresponding video scene recognition model is matched on the cloud platform, and the first video stream is used for scene recognition based on the video scene recognition model to obtain the scene recognition result;

[0023] The video scene description information is adjusted based on the scene recognition result, and an alarm is issued based on the adjusted video scene description information.

[0024] Furthermore, in one embodiment of the present invention, determining the corresponding video frame image set based on the first video stream specifically includes:

[0025] The first video stream is divided into multiple video segments according to a preset time length;

[0026] Video frames of each video segment are extracted according to a preset sampling frequency to obtain a set of video frame images corresponding to each video segment;

[0027] The image set number of the video frame image set is determined based on the start time, end time, and corresponding camera code of the video segment.

[0028] Furthermore, in one embodiment of the present invention, the step of inputting the video frame image set into a multimodal large model deployed on the edge to obtain corresponding video scene description information specifically includes:

[0029] Generate the first prompt statement based on the preset first prompt word template;

[0030] The video frame image set and the first prompt statement are input into the multimodal large model to obtain the video scene description information;

[0031] The video scene description information includes video content description, event type description, and event nature description.

[0032] Furthermore, in one embodiment of the present invention, the step of matching the corresponding video scene recognition model on the cloud platform according to the video scene description information specifically includes:

[0033] Obtain the scene recognition model library stored on the cloud platform;

[0034] The target model type is determined based on the event type description, and then matched in the scene recognition model library according to the target model type to obtain the corresponding video scene recognition model.

[0035] Furthermore, in one embodiment of the present invention, the scene recognition model library is constructed through the following steps:

[0036] Obtain pre-trained video scene recognition models for multiple sub-domains, and determine the model type label of the corresponding video scene recognition model based on the sub-domain;

[0037] The model type label is used as the key value, and the corresponding video scene recognition model is used as the value value to generate model key-value pairs;

[0038] The scene recognition model library is constructed based on the model key-value pairs.

[0039] Furthermore, in one embodiment of the present invention, adjusting the video scene description information based on the scene recognition result specifically includes:

[0040] Determine whether the scene recognition result is consistent with the event type description;

[0041] When the scene recognition result matches the event type description, the video content description is completed based on the scene recognition result to obtain the adjusted video scene description information;

[0042] When the scene recognition result is inconsistent with the event type description, the first video stream is subjected to scene recognition based on the scene recognition big model deployed on the cloud platform to obtain the adjusted video scene description information.

[0043] Furthermore, in some optional embodiments, the step of determining whether to issue an alarm based on the adjusted video scene description information specifically includes:

[0044] Generate a second prompt statement based on a preset second prompt word template;

[0045] The adjusted video scene description information and the second prompt statement are input into the alarm decision model deployed on the cloud platform to obtain the alarm decision description;

[0046] Determine whether to issue an alarm based on the alarm decision description. If an alarm is issued, determine the alarm cause description based on the alarm decision description.

[0047] On the other hand, embodiments of the present invention provide a scene monitoring device based on a multimodal large model, comprising:

[0048] The video frame extraction module is used to acquire the first video stream of the target scene and determine the corresponding video frame image set based on the first video stream.

[0049] The edge recognition module is used to input the video frame image set into the multimodal large model deployed on the edge to obtain the corresponding video scene description information, and then upload the first video stream and the video scene description information to the cloud platform;

[0050] The cloud platform recognition module is used to match the corresponding video scene recognition model on the cloud platform according to the video scene description information, and to perform scene recognition on the first video stream according to the video scene recognition model to obtain the scene recognition result;

[0051] The alarm decision module is used to adjust the video scene description information according to the scene recognition result, and determine whether to issue an alarm based on the adjusted video scene description information.

[0052] On the other hand, embodiments of the present invention provide an electronic device, which includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for enabling communication between the processor and the memory. When the program is executed by the processor, it implements the scene monitoring method based on a multimodal large model as described above.

[0053] On the other hand, embodiments of the present invention also provide a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, which can be executed by one or more processors to implement the scene monitoring method based on a multimodal large model as described above.

[0054] The advantages and beneficial effects of the present invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention:

[0055] This invention acquires a first video stream of a target scene, determines a corresponding set of video frame images based on the first video stream, inputs the set of video frame images into a multimodal large model deployed on the edge, obtains corresponding video scene description information, and then uploads the first video stream and video scene description information to a cloud platform. Based on the video scene description information, a corresponding video scene recognition model is matched on the cloud platform, and scene recognition is performed on the first video stream according to the video scene recognition model to obtain scene recognition results. The video scene description information is adjusted based on the scene recognition results, and an alarm is issued based on the adjusted video scene description information. This invention first uses a lightweight multimodal large model deployed on the edge to perform preliminary recognition of the acquired video images to obtain video scene description information. Then, based on this video scene description information, a video scene recognition model corresponding to a specific sub-domain is matched on the cloud platform. A high-precision video scene recognition model deployed on the cloud platform is used to further recognize the scene in the video images. This allows for adjustment of the video scene description information based on the scene recognition results, resulting in more accurate and detailed video scene description information. Finally, alarm decisions are made based on the adjusted video scene description information, improving the efficiency and accuracy of scene monitoring. Attached Figure Description

[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the embodiments of the present invention are described below. It should be understood that the drawings described below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0057] Figure 1 A flowchart illustrating the steps of a scene monitoring method based on a multimodal large model provided in an embodiment of the present invention;

[0058] Figure 2 A flowchart of step S101 provided in an embodiment of the present invention;

[0059] Figure 3 A flowchart of step S102 provided in an embodiment of the present invention;

[0060] Figure 4 A flowchart of step S103 provided in an embodiment of the present invention;

[0061] Figure 5 A flowchart illustrating the steps involved in constructing a scene recognition model library, as provided in an embodiment of the present invention.

[0062] Figure 6 A flowchart of step S104 provided in an embodiment of the present invention;

[0063] Figure 7 Another flowchart of step S104 provided in an embodiment of the present invention;

[0064] Figure 8 A schematic diagram of the structure of the scene monitoring device based on a multimodal large model provided in an embodiment of the present invention;

[0065] Figure 9 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention;

[0066] Figure 10 This is a schematic diagram of the structure of the storage medium provided in an embodiment of the present invention. Detailed Implementation

[0067] The embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application. It should be noted that although functional modules are divided in the system schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the system schematic diagram or the order in the flowchart. The step numbers in the following embodiments are only set for ease of explanation and do not limit the order between steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0068] In the description of this invention, "multiple" means two or more. The use of "first" and "second" is for distinguishing technical features only and should not be construed as indicating or implying relative importance, the number of indicated technical features, or the order of the indicated technical features. Furthermore, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0069] The scene monitoring method based on a multimodal large model provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, set-top box, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the scene monitoring method based on a multimodal large model, etc., but is not limited to the above forms.

[0070] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0071] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards of the relevant countries and regions. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirects to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data for the proper functioning of the embodiments of this application obtained.

[0072] like Figure 1 The diagram shown is a flowchart of a scene monitoring method based on a multimodal large model provided by an embodiment of the present invention. (Refer to...) Figure 1This invention provides a scene monitoring method based on a multimodal large model, specifically including the following steps:

[0073] S101. Obtain the first video stream of the target scene, and determine the corresponding video frame image set based on the first video stream;

[0074] S102. Input the video frame image set into the multimodal large model deployed on the terminal side to obtain the corresponding video scene description information, and then upload the first video stream and video scene description information to the cloud platform.

[0075] S103. Match the corresponding video scene recognition model on the cloud platform according to the video scene description information, and perform scene recognition on the first video stream according to the video scene recognition model to obtain the scene recognition result;

[0076] S104. Adjust the video scene description information according to the scene recognition results, and determine whether to issue an alarm based on the adjusted video scene description information.

[0077] This invention enhances the detection and cognitive capabilities of AI camera surveillance scenes by coupling a small-scale, lightweight multimodal model deployed on the edge device with a high-precision video scene recognition model deployed on a cloud platform across different sub-domains. This allows for full utilization of video stream data on the edge device, enabling the understanding and recognition of complex scenes. Through a prompt project, adaptive algorithms are matched between key and general scene monitoring. This allows for further identification of key single scenes by calling the high-precision video scene recognition model on the cloud platform, thereby verifying the edge monitoring results and enabling the identification and detection of undefined sudden disaster events in complex and unpredictable scenarios. Furthermore, this invention utilizes the recognition results from both the edge and cloud platforms for comprehensive analysis and judgment, providing alarm strategies. Compared to manual judgment, this significantly increases processing efficiency and supports the discriminative analysis of multi-service monitoring and early warning for massive monitoring devices.

[0078] In this embodiment of the invention, a lightweight multimodal large model deployed on the edge is first used to perform preliminary identification on the acquired video images to obtain video scene description information. Then, based on the video scene description information, a video scene recognition model corresponding to the subdivided field is matched on the cloud platform. The high-precision video scene recognition model deployed on the cloud platform is used to further identify the scene in the video images. Thus, the video scene description information can be adjusted according to the scene recognition results to obtain more accurate and detailed video scene description information. In turn, alarm decisions can be made based on the adjusted video scene description information, thereby improving the efficiency and accuracy of scene monitoring.

[0079] It should be noted that the focus of this invention is on the deployment of models on the cloud platform and the edge, as well as the process of mutual cooperation to achieve scene monitoring. The specific model training process is not the focus of this invention. The training of relevant models can adopt the training methods of corresponding models in the prior art, or can directly utilize existing relevant models in the prior art. This invention will not elaborate on these aspects.

[0080] like Figure 2 The diagram shown is a flowchart of step S101 provided in an embodiment of the present invention. (Refer to...) Figure 2 As an optional implementation, the corresponding video frame image set is determined based on the first video stream, specifically including:

[0081] S1011. Divide the first video stream into multiple video segments according to the preset time length;

[0082] S1012. Extract video frames of each video segment according to the preset sampling frequency to obtain the video frame image set corresponding to each video segment;

[0083] S1013. Determine the image set number of the video frame image set based on the start time, end time of the video segment and the corresponding camera code.

[0084] Specifically, the video stream is divided into multiple video segments according to the time length T, where T is less than or equal to 24h; video frames of the video segments are extracted at 15 frames / s and packaged into a set of images; the image set is numbered according to the rule of "start time + end time + camera code + unique value + number", for example, images-202409111420-111520-vabc-123.

[0085] like Figure 3 The diagram shown is a flowchart of step S102 provided in an embodiment of the present invention. (Refer to...) Figure 3 As an optional implementation, the video frame image set is input into a multimodal large model deployed on the edge to obtain the corresponding video scene description information, which specifically includes:

[0086] S1021. Generate the first prompt statement according to the preset first prompt word template;

[0087] S1022. Input the video frame image set and the first prompt statement into the multimodal large model to obtain video scene description information;

[0088] The video scene description information includes video content description, event type description, and event nature description.

[0089] Specifically, the image set is fed into the multimodal large model, the preset prompt words are edited, the model is guided to output its understanding of the video, generate a detailed description of the video, and then the first video stream and video scene description information are uploaded to the cloud platform.

[0090] For example, inputting the image set "images" and the prompt "msgs":[{"role":"user","content":"Classify the video event as neutral, negative, or positive\nStrictly follow the format below to answer. For example: [Video Content]: A fight broke out; [Event Type]: Fire; [Event Nature]: Negative. \n[Video Content]: XXXXXXXX (describe the video content in detail); [Tags]: XXXX (choose only one: a. Fire, b. Traffic congestion, c. Flooding, d. Injuries, e. Garbage accumulation, f. Cannot be defined); [Event Classification]: XX (answer only positive, negative, or neutral)." The model returns the information: "result":["[Video Content]: A fire broke out, grass and trees along the road were burning, no casualties reported; [Event Type]: Fire; [Event Nature]: Negative."

[0091] In the example above, the model returns information including video content description, event type description, and event nature description. Then, by using "images-202409111420-111520-vabc-123" as the text number and "result" as the text information, the video scene description information of the corresponding video segment can be obtained.

[0092] like Figure 4 The diagram shown is a flowchart of step S103 provided in an embodiment of the present invention. (Refer to...) Figure 4 As an optional implementation, the corresponding video scene recognition model is matched on the cloud platform based on the video scene description information, which specifically includes:

[0093] S1031. Obtain the scene recognition model library stored on the cloud platform;

[0094] S1032. Determine the target model type based on the event type description, and match it in the scene recognition model library according to the target model type to obtain the corresponding video scene recognition model.

[0095] Specifically, the cloud platform deploys a scene recognition model library, which includes high-precision video scene recognition models of different model types, such as fire recognition models, traffic congestion recognition models, flood and waterlogging recognition models, personnel injury recognition models, and waste recognition models. These high-precision video scene recognition models can not only accurately identify the corresponding video scenes, but also quantify their severity, such as the severity of a fire or the degree of traffic congestion.

[0096] Based on the video scene description information uploaded by the client, the corresponding high-precision video scene recognition model is matched. For example, if the event type is described as "fire", the target model type is determined to be "fire recognition". Then, the high-precision video scene recognition model of the corresponding model type is searched in the scene recognition model library for further scene recognition, thereby determining whether a fire has occurred and the scope, source, severity, and whether there are any casualties, etc.

[0097] like Figure 5 The diagram shown is a flowchart illustrating one step in constructing a scene recognition model library according to an embodiment of the present invention. (Refer to...) Figure 5 As an optional implementation, the scene recognition model library is constructed through the following steps:

[0098] S201. Obtain pre-trained video scene recognition models for multiple sub-domains, and determine the model type label of the corresponding video scene recognition model according to the sub-domain.

[0099] S202. Use the model type label as the key value and the corresponding video scene recognition model as the value value to generate model key-value pairs;

[0100] S203. Construct a scene recognition model library based on model key-value pairs.

[0101] like Figure 6 The diagram shown is a flowchart of step S104 provided in an embodiment of the present invention. (Refer to...) Figure 6 Furthermore, as an optional implementation, the video scene description information is adjusted based on the scene recognition results, specifically including:

[0102] S1041. Determine whether the scene recognition result is consistent with the event type description;

[0103] S1042. When the scene recognition result is consistent with the event type description, the video content description is completed according to the scene recognition result to obtain the adjusted video scene description information.

[0104] S1043. When the scene recognition result is inconsistent with the event type description, the first video stream is subjected to scene recognition based on the scene recognition big model deployed on the cloud platform to obtain the adjusted video scene description information.

[0105] Specifically, the process involves determining whether the scene recognition result output by the high-precision video scene recognition model matches the event type description uploaded by the client. This allows for adjustments to the video scene description information. For example, if the uploaded video scene description is "[Video Content]: A fire has occurred; grass and trees along the road are burning; no casualties have been reported; [Event Type]: Fire; [Event Nature]: Negative," and the high-precision video scene recognition model outputs a multi-label result "A fire has occurred; the fire is large; it is a natural fire; it is a level two fire; there are no casualties," then the scene recognition result matches the event type description, and adjustments can be made based on the scene recognition result. Other tags are used to complete the video content description, resulting in adjusted video scene description information. For example, if the video scene description information uploaded by the client is "[Video Content]: A fire has occurred, grass and trees along the road are burning, and there are no casualties; [Event Type]: Fire; [Event Nature]: Negative", and the high-precision video scene recognition model outputs a single-tag result "No fire has occurred", it means that the scene recognition result is inconsistent with the event type description. In this case, it is necessary to re-recognize the first video stream based on the high-precision, general-purpose scene recognition model deployed on the cloud platform to obtain adjusted video scene description information.

[0106] like Figure 7 The diagram shown is another flowchart of step S104 provided in an embodiment of the present invention. (Refer to...) Figure 7 As an optional implementation, the decision to issue an alarm is made based on the adjusted video scene description information, specifically including:

[0107] S1044. Generate a second prompt statement based on the preset second prompt word template;

[0108] S1045. Input the adjusted video scene description information and the second prompt statement into the alarm decision model deployed on the cloud platform to obtain the alarm decision description.

[0109] S1046. Determine whether to issue an alarm based on the alarm decision description. If an alarm is issued, determine the alarm cause description based on the alarm decision description.

[0110] Specifically, the adjusted video scene description information is input into the comprehensive decision-making model for analysis. For example, a prompt is constructed: "[Video Content]: A fire has occurred; grass and trees along the road are burning. The fire is large, a natural fire, level two, and there are currently no casualties; [Event Type]: Fire; [Event Nature]: Negative." Based on the above content, a decision is made on whether to issue an alarm, answering in the following format: Alarm: XXX (only answer "Yes" or "No"), Alarm Reason: XXXXXX (Description of Alarm Reason). For example: Alarm: Yes, Alarm Reason: A fight has occurred, resulting in casualties." The decision-making model returns the result: "Alarm: Yes, Alarm Reason: The fire is spreading and will cause loss of life and property; relevant departments need to be notified immediately for handling." Then, a work order is generated, relevant department personnel accept the order and dispatch personnel to handle it, and the camera footage is used to verify whether the handling has been completed.

[0111] The method steps of the embodiments of the present invention have been described above. It can be understood that the embodiments of the present invention first utilize a lightweight multimodal large model deployed on the edge to perform preliminary identification of the acquired video images, obtaining video scene description information. Then, based on this video scene description information, a video scene recognition model corresponding to the subdivided domain is matched on the cloud platform. The high-precision video scene recognition model deployed on the cloud platform is then used to further identify the scene in the video images. This allows for adjustment of the video scene description information based on the scene recognition results, resulting in more accurate and detailed video scene description information. Furthermore, alarm decisions can be made based on the adjusted video scene description information, improving the efficiency and accuracy of scene monitoring.

[0112] like Figure 8 The diagram shown is a structural schematic of a scene monitoring device based on a multimodal large model provided in an embodiment of the present invention. (Refer to...) Figure 8 This invention provides a scene monitoring device based on a multimodal large model, comprising:

[0113] The video frame extraction module is used to acquire the first video stream of the target scene and determine the corresponding video frame image set based on the first video stream.

[0114] The edge recognition module is used to input the video frame image set into the multimodal large model deployed on the edge to obtain the corresponding video scene description information, and then upload the first video stream and video scene description information to the cloud platform;

[0115] The cloud platform recognition module is used to match the corresponding video scene recognition model on the cloud platform according to the video scene description information, and to perform scene recognition on the first video stream according to the video scene recognition model to obtain the scene recognition result;

[0116] The alarm decision module is used to adjust the video scene description information based on the scene recognition results, and to determine whether to issue an alarm based on the adjusted video scene description information.

[0117] The content of the above method embodiments is applicable to the device embodiments. The specific functions implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0118] This invention also provides an electronic device, comprising: a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for communication between the processor and the memory. When the program is executed by the processor, it implements the aforementioned scene monitoring method based on a multimodal large model. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0119] like Figure 9 The diagram shown is a hardware structure schematic of an electronic device provided in an embodiment of the present invention. (Refer to...) Figure 9 This invention provides an electronic device, comprising:

[0120] The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention.

[0121] The memory 902 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 902 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 to implement the scene monitoring method based on a multimodal large model according to the embodiments of this invention.

[0122] The input / output interface 903 is used to implement information input and output;

[0123] The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0124] Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904);

[0125] The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.

[0126] like Figure 10 The diagram shown is a structural schematic of the storage medium provided in an embodiment of the present invention. (Refer to...) Figure 10 The present invention also provides a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs 1001, which can be executed by one or more processors to implement the above-described scene monitoring method based on a multimodal large model.

[0127] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0128] This invention also discloses a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform... Figure 1 The method shown.

[0129] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order shown in the operation diagrams. For example, depending on the functions / operations involved, two consecutively shown blocks may actually be executed substantially simultaneously, or the aforementioned blocks may sometimes be executed in reverse order. Furthermore, the embodiments presented and described in the flowcharts of this invention are provided by way of example to provide a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented herein. Alternative embodiments are contemplated in which the order of various operations is changed and sub-operations described as part of a larger operation are executed independently.

[0130] Furthermore, although the invention has been described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the aforementioned functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the invention. Rather, given the properties, functions, and internal relationships of the various functional modules in the apparatus disclosed herein, the actual implementation of the module will be understood within the scope of conventional skill of an engineer. Therefore, those skilled in the art can implement the invention as set forth in the claims using ordinary techniques without excessive experimentation. It is also understood that the specific concepts disclosed are merely illustrative and not intended to limit the scope of the invention, which is determined by the full scope of the appended claims and their equivalents.

[0131] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0132] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0133] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the aforementioned program can be printed, because the aforementioned program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or, if necessary, processing in other suitable ways, and then stored in computer memory.

[0134] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0135] In the foregoing description of this specification, references to terms such as "one embodiment," "another embodiment," or "some embodiments" indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least one embodiment or example of the present invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0136] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

[0137] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.

Claims

1. A scene monitoring method based on a multimodal large model, characterized in that, Includes the following steps: Acquire the first video stream of the target scene, and determine the corresponding video frame image set based on the first video stream; The video frame image set is input into a multimodal large model deployed on the edge to obtain the corresponding video scene description information, and then the first video stream and the video scene description information are uploaded to the cloud platform. Based on the video scene description information, a corresponding video scene recognition model is matched on the cloud platform, and the first video stream is used for scene recognition based on the video scene recognition model to obtain the scene recognition result; The video scene description information is adjusted based on the scene recognition results, and an alarm is issued based on the adjusted video scene description information. The step of inputting the video frame image set into a multimodal large model deployed on the edge to obtain the corresponding video scene description information specifically includes: Generate the first prompt statement based on the preset first prompt word template; The video frame image set and the first prompt statement are input into the multimodal large model to obtain the video scene description information; The video scene description information includes video content description, event type description, and event nature description; The adjustment of the video scene description information based on the scene recognition result specifically includes: Determine whether the scene recognition result is consistent with the event type description; When the scene recognition result matches the event type description, the video content description is completed based on the scene recognition result to obtain the adjusted video scene description information; When the scene recognition result is inconsistent with the event type description, the first video stream is subjected to scene recognition based on the scene recognition big model deployed on the cloud platform to obtain the adjusted video scene description information.

2. The scene monitoring method based on a multimodal large model according to claim 1, characterized in that, The step of determining the corresponding video frame image set based on the first video stream specifically includes: The first video stream is divided into multiple video segments according to a preset time length; Video frames of each video segment are extracted according to a preset sampling frequency to obtain a set of video frame images corresponding to each video segment; The image set number of the video frame image set is determined based on the start time, end time, and corresponding camera code of the video segment.

3. The scene monitoring method based on a multimodal large model according to claim 1, characterized in that, The step of matching the corresponding video scene recognition model on the cloud platform based on the video scene description information specifically includes: Obtain the scene recognition model library stored on the cloud platform; The target model type is determined based on the event type description, and then matched in the scene recognition model library according to the target model type to obtain the corresponding video scene recognition model.

4. The scene monitoring method based on a multimodal large model according to claim 3, characterized in that, The scene recognition model library is constructed through the following steps: Obtain pre-trained video scene recognition models for multiple sub-domains, and determine the model type label of the corresponding video scene recognition model based on the sub-domain; The model type label is used as the key value, and the corresponding video scene recognition model is used as the value value to generate model key-value pairs; The scene recognition model library is constructed based on the model key-value pairs.

5. A scene monitoring method based on a multimodal large model according to any one of claims 1 to 4, characterized in that, The step of determining whether to issue an alarm based on the adjusted video scene description information specifically includes: Generate a second prompt statement based on a preset second prompt word template; The adjusted video scene description information and the second prompt statement are input into the alarm decision model deployed on the cloud platform to obtain the alarm decision description; Determine whether to issue an alarm based on the alarm decision description. If an alarm is issued, determine the alarm cause description based on the alarm decision description.

6. A scene monitoring device based on a multimodal large model, characterized in that, include: The video frame extraction module is used to acquire the first video stream of the target scene and determine the corresponding video frame image set based on the first video stream. The edge recognition module is used to input the video frame image set into the multimodal large model deployed on the edge to obtain the corresponding video scene description information, and then upload the first video stream and the video scene description information to the cloud platform; The cloud platform recognition module is used to match the corresponding video scene recognition model on the cloud platform according to the video scene description information, and to perform scene recognition on the first video stream according to the video scene recognition model to obtain the scene recognition result; The alarm decision module is used to adjust the video scene description information according to the scene recognition result, and to determine whether to issue an alarm based on the adjusted video scene description information. The step of inputting the video frame image set into a multimodal large model deployed on the edge to obtain the corresponding video scene description information specifically includes: Generate the first prompt statement based on the preset first prompt word template; The video frame image set and the first prompt statement are input into the multimodal large model to obtain the video scene description information; The video scene description information includes video content description, event type description, and event nature description; The adjustment of the video scene description information based on the scene recognition result specifically includes: Determine whether the scene recognition result is consistent with the event type description; When the scene recognition result matches the event type description, the video content description is completed based on the scene recognition result to obtain the adjusted video scene description information; When the scene recognition result is inconsistent with the event type description, the first video stream is subjected to scene recognition based on the scene recognition big model deployed on the cloud platform to obtain the adjusted video scene description information.

7. An electronic device, characterized in that, The electronic device includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for establishing communication between the processor and the memory. When the program is executed by the processor, it implements the steps of the scene monitoring method based on a multimodal large model as described in any one of claims 1 to 5.

8. A storage medium, said storage medium being a computer-readable storage medium for computer-readable storage, characterized in that, The storage medium stores one or more programs, which can be executed by one or more processors to implement the steps of the scene monitoring method based on a multimodal large model as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Event processing method, device and equipment, storage medium and program product

    CN116778370A

  • Video automatic analysis method and device based on large language model control and medium

    CN116935288A