Monitoring and early warning method and device based on multi-modal large model, equipment and medium
Through the monitoring and early warning method based on multimodal large model, the problem of traditional computer vision algorithms insufficient understanding of complex scenes is solved, and accurate prediction and early warning of security videos in the monitoring area is achieved, which improves security and security.
Patent Information
- Application Number
- CN202510050744.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-06
AI Technical Summary
Traditional computer vision algorithms lack the ability to understand complex scenarios and lack the understanding and reasoning of general knowledge, which leads to the inability to accurately predict the content displayed in the video, resulting in low security security in the monitoring area.
The monitoring and early warning method based on the multimodal large model is adopted to obtain the security video of the target area in real time, and after preprocessing, the video is input into the pre-designed security model to obtain the security prediction results, and alarm information is generated based on the prediction results, sent to the user terminal for display, and the associated alarm equipment is controlled for alarm operations.
The multimodal large model predicts security videos, which improves the security and security of the monitoring area, and can accurately identify the features in the target area, thereby making effective security prediction and early warning.
Smart Images

Figure CN119946226A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular to a monitoring and early warning method, apparatus, device, and medium based on a multimodal large model. Background Art
[0002] In the field of security monitoring, traditional computer vision algorithms are usually used to identify collected data.
[0003] However, when the above method is used for security monitoring, the following technical problems often occur:
[0004] Traditional computer vision algorithms are unable to adequately comprehend complex scenes and lack the understanding and reasoning of general knowledge, which results in an inability to make accurate security predictions based on video content, leading to low security in the monitored area.
[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the invention
[0006] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.
[0007] Some embodiments of the present disclosure propose a monitoring and early warning method, device, electronic device and computer-readable medium based on a multimodal large model to solve one or more of the technical problems mentioned in the above background technology section.
[0008] In a first aspect, some embodiments of the present disclosure provide a monitoring and early warning method based on a multimodal big model, the method comprising: acquiring security videos corresponding to a target area in real time; preprocessing the security videos to generate preprocessed security videos; inputting the preprocessed security videos into a pre-deployed security big model to obtain security prediction results; in response to the security prediction results satisfying preset alarm conditions, generating alarm information corresponding to the security prediction results; sending the alarm information to an associated user terminal for display, and controlling an associated alarm device to perform an alarm operation.
[0009] In a second aspect, some embodiments of the present disclosure provide a monitoring and early warning device based on a multimodal big model, the device comprising: an acquisition unit, configured to acquire security videos corresponding to a target area in real time; a preprocessing unit, configured to preprocess the security videos to generate preprocessed security videos; an input unit, configured to input the preprocessed security videos into a pre-deployed security big model to obtain security prediction results; a generation unit, configured to generate alarm information corresponding to the security prediction results in response to the security prediction results satisfying preset alarm conditions; a sending unit, configured to send the alarm information to an associated user terminal for display, and to control an associated alarm device to perform an alarm operation.
[0010] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the above-mentioned first aspect.
[0011] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the above-mentioned first aspect is implemented.
[0012] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: through the monitoring and early warning method based on the multimodal large model of some embodiments of the present disclosure, the security of the monitoring area is improved. Specifically, the reason for the low security of the monitoring area is that the traditional computer vision algorithm has insufficient understanding ability of complex scenes and lacks understanding and reasoning of general knowledge, which makes it impossible to make security predictions through the content displayed by the video, resulting in low security in the monitoring area. Based on this, the monitoring and early warning method based on the multimodal large model of some embodiments of the present disclosure first obtains the security video corresponding to the target area in real time. Thus, the video of the target area can be obtained. Secondly, the above security video is preprocessed to generate a preprocessed security video. Thus, the security video can be preprocessed to facilitate security prediction. Then, the above preprocessed security video is input into the pre-deployed security large model to obtain the security prediction result. Thus, the content displayed by the security video can be predicted by the multimodal large model. Finally, in response to the security prediction result meeting the preset alarm condition, an alarm message corresponding to the security prediction result is generated; the alarm message is sent to the associated user terminal for display, and the associated alarm device is controlled to perform an alarm operation. Thus, the target area is monitored for security warning through the prediction result, and because the target area is monitored for security through a multimodal large model, the features displayed in the target area can be identified, thereby performing security prediction and improving the security of the monitored area. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0014] Figure 1 is a schematic diagram of an application scenario of a monitoring and early warning method based on a multimodal large model in some embodiments of the present disclosure;
[0015] Figure 2 is a flowchart of some embodiments of the monitoring and early warning method based on a multimodal large model according to the present disclosure;
[0016] Figure 3 is a flowchart of some embodiments of the monitoring and early warning method based on a multimodal large model according to the present disclosure;
[0017] Figure 4 It is a structural schematic diagram of some embodiments of the monitoring and early warning device based on the multimodal large model according to the present disclosure;
[0018] Figure 5It is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0019] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0020] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.
[0021] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0022] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0023] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0024] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0025] Figure 1 It is a schematic diagram of an application scenario of a monitoring and early warning method based on a multimodal large model in some embodiments of the present disclosure.
[0026] exist Figure 1In the application scenario, first, the computing device 101 can obtain the security video 102 corresponding to the target area in real time. Secondly, the computing device 101 can pre-process the above security video to generate a pre-processed security video 103. Then, the computing device 101 can input the above pre-processed security video 103 into the pre-deployed security big model 104 to obtain the security prediction result 105. Afterwards, the computing device 101 can generate an alarm message 106 corresponding to the above security prediction result in response to the above security prediction result 105 meeting the preset alarm condition. Finally, the computing device 101 can send the above alarm message 106 to the associated user terminal for display, and control the associated alarm device to perform an alarm operation.
[0027] It should be noted that the computing device 101 can be hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or as a single server or a single terminal device. When the computing device is embodied as software, it can be installed in the hardware devices listed above. It can be implemented as multiple software or software modules for providing distributed services, or as a single software or software module. No specific limitation is made here.
[0028] It should be understood that Figure 1 The number of computing devices in the embodiment is only illustrative. Any number of computing devices may be provided according to implementation requirements.
[0029] Continue to refer Figure 2 , shows a process 200 of some embodiments of the monitoring and early warning method based on a multimodal large model according to the present disclosure. The monitoring and early warning method based on a multimodal large model includes the following steps:
[0030] Step 201, obtaining security video corresponding to the target area in real time.
[0031] In some embodiments, the execution subject of the monitoring and early warning method based on the multimodal large model (for example Figure 1 The computing device 101 shown or the pre-deployed cloud server) can obtain the security video corresponding to the target area in real time. In practice, the security video of the target area can be obtained in real time by deploying a shooting device that can shoot the above-mentioned target area. Among them, the above-mentioned target area can be a pre-set area. The above-mentioned shooting device can be a device with a photo taking and video recording function. For example, the above-mentioned shooting device can be a camera. Here, the security video corresponding to the target area can also be obtained from a database storing security videos.
[0032] Step 202: pre-process the security video to generate a pre-processed security video.
[0033] In some embodiments, the execution entity may pre-process the security video to generate a pre-processed security video.
[0034] In practice, the above security video can be preprocessed through the following steps to generate a preprocessed security video:
[0035] The first step is to decode the security video to generate a security video frame sequence.
[0036] In the second step, according to the preset extraction configuration, the security video frame sequence is subjected to video frame extraction processing, and the extracted video frames are combined into an extracted video frame sequence.
[0037] The third step is to adjust each extracted video frame in the extracted video frame sequence to generate an adjusted video frame sequence, wherein the adjustment process includes frame rate adjustment and resolution adjustment.
[0038] In the fourth step, data enhancement and denoising are performed on the adjusted video frame sequence to generate an enhanced and denoised adjusted video frame sequence as a pre-processed security video.
[0039] Step 203, input the pre-processed security video into the pre-deployed security model to obtain the security prediction result.
[0040] In some embodiments, the execution subject may input the pre-processed security video into a pre-deployed security model to obtain a security prediction result. The security model may be a multi-modal model. Here, the security prediction result may be obtained by the following steps:
[0041] In the first step, the Transformer model is used to block the video frames in the preprocessed security video to generate a group of block-based video frames.
[0042] In the second step, feature extraction is performed on each of the segmented video frames included in the segmented video frame group to generate an extracted feature group.
[0043] The third step is to perform feature compression on each extracted feature included in the extracted feature group based on the MustDrop method to generate a compressed feature group.
[0044] The fourth step is to input the compressed feature group into the large language model included in the above-mentioned security model to obtain the security prediction result.
[0045] In some optional implementations of some embodiments, the security big model may include a security vision encoder and a big language model.
[0046] Here, the above security model can be trained through the following steps:
[0047] The first step is to obtain at least one unlabeled security video to obtain an unlabeled security video group, wherein the unlabeled security video may be a pre-stored video used to train the security big model.
[0048] The second step is to obtain a preset annotation rule set from the target database. The preset annotation rules in the preset annotation rule set are rules for executing specific annotation words. It should be noted that the execution subject can also use an open source general large model or a security large model to annotate the video.
[0049] The third step is to perform labeling and fusion processing on each unlabeled security video in the unlabeled security video group based on the preset labeling rule set to generate a labeled security video group. In practice, each unlabeled security video in the unlabeled security video group can be labeled and fused by a weak supervision method such as snorkel.
[0050] The fourth step is to use the above-mentioned annotated security video group to train the initial large model.
[0051] In step 5, in response to the training result not satisfying the preset stop condition, the above training steps are executed again.
[0052] In the process of adopting technical solutions to solve the above technical problems, the following technical problems are often accompanied: the content of visual features is redundant, and there are many invalid features, which affects the effect of the large security model, and requires more computing resources and storage resources. It also takes a long time to conduct security identification and early warning of the target area, affecting the user experience.
[0053] In some other optional implementations of some embodiments, the execution subject may generate the security prediction result through the following steps:
[0054] The first step is to determine the attention score corresponding to each visual feature sequence in each generated visual feature sequence. The attention score may be an attention score of the visual feature sequence determined by a self-attention mechanism.
[0055] The second step is to sort the visual feature sequences according to the corresponding attention scores, and select a preset number of visual feature sequences as the first target visual feature sequence set. The sorting process may be in descending order according to the corresponding attention scores.
[0056] In the third step, each visual feature sequence removed from the first target visual feature sequence set is determined as a second target visual feature sequence set.
[0057] Step 4: for each second target visual feature sequence in the second target visual feature sequence set, perform the following determination steps:
[0058] The first determination step is to determine the correlation between the second target visual feature sequence and each first target visual feature sequence in the first target visual feature sequence set to obtain a correlation set. Here, the correlation between the second target visual feature sequence and each first target visual feature sequence in the first target visual feature sequence set can be determined by determining the correlation between corresponding rows and columns in the attention matrix.
[0059] The second determination step is to determine the second target visual feature sequence as a supplementary feature sequence in response to the relevance set satisfying a preset relevance condition, wherein the preset relevance condition may be that there is a relevance in the relevance set that is greater than or equal to a preset relevance threshold.
[0060] The third determination step is to delete the second target visual feature sequence in response to the relevance set not satisfying the preset relevance condition.
[0061] The fifth step is to select at least one first target visual feature sequence from the above-mentioned first target visual feature sequence set for each of the determined supplementary feature sequences according to the corresponding correlation set, and compress the selected first target visual feature sequence and the above-mentioned supplementary feature sequence to generate a compressed feature sequence as the first target visual feature sequence.
[0062] The sixth step is to input each first target visual feature sequence into the large language model included in the above-mentioned security large model to generate a security prediction result.
[0063] The above-mentioned first step to the sixth step, as an inventive point of an embodiment of the present disclosure, combined with the following step "step 205", solves the technical problem "the content of the visual features is redundant, there are many invalid features, which affects the effect of the large security model, and leads to the need to consume more computing resources and storage resources, and it also takes a long time to perform security identification and warning on the target area, affecting the user experience.". The reasons for the need to consume more computing resources and storage resources, and the need to consume a long time for security identification and warning on the target area are as follows: the content of the visual features is redundant, there are many invalid features, which affects the effect of the large security model, and leads to the need to consume more computing resources and storage resources, and it also takes a long time to perform security identification and warning on the target area, affecting the user experience. If the above factors are solved, the effect of reducing the waste of computing resources and storage resources, and reducing the time for security identification and warning on the target area can be achieved. In order to achieve this effect, the present disclosure first determines the attention score corresponding to each visual feature sequence in each generated visual feature sequence; according to the corresponding attention score, each visual feature sequence is sorted and processed, and a preset number of visual feature sequences are selected as the first target visual feature sequence set. Thus, important visual features can be selected. Second, each visual feature sequence removed from the first target visual feature sequence set is determined as a second target visual feature sequence set; for each second target visual feature sequence in the second target visual feature sequence set, the following determination step is performed: the correlation between the second target visual feature sequence and each first target visual feature sequence in the first target visual feature sequence set is determined to obtain a correlation set. Thus, the feature correlation between unimportant visual features and important visual features can be determined. Third, in response to the correlation set satisfying a preset correlation condition, the second target visual feature sequence is determined as a supplementary feature sequence; in response to the correlation set not satisfying the preset correlation condition, the second target visual feature sequence is deleted. Thus, visual features with higher correlation can be fused, and associated visual features are retained, thereby reducing the input visual features while avoiding the loss of associated features. Fourth, for each of the determined supplementary feature sequences, at least one first target visual feature sequence is selected from the above-mentioned first target visual feature sequence set according to the corresponding correlation set, and the selected first target visual feature sequence and the above-mentioned supplementary feature sequence are compressed to generate a compressed feature sequence as the first target visual feature sequence; each first target visual feature sequence is input into the above-mentioned multimodal large model to generate a security prediction result. In this way, the security prediction of the target area is completed, and by deleting unimportant visual features, the time for security identification and early warning of the target area is reduced, and the waste of computing resources and storage resources is reduced.
[0064] Step 204: in response to the security prediction result satisfying the preset alarm condition, generate alarm information corresponding to the security prediction result.
[0065] In some embodiments, the execution subject may generate alarm information corresponding to the security prediction result in response to the security prediction result satisfying a preset alarm condition, wherein the preset alarm condition may be a pre-set condition for alarming the alarm information.
[0066] Step 205: Send the alarm information to the associated user terminal for display, and control the associated alarm device to perform an alarm operation.
[0067] In some embodiments, the execution entity may send the alarm information to an associated user terminal for display, and control an associated alarm device to perform an alarm operation.
[0068] Optionally, after step 205, in response to receiving the feedback result sent by the user terminal, the security big model is optimized according to the feedback result.
[0069] In some embodiments, the execution entity may optimize the security big model in response to receiving the feedback result sent by the user terminal according to the feedback result.
[0070] Optionally, after step 205, based on the feedback result sent by the user terminal, the associated security device is controlled to perform a preset security operation.
[0071] In some embodiments, the above-mentioned execution entity can control the associated security equipment to perform preset security operations based on the feedback results sent by the above-mentioned user terminal.
[0072] As an example, in response to the target area being a school gate, the feedback result may indicate that there is a security risk, and the security device may close the school gate, and the preset security operation may be controlling the school gate to close.
[0073] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: through the monitoring and early warning method based on the multimodal large model of some embodiments of the present disclosure, the security of the monitoring area is improved. Specifically, the reason for the low security of the monitoring area is that the traditional computer vision algorithm has insufficient understanding ability of complex scenes and lacks understanding and reasoning of general knowledge, which makes it impossible to make security predictions through the content displayed by the video, resulting in low security in the monitoring area. Based on this, the monitoring and early warning method based on the multimodal large model of some embodiments of the present disclosure first obtains the security video corresponding to the target area in real time. Thus, the video of the target area can be obtained. Secondly, the above security video is preprocessed to generate a preprocessed security video. Thus, the security video can be preprocessed to facilitate security prediction. Then, the above preprocessed security video is input into the pre-deployed security large model to obtain the security prediction result. Thus, the content displayed by the security video can be predicted by the multimodal large model. Finally, in response to the security prediction result meeting the preset alarm condition, an alarm message corresponding to the security prediction result is generated; the alarm message is sent to the associated user terminal for display, and the associated alarm device is controlled to perform an alarm operation. Thus, the target area is monitored for security warning through the prediction result, and because the target area is monitored for security through a multimodal large model, the features displayed in the target area can be identified, thereby performing security prediction and improving the security of the monitored area.
[0074] Further references Figure 3 , which shows a process 300 of another embodiment of a monitoring and early warning method based on a multimodal large model. The process 300 of the monitoring and early warning method based on a multimodal large model includes the following steps:
[0075] Step 301, obtaining security video corresponding to the target area in real time.
[0076] Step 302: pre-process the security video to generate a pre-processed security video.
[0077] In some embodiments, the specific implementation of steps 301-302 and the technical effects thereof can be referred to in Figure 2 The steps 201-202 in the corresponding embodiment are not described in detail here.
[0078] The above security video frames are visually encoded based on the visual model (Vision Transformer) obtained by contrastive learning. Here, a contrastive learning method based on SigLIP (Sigmoid Loss for Language-Image Pre-training) and CLIP (Contrastive Language-Image Pre-training) can be used, including the following steps:
[0079] Step 303, executing the following processing steps for each security video frame in the pre-processed security video:
[0080] Step 3031, block processing is performed on the security video frame to generate a group of security video frames after block processing. Here, the security video frame can be block processed according to a preset block size to generate a group of security video frames after block processing.
[0081] Step 3032: Tiling the blocked security video frames included in the blocked security video frame group to generate a blocked security video frame sequence.
[0082] Step 3033: Perform visual feature extraction processing on each of the segmented security video frames in the segmented security video frame sequence to generate a visual feature sequence. In practice, the visual feature extraction processing can be performed on each of the segmented security video frames in the segmented security video frame sequence by using a Vision Transformer (ViT).
[0083] Step 3034, mapping the output space including the visual feature sequence to the input space of the large language model included in the above-mentioned security large model.
[0084] Step 304: compress each generated visual feature sequence to generate a compressed visual feature sequence set.
[0085] In some embodiments, the execution subject can extract features of the two paths through the Slow pathway and the Fast pathway to capture static features and dynamic features for fusion to obtain security prediction results. Among them, the Slow pathway runs at a low frame rate and focuses on the spatial semantic information in the video, that is, the shape, color, texture and other features of the object in each frame. The Slow pathway can use spatial pooling with a smaller kernel / stride as the backbone network, usually with a higher time dimension step size τ, that is, one frame is selected from the input video frame sequence for every τ frames for processing. The Fast pathway runs at a high frame rate and focuses on capturing fine temporal resolution motion information in the video, that is, the motion trajectory and speed change of the object between different frames. The Fast pathway can use spatial pooling with a larger kernel / stride as the backbone network, with a higher frame rate, that is, a smaller step size τ / α, where α>1.
[0086] Step 305: input the compressed visual feature sequence set into the large language model included in the large security model to obtain a security prediction result.
[0087] In some embodiments, the execution entity may input the compressed visual feature sequence set into the large language model included in the large security model to obtain a security prediction result.
[0088] Step 306: In response to the security prediction result satisfying the preset alarm condition, generate alarm information corresponding to the security prediction result.
[0089] Step 307: Send the alarm information to the associated user terminal for display, and control the associated alarm device to perform an alarm operation.
[0090] In some embodiments, the specific implementation of steps 306-307 and the technical effects thereof can be referred to in Figure 2 Steps 204-205 in the corresponding embodiment will not be repeated here.
[0091] from Figure 3 It can be seen that Figure 2 Compared with the description of some corresponding embodiments, Figure 3The process 300 of the monitoring and early warning method based on the multimodal large model in some corresponding embodiments further highlights the specific steps of inputting the above-mentioned pre-processed security video into the pre-deployed security large model to obtain the security prediction results. Therefore, the scheme described in these embodiments captures the static features and dynamic features of the security video through comparative learning for fusion, thereby obtaining the security prediction results. In turn, it can avoid capturing invalid features, improve the effect of the large model, reduce the time for security identification and early warning of the target area, and reduce the waste of computing resources and storage resources.
[0092] Further references Figure 4 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a monitoring and early warning device based on a multi-modal large model. These device embodiments are similar to Figure 1 Corresponding to the method embodiments shown, the monitoring and early warning device based on the multimodal large model can be specifically applied to various electronic devices.
[0093] like Figure 4 As shown, the monitoring and early warning device 400 based on the multimodal big model of some embodiments includes: an acquisition unit 401, a preprocessing unit 402, an input unit 403, a generation unit 404 and a sending unit 405. Among them, the acquisition unit 401 is configured to acquire the security video corresponding to the target area in real time; the preprocessing unit 402 is configured to preprocess the security video to generate a preprocessed security video; the input unit 403 is configured to input the preprocessed security video into the pre-deployed security big model to obtain a security prediction result; the generation unit 404 is configured to generate an alarm information corresponding to the security prediction result in response to the security prediction result meeting the preset alarm condition; the sending unit 405 is configured to send the alarm information to the associated user terminal for display, and control the associated alarm device to perform an alarm operation.
[0094] It is understandable that the units recorded in the monitoring and early warning device 400 based on the multi-modal large model are similar to the reference Figure 2 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the monitoring and early warning device 400 based on the multimodal large model and the units contained therein, and will not be described in detail here.
[0095] Reference below Figure 5, which shows a schematic diagram of the structure of an electronic device 500 suitable for implementing some embodiments of the present disclosure. The electronic devices in some embodiments of the present disclosure may include but are not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0096] like Figure 5 As shown, the electronic device 500 may include a processing device 501 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0097] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 5 The electronic device 500 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 5 Each block shown in the figure may represent one device, or may represent multiple devices as required.
[0098] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.
[0099] It should be noted that the computer-readable medium recorded in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0100] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0101] The computer-readable medium may be included in the electronic device; or it may exist independently without being installed in the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: acquires the security video corresponding to the target area in real time. Preprocesses the security video to generate a preprocessed security video. Inputs the preprocessed security video into the pre-deployed security model to obtain a security prediction result. In response to the security prediction result satisfying the preset alarm condition, generates an alarm message corresponding to the security prediction result. Sends the alarm message to the associated user terminal for display, and controls the associated alarm device to perform an alarm operation.
[0102] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0103] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0104] The units described in some embodiments of the present disclosure may be implemented by software or hardware. The units described may also be set in a processor, for example, it may be described as: a processor includes an acquisition unit, a preprocessing unit, an input unit, a generation unit, and a sending unit. The names of these units do not constitute a limitation on the units themselves in some cases. For example, the acquisition unit may also be described as a "unit for acquiring security video corresponding to the target area in real time."
[0105] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0106] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) and the technical solutions formed.
Claims
1. A monitoring and early warning method based on a multimodal large model, comprising: Obtain security video corresponding to the target area in real time; Preprocessing the security video to generate a preprocessed security video; Inputting the pre-processed security video into the pre-deployed security model to obtain a security prediction result; In response to the security prediction result satisfying a preset alarm condition, generating alarm information corresponding to the security prediction result; The alarm information is sent to an associated user terminal for display, and the associated alarm device is controlled to perform an alarm operation.
2. The method according to claim 1, wherein: The method further comprises: In response to receiving the feedback result sent by the user terminal, the security big model is optimized according to the feedback result.
3. The method according to claim 1, wherein: The method further comprises: Based on the feedback result sent by the user terminal, the associated security equipment is controlled to perform a preset security operation.
4. The method according to claim 1, wherein: The security big model includes a security vision encoder and a big language model.
5. The method according to claim 4, wherein: The pre-processed security video is input into the pre-deployed security model to obtain the security prediction result, including: For each security video frame in the pre-processed security video, the following processing steps are performed: The security video frame is processed by dividing into blocks to generate a group of divided security video frames; Tiling each of the blocked security video frames included in the blocked security video frame group to generate a blocked security video frame sequence; Performing visual feature extraction processing on each of the segmented security video frames in the segmented security video frame sequence to generate a visual feature sequence; Mapping an output space including the visual feature sequence to an input space of a large language model included in the large security model; Performing compression processing on each generated visual feature sequence to generate a compressed visual feature sequence set; The compressed visual feature sequence set is input into the large language model included in the large security model to obtain a security prediction result.
6. The method according to claim 1, wherein: The security model is trained through the following training steps: Acquire at least one unlabeled security video to obtain an unlabeled security video group; Acquire a preset tagging rule set from a target database, wherein a preset tagging rule in the preset tagging rule set is a rule for executing a specific tagging word; Based on the preset annotation rule set, each unlabeled security video in the unlabeled security video group is annotated and fused to generate an annotated security video group; Using the annotated security video group, training an initial large model; In response to the training result not satisfying the preset stop condition, the training step is performed again.
7. The method according to claim 1, wherein: The preprocessing of the security video to generate a preprocessed security video includes: Decoding the security video to generate a security video frame sequence; According to a preset extraction configuration, performing video frame extraction processing on the security video frame sequence, and combining the extracted video frames into an extracted video frame sequence; Performing adjustment processing on each extracted video frame in the extracted video frame sequence to generate an adjusted video frame sequence, wherein the adjustment processing includes frame rate adjustment and resolution adjustment; The adjusted video frame sequence is subjected to data enhancement and denoising processing to generate an enhanced and denoised adjusted video frame sequence as a pre-processed security video.
8. A monitoring and early warning device based on a multimodal large model, comprising: An acquisition unit is configured to acquire security video corresponding to a target area in real time; A preprocessing unit, configured to preprocess the security video to generate a preprocessed security video; An input unit is configured to input the pre-processed security video into a pre-deployed security model to obtain a security prediction result; A generating unit, configured to generate alarm information corresponding to the security prediction result in response to the security prediction result satisfying a preset alarm condition; The sending unit is configured to send the alarm information to the associated user terminal for display, and to control the associated alarm device to perform an alarm operation.
9. An electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Intelligent security and protection monitoring system and method based on large model
CN120526380A
Alarm information processing method, device and system, and alarm information auxiliary processing method, device and system
CN120640083A
Monitoring and early-warning method and apparatus based on multi-modal large model, security video description information generation method and apparatus, security video description information processing method and apparatus, and device and medium
WO2026149563A1