Small place anomaly detection method and device based on large model

Through the combination of multimodal large model GLM-4V and generative adversarial network, the accuracy and real-time problems of abnormal detection in small places are solved, efficient and accurate abnormal behavior recognition and alarm are achieved, and management costs are reduced.

CN120259742APending Publication Date: 2025-07-04INSPUR SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510317843.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art has problems in small places with low accuracy, poor real-time performance, high missed detection rate and insufficient adaptability, making it difficult to realize real-time monitoring and processing of complex scenarios.

Method used

The multimodal large model GLM-4V is adopted, combining data collection, preprocessing, model fine-tuning and real-time inference, and the model is optimized to improve detection accuracy and efficiency by generating adversarial network super-resolution reconstruction and reinforcement learning.

Benefits of technology

Accurate detection of abnormal behaviors in small places is achieved, the error detection rate is reduced, detection efficiency and real-time performance is improved, the dependence on manual inspections is reduced, and management costs are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259742A_ABST
    Figure CN120259742A_ABST
Patent Text Reader

Abstract

The invention discloses a small place anomaly detection method and device based on a large model, and relates to the field of artificial intelligence and computer vision. Comprising the following steps of: 1, collecting data; 2, preprocessing the data; 3, based on the preprocessed data, learning abnormal behaviors in life and industrial scenes by utilizing a prompt project guide model, making a virtual user propose a problem, and generating scene-related anomaly detection task description, and 4, using the generated virtual user problem and real scene data to perform fine adjustment on a GLM-4V large model, so as to obtain a GLM-4V anomaly detection task description. The method comprises the following steps of: 1, detecting an abnormal behavior scene of a small place by using a GLM-4V large model, 5, evaluating and optimizing the GLM-4V large model, and 6, integrating the GLM-4V large model in an application system, performing real-time reasoning and detecting abnormal behaviors in the small place by using the GLM-4V large model, and when a problem is detected, automatically triggering an alarm and informing operation and maintenance personnel. And 7, regularly collecting comparison data of reasoning information and actual information to form feedback data, designing a reinforcement learning process by using the feedback data, and continuously optimizing the GLM-4V large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses a method and device for abnormal detection in small places based on a large model, which relates to the fields of artificial intelligence and computer vision. Background Art

[0002] With the acceleration of the urbanization process and the development of social economy, public safety issues have received increasing attention. Among them, some small places, such as shops, restaurants, hotels, Internet cafes, beauty salons, entertainment halls, processing workshops, and warehouses, etc. Due to characteristics such as small space, dense population, loose management, and simple facilities, they are extremely prone to becoming high-incidence areas of safety accidents such as fires, stampedes, and equipment failures. The existing safety management methods mostly rely on manual inspections or simple monitoring devices. However, manual inspections are costly and inefficient, and are limited by the inspection frequency and capabilities of personnel, making it difficult to achieve real-time monitoring. Video surveillance methods are often ineffective in the face of complex and changeable scenarios and are difficult to detect and handle potential abnormal behaviors in real time. Therefore, for the complex safety hazards in small places, there are still problems such as insufficient adaptability and generalization ability, lack of comprehensive coverage of abnormal behaviors, high false detection and missed detection rates, and poor real-time performance and response speed. Summary of the Invention

[0003] In view of the problems of the existing technology, the present invention provides a method and device for abnormal detection in small places based on a large model, which combines the multi-modal large model GLM-4V to solve the problems of low accuracy in detecting abnormal situations, high false detection and missed detection rates, and poor adaptability in complex scenarios.

[0004] The specific solution proposed by the present invention is as follows:

[0005] The present invention provides a method for abnormal detection in small places based on a large model, including:

[0006] Step 1: Perform data collection: Real-time collect image and video data from multiple data sources, and at the same time obtain relevant sensor data, and label the collected data to form a data set.

[0007] Step 2: Perform data preprocessing: Based on the generative adversarial network (GAN), perform super-resolution reconstruction on the image data to obtain an image with higher resolution, clearer edge contours, and more detailed image details.

[0008] Perform time synchronization and spatial alignment processing on data of different modalities to ensure that multi-source data is consistent in the time series.

[0009] Step 3: Based on the preprocessed data, use prompt engineering to guide the model to learn abnormal behaviors in life and industrial scenarios, and virtual users ask questions to generate abnormal detection task descriptions related to the scenarios.

[0010] Step 4: Use the generated virtual user questions and real - scenario data to fine - tune the GLM - 4V large model, enabling the GLM - 4V large model to detect abnormal behavior scenarios in small venues.

[0011] Step 5: Conduct the evaluation and optimization of the GLM - 4V large model.

[0012] Step 6: Integrate the GLM - 4V large model into the application system and use the GLM - 4V large model for real - time inference to detect abnormal behavior in small venues. When a problem is detected, automatically trigger an alarm to notify the operation and maintenance personnel.

[0013] Step 7: Regularly collect the comparison data between the inference information and the actual information to form feedback data. Use the feedback data to design a reinforcement learning process to continuously optimize the GLM - 4V large model.

[0014] Furthermore, in step 2 of the method for detecting abnormal behavior in small venues based on a large model, when performing data pre - processing, it also includes: cleaning and screening the data to remove noise, redundant information, and irrelevant interference data. During the cleaning process, use a filtering algorithm to remove invalid data caused by environmental changes and light fluctuations, reducing interference with the inference of the GLM - 4V large model and improving the overall detection accuracy and efficiency.

[0015] Furthermore, in step 2 of the method for detecting abnormal behavior in small venues based on a large model, when performing super - resolution reconstruction of image data based on the generative adversarial network GAN, it includes:

[0016] Process the low - resolution LR image through the generative network to reconstruct the low - resolution LR image into a super - resolution SR image. Input the generated super - resolution SR image and the real high - resolution HR image into the discriminative network together. The discriminative network distinguishes the input SR image and HR image, evaluates the authenticity and quality of the generated image, generates a judgment result, and transmits the judgment result as a feedback signal back to the generative network respectively. The generative network adjusts the parameters according to the feedback signal to make the generated SR image closer to the HR image.

[0017] Furthermore, in step 4 of the method for detecting abnormal behavior in small venues based on a large model, when fine - tuning the GLM - 4V large model, it includes: adjusting the model parameters, optimizing feature extraction, and conducting scene - adaptability training. When conducting scene - adaptability training, focus on detecting: the damage degree of fire doors and fire curtains; the situation where evacuation passages and fire - fighting passages are blocked or occupied; illegal parking behavior; the situation where fire extinguishers are not placed or are damaged; whether the safety exit signs and emergency lighting are in good condition; whether fire - fighting facilities are deactivated, demolished, or blocked; illegal parking and flying - wire charging of electric bicycles.

[0018] Furthermore, in step 5 of the method for anomaly detection in small places based on large models, the GLM-4V large model is evaluated using a test set and a validation set. The test set and the validation set contain various anomaly behavior samples under different scenarios and lighting conditions. The evaluation metrics include precision, recall rate, and F1 value.

[0019] According to the computing power and actual application requirements of the target device to be integrated, pruning, quantization, and knowledge distillation are performed on the GLM-4V large model. When pruning, the model complexity is reduced by removing redundant network connections. When quantifying, the weights of the GLM-4V large model are converted into low-precision representations, thereby reducing the occupation of computing resources.

[0020] The present invention also provides an anomaly detection device for small places based on large models, including a data acquisition module, a data preprocessing module, an anomaly detection module, an alarm output module, an evaluation and optimization module, and a decision support module.

[0021] The data acquisition module collects data: real-time collects image and video data from multiple data sources, and at the same time obtains relevant sensor data, and labels the collected data to form a data set.

[0022] The data preprocessing module performs data preprocessing: based on the generative adversarial network (GAN), super-resolution reconstruction is performed on the image data to obtain an image with higher resolution, clearer edge contours, and more detailed image details.

[0023] Perform time synchronization and spatial alignment processing on data of different modalities to ensure that multi-source data is consistent in the time series.

[0024] Based on the preprocessed data, the anomaly detection module uses prompt engineering to guide the model to learn anomaly behaviors in life and industrial scenarios. The virtual user asks questions to generate an anomaly detection task description related to the scenario.

[0025] The anomaly detection module uses the generated virtual user questions and real-scenario data to fine-tune the GLM-4V large model, enabling the GLM-4V large model to detect anomaly behavior scenarios in small places.

[0026] The evaluation and optimization module conducts the evaluation and optimization of the GLM-4V large model.

[0027] The anomaly detection module integrates the GLM-4V large model into the application system and uses the GLM-4V large model for real-time inference to detect anomaly behaviors in small places. When a problem is detected, the alarm output module automatically triggers an alarm to notify the operation and maintenance personnel.

[0028] The decision support module regularly collects comparative data between inference information and actual information to form feedback data. It uses the feedback data to design a reinforcement learning process and continuously optimize the GLM-4V large model.

[0029] Furthermore, the data preprocessing module of the small-place anomaly detection device based on a large model performs data preprocessing, and also includes: cleaning and screening the data to remove noise, redundant information and irrelevant interference data. During the cleaning process, the invalid data caused by environmental changes and light fluctuations is removed through a filtering algorithm, thereby reducing interference with the GLM-4V large model reasoning and improving the overall detection accuracy and efficiency.

[0030] Furthermore, the data preprocessing module of the small-place anomaly detection device based on a large model performs super-resolution reconstruction on the image data based on the adversarial network GAN, including: processing the low-resolution LR image through the generative network, reconstructing the low-resolution LR image into a super-resolution SR image, inputting the generated super-resolution SR image and the real high-resolution HR image into the discriminant network together, distinguishing the input SR image and HR image through the discriminant network, evaluating the authenticity and quality of the generated image, generating a judgment result, and transmitting the judgment result back to the generative network as feedback signals, and adjusting parameters according to the feedback signal through the generative network to make the generated SR image closer to the HR image.

[0031] Furthermore, the anomaly detection module of the small-venue anomaly detection device based on a large model fine-tunes the GLM-4V large model, including: adjusting model parameters, optimizing feature extraction and performing scene adaptability training, wherein during the scene adaptability training, the focus is on detecting: the degree of damage to fire doors and fire shutters; the blockage or occupation of evacuation passages and fire passages; illegal parking behavior; the absence or damage of fire extinguishers; whether the safety exit signs and emergency lighting are intact; whether the fire-fighting facilities are disabled, dismantled or blocked; illegal parking and flying wire charging of electric bicycles.

[0032] Furthermore, the evaluation and optimization module of the large-model-based small-site anomaly detection device uses a test set and a validation set to evaluate the GLM-4V large model. The test set and the validation set contain a variety of abnormal behavior samples under different scenes and lighting conditions. The evaluation indicators include precision, recall rate, and F1 value.

[0033] According to the computing power of the target device to be integrated and the actual application requirements, the GLM-4V large model is pruned, quantized and knowledge distilled. During pruning, the model complexity is reduced by removing redundant network connections. During quantization, the weights of the GLM-4V large model are converted into low-precision representations, thereby reducing the occupancy of computing resources.

[0034] The advantages of the present invention are as follows:

[0035] Through the multi-modal large model GLM-4V, precise detection of abnormal behaviors in small places is achieved, effectively solving the problems of low detection accuracy, poor real-time performance, and insufficient multi-modal data fusion ability existing in the prior art. Data from multiple devices are collected and processed, and the super-resolution reconstruction algorithm of the generative adversarial network is used to optimize images, significantly improving the parsing ability of low-quality images and enabling the detection system to maintain efficient and accurate recognition in complex environments.

[0036] In addition, through the deep fusion of multi-modal data and the real-time inference mechanism, the detection efficiency of common safety hazards in small places is significantly improved. Especially in the identification of high-risk behaviors such as blocked fire exits, damaged fire extinguishers, and illegal charging of electric bicycles, it shows excellent accuracy and rapid response ability. By optimizing the model structure and introducing edge computing technology, the detection system of the present invention can achieve rapid identification and alarm of abnormal behaviors with low computational resource consumption, effectively reducing the dependence on manual inspections, lowering management costs, and improving the overall safety management level of small places. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a schematic diagram of the application process of the device of the present invention.

[0038] Figure 2 It is a schematic diagram of the data collection and preprocessing process.

[0039] Figure 3 It is a schematic diagram of the super-resolution reconstruction process of the generative adversarial network (GAN).

[0040] Figure 4 It is a schematic diagram of the fine-tuning process of the large model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] Large language model: A large language model refers to a natural language processing model with a large number of parameters and powerful learning ability. Such models are usually based on deep learning technology and adopt a deep neural network structure, which contains hundreds of millions to hundreds of billions of parameters. This enables them to learn and understand complex patterns, grammar rules, and semantic relationships in natural language.

[0042] GLM-4V: As an open-source multi-modal large model of Zhipu AI, GLM-4V is developed based on the GLM series of models, has a huge parameter scale, and has the ability of Chinese-English bilingual multi-round dialogue at a high resolution of 1120*1120. It shows excellent performance in multi-modal evaluations such as comprehensive Chinese-English ability, perceptual reasoning, character recognition, and chart understanding.

[0043] Prompt Engineering: In the field of natural language processing and machine learning, prompt engineering refers to the input text provided by a user or system to a model, which is used to trigger the model to generate a corresponding output. In a dialogue system or a generative model, the prompt text is usually a question, request, or task description posed by the user.

[0044] Super-resolution Reconstruction: Image super-resolution reconstruction (SR) is an image processing technique that uses a computer to process a low-resolution image (LR) or a sequence of images and restore a high-resolution image (HR).

[0045] Generative Adversarial Network (GAN), whose full name is Generative adversarial networks, is a generative model proposed by Goodfellow et al. in 2014. Through two neural networks, namely the generator and the discriminator, they compete with each other to learn the data distribution. The generator is responsible for learning to generate data similar to the real data from random noise, while the discriminator tries to distinguish between the generated data and the real data.

[0046] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the illustrated embodiments are not intended to limit the present invention.

[0047] Embodiment 1

[0048] The present invention provides a method for abnormal detection in small places based on a large model, including:

[0049] Step 1: Data collection: Real-time image and video data are collected from multiple data sources, and at the same time, relevant sensor data are obtained, and the collected data are labeled to form a data set.

[0050] Among them, the data are obtained by real-time monitoring of each key area in the small place through devices such as cameras, inspection robots, and robotic dogs, and multi-modal data including images, videos, texts, etc. are collected. These data include various instances of normal and abnormal behaviors in the small place, such as the closed state of fire doors, the smoothness of evacuation channels, and the integrity of fire-fighting facilities, providing basic data support for subsequent detection.

[0051] Step 2: Data preprocessing, including: cleaning and screening the data, removing noise, redundant information, and irrelevant interference data. During the cleaning process, invalid data caused by environmental changes and light fluctuations are removed through filtering algorithms to reduce interference with the inference of the GLM-4V large model and improve the overall detection accuracy and efficiency.

[0052] Based on the generative adversarial network (GAN), super-resolution reconstruction is performed on image data to obtain images with higher resolution, clearer edge contours, and more detailed image details. Among them, the super-resolution reconstruction technology based on the generative adversarial network (GAN) optimizes low-quality or blurred images. Through the magnification and enhancement of the pixel details of the image, this process can convert blurred surveillance footage into high-resolution images, providing more accurate data input for subsequent anomaly detection. The super-resolution reconstruction of image data based on the adversarial network GAN may include:

[0053] The low-resolution LR image is processed by the generative network to reconstruct and generate a super-resolution SR image. The generated super-resolution SR image and the real high-resolution HR image are input into the discriminative network together. The discriminative network distinguishes the input SR image and HR image, evaluates the authenticity and quality of the generated image, generates a judgment result, and transmits the judgment result back to the generative network as a feedback signal respectively. The generative network adjusts the parameters according to the feedback signal to make the generated SR image closer to the HR image.

[0054] Time synchronization and spatial alignment processing are performed on data of different modalities to ensure the consistency of multi-source data in the time series. Among them, multi-modal data from different devices, such as videos, images, and texts, are synchronized and aligned to ensure the consistency of these data in time and space. The multi-modal alignment processing includes time series alignment and feature alignment, enabling data from different sources to be analyzed on the same time axis. Through the alignment operation, the system can more accurately capture the occurrence time of abnormal behaviors and their associated information, ensuring the synchronization and consistency of various types of data.

[0055] Step 3: Based on the preprocessed data, use prompt engineering to guide the model to learn abnormal behaviors in life and industrial scenarios. The virtual user asks questions to generate anomaly detection task descriptions related to the scenarios. Such as "Is the fire door closed?" "Is the evacuation route occupied?" etc. These virtual questions help improve the model's ability to identify abnormal behaviors. The generated questions are used to expand the model training set and enhance the model's understanding and adaptability to diverse abnormal scenarios.

[0056] Step 4: Use the generated virtual user questions and real scenario data to fine-tune the GLM-4V large model so that the GLM-4V large model can detect abnormal behavior scenarios in small places.

[0057] Among them, fine-tuning the GLM-4V large model includes: adjusting model parameters, optimizing feature extraction, and performing scene adaptation training. When performing scene adaptation training, key detections include: the damage degree of fire doors and fire shutters; the situation where evacuation passages and fire corridors are blocked or occupied; illegal parking behaviors; the situation where fire extinguishers are not placed or are damaged; whether the safety exit signs and emergency lighting are in good condition; whether fire-fighting facilities are deactivated, demolished, or blocked; and illegal parking and flying wire charging of electric bicycles.

[0058] Step 5: Evaluate and optimize the GLM-4V large model. Among them, use the test set and validation set to evaluate the GLM-4V large model. The test set and validation set contain various abnormal behavior samples under different scenarios and lighting conditions. The evaluation metrics include precision, recall rate, and F1 value.

[0059] According to the computing power and actual application requirements of the target device to be integrated, perform pruning, quantization, and knowledge distillation on the GLM-4V large model. When pruning, reduce the model complexity by removing redundant network connections. When quantifying, convert the GLM-4V large model weights into low-precision representations, thereby reducing the occupation of computing resources.

[0060] Step 6: Integrate the GLM-4V large model into the application system and use the GLM-4V large model for real-time inference to detect abnormal behaviors in small places. When a problem is detected, automatically trigger an alarm to notify the operation and maintenance personnel. Among them, during the system operation, receive multi-modal data in real time and perform abnormal behavior detection. Design an automated alarm mechanism. When the model detects an abnormal behavior, immediately trigger the alarm system, send alarm information to the monitoring center, and record relevant logs, retaining the detection results and relevant video clips as the basis for post-event analysis. Support a regular feedback mechanism, and perform online adjustment and optimization on the GLM-4V large model through reinforcement learning technology to continuously improve the detection accuracy and robustness.

[0061] Step 7: Regularly collect comparison data between inference information and actual information to form feedback data. Use the feedback data to design a reinforcement learning process to continuously optimize the GLM-4V large model. Specifically, use reinforcement learning technology to dynamically adjust the model. Use the detection results in the actual scenario as reward or punishment signals to adjust the model parameters and strategies, enabling it to more accurately identify diverse abnormal behaviors. This continuous learning ability ensures that the system can adapt to environmental changes and new behavior patterns. At the same time, provide intelligent decision-making support functions, such as recommending the optimal emergency treatment plan when the evacuation passage is blocked, or reminding relevant personnel to perform timely maintenance when fire-fighting equipment is damaged.

[0062] Embodiment 2

[0063] The present invention also provides a small - venue anomaly detection device based on a large model, which includes a data acquisition module, a data pre - processing module, an anomaly detection module, an alarm output module, an evaluation and optimization module, and a decision - making support module.

[0064] The data acquisition module collects data: it collects image and video data in real - time from multiple data sources, and at the same time obtains relevant sensor data, and labels the collected data to form a data set.

[0065] The data pre - processing module performs data pre - processing: based on the generative adversarial network (GAN), it performs super - resolution reconstruction on the image data to obtain images with higher resolution, clearer edge contours and image details.

[0066] It performs time synchronization and spatial alignment processing on data of different modalities to ensure that multi - source data is consistent in the time series.

[0067] Based on the pre - processed data, the anomaly detection module uses prompt engineering to guide the model to learn abnormal behaviors in life and industrial scenarios. The virtual user asks questions to generate anomaly detection task descriptions related to the scenario.

[0068] The anomaly detection module uses the generated virtual user questions and real - scenario data to fine - tune the GLM - 4V large model, so that the GLM - 4V large model can detect abnormal behaviors in small - venue scenarios.

[0069] The evaluation and optimization module evaluates and optimizes the GLM - 4V large model.

[0070] The anomaly detection module integrates the GLM - 4V large model into the application system and uses the GLM - 4V large model for real - time inference to detect abnormal behaviors in small venues. When a problem is detected, the alarm output module automatically triggers an alarm to notify the operation and maintenance personnel.

[0071] The decision - making support module regularly collects comparison data between inference information and actual information to form feedback data, and uses the feedback data to design a reinforcement learning process to continuously optimize the GLM - 4V large model.

[0072] Regarding the information interaction and execution process among the above - mentioned modules in the device, since they are based on the same concept as the method embodiment of the present invention, the specific content can be referred to the description in the method embodiment of the present invention, and will not be elaborated here.

[0073] Similarly, the device of the present invention realizes the accurate detection of abnormal behaviors in small places through the multi-modal large model GLM-4V, effectively solving the problems of low detection accuracy, poor real-time performance, and insufficient multi-modal data fusion ability existing in the prior art. It collects and processes data from multiple devices, and uses the super-resolution reconstruction algorithm of the generative adversarial network to optimize images, significantly improving the parsing ability of low-quality images and enabling the detection system to maintain efficient and accurate recognition in complex environments.

[0074] In addition, through the deep fusion of multi-modal data and the real-time inference mechanism, the present invention significantly improves the detection efficiency of common safety hazards in small places, especially showing excellent accuracy and rapid response ability in the identification of high-risk behaviors such as the occupation of fire corridors, damage to fire extinguishers, and illegal charging of electric bicycles. By optimizing the model structure and introducing edge computing technology, the detection system of the present invention can achieve rapid identification and alarm of abnormal behaviors with low computational resource consumption, effectively reducing the dependence on manual inspections, lowering management costs, and improving the overall safety management level of small places.

[0075] It should be noted that not all steps and modules in the above-mentioned processes and device structures are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted according to needs. The system structure described in the above-mentioned embodiments can be a physical structure or a logical structure, that is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities respectively, or some components in multiple independent devices can be jointly implemented.

[0076] The above-described embodiments are only preferred embodiments given to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are within the protection scope of the present invention. The protection scope of the present invention is subject to the claims.

Claims

1. A method for detecting anomalies in small venues based on large models, characterized in that Including: Step 1: Conduct data collection: Real-time collect image and video data from multiple data sources, and at the same time obtain relevant sensor data. Label the collected data to form a data set. Step 2: Conduct data preprocessing: Based on the generative adversarial network (GAN), perform super-resolution reconstruction on the image data to obtain images with higher resolution, clearer edge contours, and more detailed image details. Perform time synchronization and spatial alignment processing on data of different modalities to ensure that multi-source data is consistent in the time series. Step 3: Based on the preprocessed data, use prompt engineering to guide the model to learn abnormal behaviors in life and industrial scenarios. The virtual user asks questions to generate abnormal detection task descriptions related to the scenarios. Step 4: Use the generated virtual user questions and real-scenario data to fine-tune the large GLM-4V model so that the GLM-4V large model can detect abnormal behavior scenarios in small places. Step 5: Conduct evaluation and optimization of the GLM-4V large model. Step 6: Integrate the GLM-4V large model into the application system and use the GLM-4V large model for real-time inference to detect abnormal behaviors in small places. When a problem is detected, automatically trigger an alarm to notify the operation and maintenance personnel. Step 7: Regularly collect comparison data between inference information and actual information to form feedback data. Use the feedback data to design a reinforcement learning process to continuously optimize the GLM-4V large model.

2. The method for abnormal detection of small venues based on large models according to claim 1 is characterized in that When conducting data preprocessing in Step 2, it also includes: cleaning and screening the data to remove noise, redundant information, and irrelevant interference data. During the cleaning process, use a filtering algorithm to remove invalid data caused by environmental changes and light fluctuations, reduce interference with the inference of the GLM-4V large model, and improve the overall detection accuracy and efficiency.

3. The method for abnormal detection of small places based on a large model according to claim 1, characterized in that When performing super-resolution reconstruction on the image data based on the generative adversarial network (GAN) in Step 2, it includes: Process the low-resolution (LR) image through the generator network to reconstruct the low-resolution LR image into a super-resolution (SR) image. Input the generated super-resolution SR image and the real high-resolution (HR) image into the discriminator network. The discriminator network distinguishes the input SR image and HR image, evaluates the authenticity and quality of the generated image, generates a judgment result, and transmits the judgment result as a feedback signal back to the generator network respectively. The generator network adjusts the parameters according to the feedback signal to make the generated SR image closer to the HR image.

4. A method for abnormal detection of small places based on a large model according to claim 1, The feature is that when fine-tuning the GLM-4V large model in Step 4, it includes: adjusting the model parameters, optimizing feature extraction, and conducting scene adaptability training. When conducting scene adaptability training, focus on detecting: the damage degree of fire doors and fire curtains; the situation where evacuation channels and fire channels are blocked or occupied; illegal parking behaviors; the situation where fire extinguishers are not placed or are damaged; whether the safety exit indicator signs and emergency lighting are in good condition; whether fire-fighting facilities are deactivated, demolished, or blocked; illegal parking and flying wire charging of electric bicycles.

5. The method for abnormal detection of small places based on large models according to claim 1, characterized in that In step 5, the GLM-4V large model is evaluated using the test set and the validation set. The test set and the validation set contain various abnormal behavior samples under different scenarios and lighting conditions. The evaluation metrics include precision, recall, and F1 score. According to the computing power of the target device to be integrated and the actual application requirements, pruning, quantization, and knowledge distillation are performed on the GLM-4V large model. During pruning, the model complexity is reduced by removing redundant network connections. During quantization, the weights of the GLM-4V large model are transformed into low-precision representations, thereby reducing the consumption of computing resources.

6. An abnormal detection device for small places based on a large model, characterized in that It includes a data acquisition module, a data preprocessing module, an anomaly detection module, an alarm output module, an evaluation and optimization module, and a decision support module. The data acquisition module collects data: real-time collects image and video data from multiple data sources, and at the same time obtains relevant sensor data, and labels the collected data to form a data set. The data preprocessing module performs data preprocessing: based on the generative adversarial network (GAN), super-resolution reconstruction is performed on the image data to obtain images with higher resolution, clearer edge contours, and more detailed image details. Perform time synchronization and spatial alignment processing on data of different modalities to ensure that multi-source data is consistent in the time series. Based on the preprocessed data, the anomaly detection module uses prompt engineering to guide the model to learn abnormal behaviors in life and industrial scenarios. The virtual user asks questions and generates anomaly detection task descriptions related to the scenario. The anomaly detection module uses the generated virtual user questions and real-scenario data to fine-tune the GLM-4V large model, enabling the GLM-4V large model to detect abnormal behavior scenarios in small places. The evaluation and optimization module evaluates and optimizes the GLM-4V large model. The anomaly detection module integrates the GLM-4V large model into the application system and uses the GLM-4V large model for real-time inference to detect abnormal behaviors in small places. When a problem is detected, the alarm output module automatically triggers an alarm to notify the operation and maintenance personnel. The decision support module regularly collects comparison data between the inference information and the actual information to form feedback data. Using the feedback data, a reinforcement learning process is designed to continuously optimize the GLM-4V large model.

7. The small venue anomaly detection device based on a large model according to claim 6, characterized in that The data preprocessing module also includes: cleaning and screening the data, removing noise, redundant information, and irrelevant interference data. During the cleaning process, invalid data caused by environmental changes and light fluctuations is removed through filtering algorithms to reduce interference with the inference of the GLM-4V large model and improve the overall detection accuracy and efficiency.

8. The small venue anomaly detection device based on a large model according to claim 6, characterized in that The data preprocessing module performs super-resolution reconstruction on image data based on the generative adversarial network GAN, including: processing the low-resolution LR image through the generative network, reconstructing the low-resolution LR image into a super-resolution SR image, inputting the generated super-resolution SR image and the real high-resolution HR image into the discriminative network together, distinguishing the input SR image and HR image through the discriminative network, evaluating the authenticity and quality of the generated image, generating a judgment result, and transmitting the judgment result as a feedback signal back to the generative network respectively. The generative network adjusts parameters according to the feedback signal to make the generated SR image closer to the HR image.

9. The abnormal detection device for small places based on a large model according to claim 6, wherein the abnormality The detection module fine-tunes the GLM-4V large model, including: adjusting model parameters, optimizing feature extraction, and performing scene adaptation training. When performing scene adaptation training, the key detections include: the damage degree of fire doors and fire curtains; the situation where evacuation passages and fire corridors are blocked or occupied; illegal parking behaviors; the situation where fire extinguishers are not placed or are damaged in configuration; whether the safety exit indicator signs and emergency lighting are in good condition; whether fire-fighting facilities are deactivated, demolished, or blocked; illegal parking and flying wire charging of electric bicycles.

10. The small venue anomaly detection device based on a large model according to claim 6, characterized in that The evaluation and optimization module evaluates the GLM-4V large model using the test set and the validation set. The test set and the validation set contain various abnormal behavior samples under different scenarios and lighting conditions. The evaluation metrics include precision, recall rate, and F1 value. According to the computing power and actual application requirements of the target device to be integrated, pruning, quantization, and knowledge distillation processing are performed on the GLM-4V large model. When pruning, the model complexity is reduced by removing redundant network connections. When quantizing, the weights of the GLM-4V large model are transformed into low-precision representations, thereby reducing the occupation of computing resources.