Method and system for detecting an intruding object based on machine vision

CN122821484APending Publication Date: 2026-09-25ZHONGKE RUNCHENG BEIJING INTERNET OF THINGS SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611275503.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-21
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,实践应用表明,在室外复杂环境下,尤其是低照度、强光干扰、雨雪、大雾、沙尘等恶劣天气条件下,现有CNN模型的识别准确率往往出现显著下降,难以保持稳定的检测性能,严重制约了其在真实安防场景中的可靠性与实用性

Benefits of technology

[0015]本申请实施例中,第一监控模型输出的置信度用于表征对推理结果的信心,因此,当第一监控模型输出的置信度较低时,由第二监控模型进行精确复核以提升识别准确性,降低漏检与误报的可能性;另一方面,对于虽然置信度高,但涉及入侵对象的高风险事件,则强制触发复核,从而在保障安全性的前提下,最大限度降低大模型的调用次数与计算成本,实现了检测精度、响应速度与资源消耗之间的动态平衡。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821484A_ABST
    Figure CN122821484A_ABST
Patent Text Reader

Abstract

The application provides an intrusion object detection method and system based on machine vision. It relates to the field of security technology. The method comprises: acquiring monitoring data collected in a target monitoring area; analyzing the monitoring data using a first monitoring model to obtain a first analysis result; in the case that the first analysis result meets review conditions, generating input information according to the monitoring data and inference results; sending the input information to a backend processor; wherein the backend processor is deployed with a second monitoring model, the second monitoring model is used for reanalyzing the input information to obtain a second analysis result; the parameter quantity of the first monitoring model is less than that of the second monitoring model, and the performance of the first monitoring model is lower than that of the second monitoring model. The application can improve the accuracy of detection by combining the first monitoring model and the second monitoring model for intrusion object detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of security technology, and more specifically, to a machine vision-based method and system for detecting intrusion objects. Background Technology

[0002] In security scenarios, to achieve the classification and identification of intrusion targets, current mainstream technologies typically employ Convolutional Neural Network (CNN) models (such as ResNet and YOLO) for training. In actual deployments, the system directly inputs the image to be detected into the pre-trained CNN model, obtaining classification results through forward inference, including the target's location in the image, category label, and confidence score. However, practical applications show that in complex outdoor environments, especially under adverse weather conditions such as low light, strong light interference, rain, snow, fog, and sandstorms, the recognition accuracy of existing CNN models often drops significantly, making it difficult to maintain stable detection performance and severely limiting their reliability and practicality in real-world security scenarios. Summary of the Invention

[0003] The purpose of this application is to provide a machine vision-based intrusion detection method and system to improve the accuracy of intrusion detection in monitored areas.

[0004] In a first aspect, embodiments of this application provide a machine vision-based intrusion object detection method applied to a front-end processor, the method comprising: Acquire monitoring data collected from the target monitoring area; The monitoring data is analyzed using the first monitoring model to obtain the first analysis result; the first analysis result includes the inference result; the inference result is used to characterize whether there is an intrusion object in the target monitoring area, and related information about the intrusion object; If the review conditions are met based on the first analysis results, input information is generated according to the monitoring data and reasoning results. Input information is sent to the backend processor; wherein, a second monitoring model is deployed on the backend processor, and the second monitoring model is used to analyze the input information again to obtain a second analysis result; wherein, the second analysis result is used to characterize whether there is an intrusion object in the target monitoring area, and related information of the intrusion object; the number of parameters of the first monitoring model is less than the number of parameters of the second monitoring model, and the performance of the first monitoring model is lower than that of the second monitoring model.

[0005] In this embodiment, a first monitoring model is used to analyze monitoring data to obtain a first analysis result. If the first analysis result meets the verification conditions, a second monitoring model is activated for verification. Since the second monitoring model has higher performance and detection accuracy than the first monitoring model, the detection accuracy can be improved by combining the first and second monitoring models to detect intrusion objects.

[0006] In one possible implementation of the first aspect, the monitoring data includes video data; the first monitoring model includes a classification and recognition model and an action recognition model; and the monitoring data is analyzed using the first monitoring model, including: Decode the video data to obtain multiple frames of images; The classification and recognition model is used to classify and recognize multiple frames of images to obtain classification results. The classification results include whether the target object is contained. If the target object is contained, the classification results also include the type of the target object. If the target object is a person, then the action recognition model is used to perform action recognition on the image to obtain the position of the person's limb joints in the image, and the action of the person in the image is determined based on the position of the limb joints.

[0007] In this embodiment of the application, since the intrusion target is often a person, a classification and recognition model can be used to classify and recognize multiple frames of images. If the target object in the image is determined to be a person, an action recognition model can be used to identify its actions. Based on the actions, it can be preliminarily determined whether there is an intrusion behavior.

[0008] In one possible implementation of the first aspect, input information is generated based on monitoring data and inference results, including: The reasoning results are described in a structured manner to obtain structured information; The structured information and the corresponding image are used as input information.

[0009] The embodiments of this application describe the reasoning results in a structured manner, thereby meeting the format requirements of the second monitoring model for input information.

[0010] In one possible implementation of the first aspect, the monitoring data further includes vibration data, Doppler data, and laser point cloud data; the method also includes: Vibration data, Doppler data, and laser point cloud data are projected onto the visual-language joint latent space of the second monitoring model through the projection module to obtain pseudo-visual feature maps. The structured information and the corresponding image are used as input information, including: The structured information, pseudo-visual feature maps, and corresponding images are used as input information.

[0011] This application embodiment uses multimodal data for deep semantic fusion, enabling the second monitoring model to perform joint reasoning using complementary information provided by vibration, Doppler and laser point cloud data, which greatly improves the target recognition accuracy and environmental robustness of the system in complex environments.

[0012] In one possible implementation of the first aspect, input information is generated based on monitoring data and inference results, including: The reasoning results are described in a structured manner to obtain structured information; Extract video clips of preset duration before and after an image containing an intrusion target; Structured information and video clips are used as input information.

[0013] This application embodiment describes the reasoning results in a structured way and uses video clips of a preset duration before and after the image containing the intrusion target as input information. This method can expand the detection results at a single moment into a continuous event context that includes the time dimension. This allows the second monitoring model to not only grasp the static information such as the target's category and location, but also capture its behavioral trajectories and spatiotemporal evolution patterns such as entry, movement, stay, and departure. This effectively avoids the misjudgment or omission that may occur based solely on a single frame image and improves the accuracy of recognizing complex dynamic events.

[0014] In one possible implementation of the first aspect, the first analysis result also includes the confidence level corresponding to the reasoning result; the verification conditions include the confidence level being less than a preset value, and / or the reasoning result indicating the existence of an intrusion object.

[0015] In this embodiment, the confidence level output by the first monitoring model is used to characterize the confidence in the reasoning result. Therefore, when the confidence level output by the first monitoring model is low, the second monitoring model performs a precise review to improve the identification accuracy and reduce the possibility of missed detections and false alarms. On the other hand, for high-risk events involving intrusion targets even with high confidence levels, a forced review is triggered. This ensures security while minimizing the number of times the large model is called and the computational cost, achieving a dynamic balance between detection accuracy, response speed, and resource consumption.

[0016] In one possible implementation of the first aspect, the first analysis result further includes the confidence level corresponding to the inference result; after obtaining the first analysis result, the method further includes: Obtain the feature map of the last layer of the first monitoring model; Input the feature map and confidence level into the routing network module to obtain the result of whether the verification conditions are met.

[0017] In this embodiment, the feature map and confidence score of the last layer of the first monitoring model are input into the routing network module. The routing network module intelligently determines whether the verification conditions are met. Since the feature map contains the model's deep representation information of the input data, it can reflect the intrinsic image quality attributes that are difficult to fully express by a single confidence score value, such as the edge clarity of the target, the degree of occlusion, and the distinction between the target and the background. Therefore, the routing network module learns the complex mapping relationship between the feature map and the confidence score to determine whether verification is required. This further reduces the number of invalid calls to the large model and the consumption of computing resources while ensuring that no one is missed in high-risk scenarios.

[0018] In one possible implementation of the first aspect, the method further includes: Receive the second analysis result sent by the backend processor; Training samples are generated based on the results of the second analysis and the corresponding monitoring data. The first monitoring model is optimized using training samples.

[0019] This application embodiment utilizes the second analysis results output by the second monitoring model to form training samples, continuously optimizing the first monitoring model and improving its performance. After the performance of the first monitoring model is improved, the number of calls to the second monitoring model is reduced, achieving a positive loop.

[0020] In one possible implementation of the first aspect, after acquiring the monitoring data collected from the target monitoring area, the method further includes: The system checks the quality of the monitoring data. If the data quality does not meet the analysis requirements, an alarm is output.

[0021] In this embodiment of the application, when monitoring data with poor quality is collected, this data cannot be analyzed by the intelligent model. At this time, an alarm prompt is output to allow manual verification, thus ensuring the safety of the target monitoring area.

[0022] Secondly, embodiments of this application provide another machine vision-based intrusion object detection method, including: Acquire monitoring data collected from the target monitoring area; The monitoring data is analyzed using a first monitoring model to obtain a first analysis result; the first analysis result includes an inference result; the inference result is used to characterize whether there is an intrusion object in the target monitoring area, and related information about the intrusion object; If the first analysis result meets the verification conditions, input information is generated based on the monitoring data and the reasoning result. The input information is analyzed again by the second monitoring model to obtain a second analysis result; wherein, the second analysis result is used to characterize whether there is an intrusion object in the target monitoring area, and related information about the intrusion object; the number of parameters of the first monitoring model is less than the number of parameters of the second monitoring model, and the performance of the first monitoring model is lower than that of the second monitoring model.

[0023] In this embodiment, a first monitoring model is used to analyze monitoring data to obtain a first analysis result. If the first analysis result meets the verification conditions, a second monitoring model is activated for verification. Since the second monitoring model has higher performance and detection accuracy than the first monitoring model, the detection accuracy can be improved by combining the first and second monitoring models to detect intrusion objects.

[0024] In one possible implementation of the second aspect, the first monitoring model is a neural network model, and the second monitoring model is a model based on the Transformer architecture.

[0025] In this embodiment, the first monitoring model is a neural network model, which can achieve real-time preliminary screening of multi-source sensor data. It has low computational cost and high speed, but relatively low accuracy. The second monitoring model is a Transformer-based model, which has powerful global context modeling, long-distance dependency capture, and cross-modal semantic alignment capabilities. By combining the first and second monitoring models, an architecture of local perception and global cognition is constructed, achieving a balance between detection efficiency, resource consumption, and recognition accuracy.

[0026] In one possible implementation of the second aspect, the first monitoring model is deployed on the front-end processor, and the second monitoring model is deployed on the back-end processor.

[0027] This application embodiment fully leverages the local processing advantages of edge computing, such as low latency, high bandwidth utilization, and data privacy protection, by deploying the first monitoring model on the front-end processor and the second monitoring model on the back-end processor. This enables the first monitoring model to perform real-time preliminary screening of multi-sensor data, while the second monitoring model performs in-depth verification, significantly reducing the continuous occupation of network bandwidth and the ineffective consumption of cloud computing power.

[0028] Thirdly, embodiments of this application provide an intrusion object detection device based on machine vision, comprising: The data acquisition module is used to acquire monitoring data collected from the target monitoring area; The first analysis module is used to analyze the monitoring data using the first monitoring model to obtain the first analysis result; the first analysis result includes the reasoning result; the reasoning result is used to characterize whether there is an intrusion object in the target monitoring area, and related information about the intrusion object; The input information generation module is used to generate input information for the second monitoring model based on monitoring data and inference results, when the review conditions are met according to the first analysis result. The second analysis module is used to further analyze the input information through the second monitoring model to obtain a second analysis result. The second analysis result is used to characterize whether there is an intrusion object in the target monitoring area and related information about the intrusion object. The number of parameters of the first monitoring model is less than that of the second monitoring model, and the performance of the first monitoring model is lower than that of the second monitoring model.

[0029] Fourthly, embodiments of this application provide an electronic device, including: a processor, a memory, and a bus, wherein: The processor and memory communicate with each other via a bus; The memory stores program instructions that can be executed by the processor, and the processor can execute the method of the first aspect by calling the program instructions.

[0030] Fifthly, embodiments of this application provide a non-transitory computer-readable storage medium, comprising: A non-transitory computer-readable storage medium stores computer instructions that cause the computer to perform the methods in various possible implementations of the first aspect.

[0031] In a sixth aspect, embodiments of this application provide a computer program product, including computer program instructions, which, when read and executed by a processor, perform the methods in various possible implementations of the first aspect.

[0032] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing embodiments of this application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings. Attached Figure Description

[0033] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1The overall structure of the intrusion object detection system provided in the embodiments of this application; Figure 2 This is a schematic flowchart of an intrusion object detection method provided in an embodiment of this application; Figure 3 This is a schematic diagram of another intrusion object detection method provided in an embodiment of this application; Figure 4 This is a schematic diagram of an intrusion object detection device provided in an embodiment of this application; Figure 5 This is a schematic diagram of another intrusion object detection device provided in an embodiment of this application; Figure 6 This is a schematic diagram of the physical structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0035] The embodiments of the technical solution of this application will now be described in detail with reference to the accompanying drawings. These embodiments are only used to more clearly illustrate the technical solution of this application and are therefore merely examples, and should not be used to limit the scope of protection of this application.

[0036] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this application; the terms “comprising” and “having”, and any variations thereof, in the specification and the foregoing description of the drawings are intended to cover non-exclusive inclusion.

[0037] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.

[0038] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0039] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.

[0040] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), similarly, "multiple sets" refers to two or more (including two sets), and "multiple pieces" refers to two or more (including two pieces).

[0041] In the description of the embodiments of this application, unless otherwise expressly specified and limited, technical terms such as "installation," "connection," "joining," and "fixing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. For those skilled in the art, the specific meaning of the above terms in the embodiments of this application can be understood according to the specific circumstances.

[0042] In security scenarios, to achieve the classification and identification of intrusion targets, current technologies typically employ CNN models (such as ResNet or YOLO) for training. The image to be detected is then directly input into the pre-trained CNN model to obtain the classification result (the target's location, classification, and confidence probability). The advantages of CNN models are their low computational requirements and fast processing speed, fully meeting the real-time requirements of security scenarios. However, their disadvantages include reduced accuracy and decreased usability in specific scenarios. For example, in specific outdoor scenarios such as low light, strong light interference, rain, snow, fog, and dust storms, the accuracy of commonly used CNN models drops rapidly, severely impacting their effectiveness. Furthermore, the models have poor generalization ability; adding a new classification task requires extensive data collection, annotation, and training, which is very time-consuming and labor-intensive.

[0043] Multimodal large-scale models based on Transformer technology (focusing only on image and video inference here) are rapidly discovered. Their advantages include strong generalization ability due to massive data training, making them readily usable in security scenarios without additional training (fine-tuning and alignment). Large models can maintain high accuracy even in specific scenarios, sometimes surpassing human recognition capabilities in certain areas. However, their disadvantages include large model size, high computational requirements, and currently long processing times (latency and throughput; for example, with a 1080P image input, the first token takes several seconds, and the complete result takes tens of seconds), failing to meet the real-time requirements of security scenarios.

[0044] In real-world projects, the number of cameras is large, ranging from dozens to hundreds (visual sensors attached to video surveillance systems or main units) or even thousands (visual sensors on detectors). Even after optimization with prompt words, a large video model typically takes several seconds or longer to obtain a complete text result from a single 1080P image, making it unsuitable for high-concurrency processing. Furthermore, large models often perform frame extraction during video processing, not processing all frames. For example, extracting 5 frames from 25 frames per second for processing can easily cause the large model to miss crucial information. Smaller CNN models, on the other hand, are faster and can quickly process multiple frames. In security scenarios, most of the time the system is in a non-intrusive (normal) state, with only a small portion in an abnormal state. Therefore, a small model can be used to filter out images from abnormal periods and feed them into a larger model for secondary verification. Assuming the system generates one alarm and one image requiring verification every N seconds (e.g., 5 seconds), a large model optimized with AI Infrastructure can support this.

[0045] Therefore, this application proposes a machine vision-based intrusion detection system, which includes multiple detector devices, a detection host (also known as a front-end processor), and a back-end processor. The multiple detector devices can communicate with the detection host, and the detection host can communicate with the back-end processor. During intrusion detection analysis, the detection host can share some or all of the data processing and intrusion detection work of the back-end processor based on its status, thereby reducing data transmission and processing pressure on the back-end processor, improving intrusion detection efficiency, and meeting the real-time requirements of security scenarios.

[0046] The following is in conjunction with the appendix Figure 1 The overall structure of an intrusion object detection system provided by some embodiments of this application is illustrated by way of example.

[0047] like Figure 1 As shown, some embodiments of this application provide a system for detecting intrusion targets. This system includes multiple detection devices 100, a detection host 200, and a post-processing terminal 300. Multiple detection devices 100 (e.g., 20 detection devices) are connected to one detection host 200, and multiple detection hosts 200 are connected to the post-processing terminal 300. In a perimeter security system, the detection devices 100 and detection host 200 can be deployed sequentially at certain intervals. Each detection device can correspond to an independent number and location information. The post-processing terminal 300 pre-stores the number and location information of each detection device and establishes a correspondence between the number and location information of each detection device and the detection host 200. Although... Figure 1The diagram only shows one probe host 200 connected to the post-processing terminal 300, but in actual operating environments, there may be multiple probe hosts 200 connected to the post-processing terminal 300.

[0048] The functions of each unit in the intrusion detection system are illustrated below.

[0049] Each of the multiple detection devices 100 is used to: collect at least the monitoring data of the detection area.

[0050] The detection device 100, also known as a detector, can be installed on the fence of a perimeter security system. Its maximum detection range is called the detection area. In some implementations, a single detection device 100 integrates multiple sensors, such as a vibration sensor, a wide-angle camera, and a Doppler radar. Vibration detection is a passive detection method, while Doppler microwave radar (Doppler radar for short) and visual sensors (such as a wide-angle camera) form an early warning zone outside the fence, which is an active detection method. When a target enters the early warning zone, a suspicious target can be detected through Doppler data analysis or video image analysis, generating an intrusion warning signal.

[0051] Specifically, vibration sensors can collect vibration detection signals generated by the fence due to external factors; wide-angle cameras can collect image detection data within the detection area; and Doppler radar can collect spectral signals within the detection area. Each data segment collected by different detection devices 100 carries the corresponding device number. Detection device 100 can also preprocess the collected vibration detection signals and spectral signals (e.g., signal scaling, filtering, or encoding) to obtain monitoring data that meets the requirements of subsequent intrusion detection analysis. The monitoring data includes the raw vibration data corresponding to the vibration detection signals, the raw spectral data corresponding to the spectral signals, and the image detection data.

[0052] In some implementations, in addition to data acquisition and data preprocessing functions, the detection device 100 can also perform data splicing and feature extraction operations on monitoring data under special circumstances (e.g., when the detection host 200 is under high load or data processing is abnormal) to obtain detection data.

[0053] Optionally, multiple sensors can be integrated into the detection device 100 in a pluggable manner, allowing for flexible deployment of different sensors as needed.

[0054] Alternatively, while one detection device 100 may be a single sensor, multiple different types of detection devices 100 can be deployed at a single location. For example, a vibration sensor, a wide-angle camera, and a Doppler radar can be deployed at the same location. These three sensors are discrete components and are not integrated together. Of course, the deployment locations of different types of detection devices 100 do not have to be at the same point, but rather at nearby points.

[0055] The detection host 200 is used to: acquire point cloud data and detection data of the detection area; wherein, the detection data includes vibration detection data, Doppler spectrum data or image detection data; the detection data is acquired based on the monitoring data.

[0056] The image detection data can be collected by the detection device 100 or by a visual sensor deployed on the detection host 200. The specific method can be determined based on the actual situation.

[0057] The detection host 200 can connect to multiple detection devices 100 via a bus topology to acquire detection data related to the detection areas of the multiple detection devices 100; the detection data includes vibration detection data, Doppler spectrum data, and image detection data. The detection host 200 acquires the detection data in several ways: First, the detection host 200 directly receives the monitoring data sent by the detection device 100, and processes the received monitoring data (e.g., data splicing and feature extraction operations) to obtain the detection data.

[0058] Secondly, the detection device 100 processes the monitoring data, generates detection data, and sends it to the detection host 200.

[0059] Third, for a portion of the monitoring data, the detection device 100 processes it to generate detection data and sends it to the detection host 200; while for the remaining monitoring data, the detection host 200 processes it to generate detection data.

[0060] For example, a stable and unchanging algorithm is embedded in the detection device 100, while potentially variable algorithms are placed on the detection host 200. Exemplarily, quantization filtering of the vibration sensor is performed by the detection device 100, while feature extraction is performed by the detection host 200.

[0061] For example, preprocessing that can significantly compress the data volume is performed by the probe device 100, while preprocessing with a low compression ratio is performed by the probe host 200. Example: FFT compresses the raw IQ data into a spectrogram (high compression ratio), which is performed by the probe device 100.

[0062] For example, multimodal data from the same physical location that requires strict temporal and spatial alignment can be jointly processed by the same device to avoid asynchrony caused by transmission. For instance, vibration data and Doppler data can be jointly encoded by the detection device 100.

[0063] Fourth, by default, the monitoring data is processed by the detection host 200. However, when the detection host 200 is overloaded, some or all of the data processing will be transferred to the detection device 100 or backend processing.

[0064] In some embodiments of this application, a first monitoring model may be deployed on the detection host 200. This first monitoring model performs intrusion detection on the monitoring data and obtains a first analysis result. Furthermore, if the first analysis result meets the verification requirements, input information is generated and sent to the backend processor. It should be noted that the specific implementation method of intrusion detection by the detection host 200 will be described later.

[0065] In some embodiments of this application, the post-processing end 300 is used to: detect the detection data and the data to be detected obtained by filtering the point cloud data, and determine the intrusion detection result; wherein, the intrusion detection result includes whether there is an intrusion behavior and an intrusion object, and the intrusion object category and intrusion action when the intrusion object exists.

[0066] The post-processor 300, also known as the back-end processor, is communicatively connected to the detection host 200. It can receive input information sent by the detection host and analyze the data to determine the final intrusion detection result. The specific analysis method is described in subsequent embodiments. The back-end can store the priority order of multiple detection areas. During data filtering and analysis, data can be processed sequentially according to the priority of the detection area to which it belongs; that is, higher-priority detection areas are processed first, and lower-priority detection areas are processed later.

[0067] In some embodiments, the backend processor can execute the intrusion object detection method provided in this application embodiment. That is, the backend processor can be deployed with a pre-trained monitoring model, which includes a first monitoring model and a second monitoring model. For a specific implementation of the intrusion object detection method, please refer to the following embodiments. In this case, the main function of the detection host 200 is to receive monitoring data from the detection device 100 and process the monitoring data.

[0068] In addition, the backend processor can be integrated with the electronic map to display the number, location information, and whether there are alarms or warnings for each detection device 100 on the electronic map; at the same time, when users interact with the electronic map, they can also view the video or image data collected by each detection device 100 or surrounding detection devices in real time.

[0069] Based on the above embodiments, optionally, the detection host 200 may be equipped with a laser point cloud radar, which can collect point cloud data of the detection area. Alternatively, a visual sensor may also be deployed on the detection host 200 to collect point cloud data of the detection area.

[0070] The difference between the laser point cloud radar or visual sensor on the detection host 200 and the laser point cloud radar or visual sensor (if any) on the detector (detection device 100) is that the former has a longer coverage distance (for example, the detection host can cover 100 meters, while the detector covers 10 meters). Since the detection host 200 is generally installed inside the fence, an emergency alarm zone is formed inside the fence; when a target climbs over the fence and enters this zone, the system generates an emergency alarm and reports the target's specific information; the target's specific information includes the associated detector number, location information (e.g., installation location or latitude and longitude coordinates), and the target's orientation measured by the laser radar on the detection host 200.

[0071] The following describes a scenario where the first monitoring model is deployed on the front-end processor and the second monitoring model is deployed on the back-end processor. Figure 2 This is a schematic flowchart of an intrusion object detection method provided in an embodiment of this application. The method includes: Step 201: The front-end processor acquires the monitoring data collected from the target monitoring area.

[0072] The target monitoring area refers to the area that can be detected by any one of the multiple detection devices. Monitoring data can include video data collected by the detection devices. After acquiring the video data, the front-end processor decodes it to obtain single-frame images. It can be understood that the video acquisition module in the detection device can continuously acquire video streams at a fixed frame rate (e.g., 25fps or 30fps). After receiving the video stream, the front-end processor decodes it into raw image frames. During decoding, a frame skipping strategy can be selected, such as processing one frame every two frames, to balance real-time performance and computational power consumption.

[0073] In some embodiments, after obtaining a single-frame image, the front-end processor can preprocess the image, for example, by loading an image quality model (e.g., a ResNet model) to determine if the image has quality problems, such as a black screen or a distorted image. For images that fail to meet quality standards, subsequent steps are not performed; instead, an alarm is output for manual verification, ensuring the safety of the target monitoring area. To avoid false alarms, an alarm can be triggered when multiple consecutive frames show quality problems, such as 10 or 20 consecutive frames.

[0074] Step 202: The front-end processor uses the first monitoring model to analyze the monitoring data and obtain the first analysis result.

[0075] The first monitoring model can be a lightweight convolutional neural network architecture, such as YOLOv5s or MobileNet-SSD. The number of parameters is typically in the millions to tens of millions. This first monitoring model is pre-trained using training samples, which can include images containing intrusion targets and images without intrusion targets. The first analysis result includes an inference result, which characterizes whether an intrusion target exists in the target monitoring area. Intrusion targets can include people, vehicles, animals, etc. If the inference result indicates the presence of an intrusion target, the inference result also includes relevant information about the intrusion target, such as its actions, physical characteristics, and location information.

[0076] In some embodiments, the first monitoring model may include a classification and recognition model and an action recognition model. The classification and recognition model is responsible for object detection and classification of the decoded multi-frame images, quickly filtering out frames containing objects such as people, vehicles, and objects. When the presence of a person is detected, the action recognition model is triggered to estimate the posture of the human body in that frame (or multiple consecutive frames), output the position of the limb joints, and further infer the type of human action, such as walking, running, climbing, falling, waving, etc.

[0077] The classification and recognition model employs a lightweight neural network and is capable of detecting common categories such as people, vehicles, animals, and objects. After an image is input into the model, it outputs whether a target object exists, the bounding box of each target object, the category label of the target object, and a confidence score for each category.

[0078] Action recognition models can also employ lightweight pose estimation networks, such as MediaPipe Pose, MoveNet, lightweight versions of OpenPose, or small variants of HRNet. These models can output 2D coordinates of 17, 25, 33, or more key joints of the human body (e.g., nose, left and right shoulders, elbows, wrists, hips, knees, ankles, etc.), with a parameter count typically between 1M and 5M, and can run in real-time on a front-end processor. When the classification model detects a human-classified target object in a frame, it inputs that frame into the action recognition model. This model infers the coordinates of the human joints, and based on these coordinates, the action classifier maps the original pose to semantic action labels. The action classifier can determine the action based on the set of relationships between key points, such as angles, distances, and relative positions.

[0079] Since the targets of intrusion are often people, a classification and recognition model can be used to classify and recognize multiple frames of images. If the target object in the image is determined to be a person, an action recognition model can be used to identify its actions. Based on the actions, it can be preliminarily determined whether there is an intrusion behavior.

[0080] Step 203: If the verification conditions are met based on the first analysis result, the front-end processor generates input information according to the monitoring data and inference results.

[0081] Due to the limited computing power of the front-end processor and the small model capacity, the first monitoring model performs well in normal weather and simple backgrounds. However, it suffers from low detection accuracy in severe weather (such as low light, rain, fog, and sandstorms) or when the target is partially obscured. Therefore, after obtaining the first analysis result, it is possible to determine whether the verification conditions are met based on this result.

[0082] In some embodiments, the review criteria can be a combination of the following logic: (1) The confidence level is less than the preset threshold (e.g., 0.85); where the confidence level is the one included in the first analysis result output by the first monitoring model, and this confidence level is used to characterize the credibility of the inference result output by the first monitoring model. When the confidence level is less than the preset threshold, it indicates that the first monitoring model lacks confidence, and the first analysis result needs to be reviewed. It should be noted that the preset threshold can be a fixed value. The lower the preset threshold is set, the lower the security requirement is, and vice versa. In some embodiments, the preset threshold can also be dynamically changed. Specifically, the specific value of the preset threshold can be adaptively adjusted according to the load of the second monitoring model. For example, when the current load of the second monitoring model is large, the preset threshold can be lowered; otherwise, the preset threshold can be raised. However, the preset threshold cannot be increased or decreased indefinitely, and needs to be dynamically adjusted within a reasonable range.

[0083] (2) The reasoning result indicates the existence of an intrusion target; this condition ensures that any event involving security risks must be confirmed by the second monitoring model, regardless of the confidence level of the first monitoring model.

[0084] When conditions (1) and / or (2) are met, it is determined that a review is required; otherwise, the first analysis result is directly used as the final output, and the back-end processor is no longer triggered for processing.

[0085] In some embodiments, a method for determining whether the verification conditions are met can also be based on the routing network module, as follows: Obtain the feature map of the last layer of the first monitoring model; Input the feature map and confidence level into the routing network module to obtain the result of whether the verification conditions are met.

[0086] The routing network module can be pre-trained. During training, the inputs to the routing network module are feature maps and confidence scores, with a label indicating whether a review is needed. Since the feature maps contain deep representations of the input data by the model, reflecting intrinsic image quality attributes that are difficult to fully express through a single confidence score, such as the edge sharpness of the target, the degree of occlusion, and the distinction from the background, the routing network module learns the complex mapping relationship between feature maps and confidence scores to determine whether a review is needed, thus obtaining more accurate judgment results.

[0087] In addition, the routing network module can also have context memory, that is, the routing network module can refer to the judgment results of the past few frames. If it finds that the confidence level of a certain area is low for several consecutive frames, then the probability of needing to be reviewed is also relatively high.

[0088] In another embodiment, the system can also be configured with a working mode, i.e., whether to switch to automatic review of the large model, including manual switching or timed automatic switching. For example, automatic review is performed during the day when the confidence probability is high, while human-machine collaboration is used at night or in special weather conditions. Therefore, the first analysis result can also include a timestamp, which is used to determine whether the current time is within the time period that needs to be reviewed. If so, the review process begins.

[0089] If the verification is satisfactory, the front-end processor constructs the input information for the second monitoring model based on the monitoring data and inference results. In some embodiments, the monitoring data and inference results can be structured, that is, the inference results can be converted into natural language or standardized text descriptions.

[0090] The input information also includes prompts, such as: Please analyze how many people are in the image and what they are doing? The large model outputs text information, such as how many people / what they are doing / what clothes they are wearing, etc.

[0091] In some embodiments, in addition to the structured information described above, the input information may also include video clips of a preset duration before and after the image containing the intrusion target, and the structured information and video clips may be used as input information.

[0092] Step 204: Send input information to the backend processor.

[0093] The front-end processor can send the constructed input information to the back-end processor via wired or wireless networks. The back-end processor can be deployed on cloud servers or in a central data center, equipped with high-performance GPU clusters, large-capacity memory, and elastic computing resources. A second monitoring model is deployed on the back-end processor to further analyze the input information and obtain a second analysis result. This second analysis result characterizes whether an intrusion target exists in the target monitoring area and provides information about the intrusion target. The first monitoring model has fewer parameters than the second monitoring model, and its performance is lower. The second monitoring model can be a large language model based on the Transformer architecture, such as Qwen2.5-VL, GPT-4V, or LLaVA, with a parameter count typically in the billions to tens of billions, far exceeding that of the first monitoring model.

[0094] The second analysis result also indicates whether there is an intrusion object in the target monitoring area, the specific information of the intrusion object (category, location, behavior description), and can be accompanied by confidence level and explanatory text. For example, if it is confirmed after verification that the target is a pedestrian, the confidence level of the small model is low due to dense fog, and the point cloud contour and walking trajectory are combined to determine that it is a real intrusion.

[0095] In this embodiment, a first monitoring model is used to analyze monitoring data to obtain a first analysis result. If the first analysis result meets the verification conditions, a second monitoring model is activated for verification. Since the second monitoring model has higher performance and detection accuracy than the first monitoring model, the detection accuracy can be improved by combining the first and second monitoring models to detect intrusion objects.

[0096] Based on the above embodiments, the monitoring data may also include vibration data, Doppler data, and laser point cloud data. It should be noted that the raw data from different sensors have different formats and need to be converted into a two-dimensional matrix or tensor format suitable for the projection module input.

[0097] For vibration data, a bandpass filter can be used to remove low-frequency environmental noise and high-frequency electromagnetic interference, retaining the frequency bands related to the movement of people and vehicles. Then, the continuous signal is segmented into segments according to fixed time windows. A short-time Fourier transform is performed on each window to obtain a time-spectrum. This time-spectrum is a two-dimensional matrix of time × frequency. Its horizontal axis represents the temporal position of the time window, the vertical axis represents the frequency components, and the pixel value represents the energy intensity of that frequency at that moment. Finally, the values ​​of the time-spectrum are normalized so that they are between [0,1], and used as input for subsequent encoding.

[0098] Doppler data is typically presented as range-Doppler or micro-Doppler spectra, allowing us to obtain Doppler images for each frame. The vertical axis represents radial velocity, the horizontal axis represents time or distance, and the intensity represents echo energy. Logarithmic compression is applied to intensity values ​​with large dynamic ranges to ensure both weak and strong signals are clearly represented in the spectrum. Finally, the spectrum is normalized.

[0099] For laser point cloud data, ground points can be removed using the RANSAC algorithm or a plane fitting-based method, while retaining non-ground target points. To reduce computational load, the point cloud space is divided into a voxel grid of a preset size (e.g., 0.1m × 0.1m × 0.1m), retaining the centroid point within each voxel and filtering out redundant points. The filtered point cloud is then projected onto a bird's-eye view (BEV) plane, and finally, the data from each channel is normalized.

[0100] The projection module is responsible for mapping the three heterogeneous intermediate representations (vibrational spectrogram, Doppler spectrogram, and point cloud BEV map) to the visual-language joint latent space of the second monitoring model. This space is established by the large model during the pre-training phase through contrastive learning (such as CLIP) or generative tasks, enabling the alignment of image patches and text fragments into the same embedding space. The output of the projection module is called a pseudo-visual feature map, whose dimensions are completely consistent with the visual feature maps extracted by the large model from real images (e.g., for the ViT architecture, the output is N visual tokens, each with a dimension of D; for a CNN-based large model, the output is an H×W×C feature map).

[0101] The projection module consists of three parallel encoder-projector branches and a fusion layer, namely a vibration encoder, a Doppler encoder, and a point cloud encoder. Each encoder is followed by a learnable multilayer perceptron (MLP) to map the feature vectors to the visual token space of the large model. The fusion layer concatenates the pseudo-visual token sequences from the three outputs along the feature dimensions, and then compresses them through a 1×1 convolutional or linear layer to obtain the final pseudo-visual feature map.

[0102] It should be noted that the pseudo-visual feature map and the real image feature map have the same number of channels in a tensor. For example, if the dimension of the real image feature map is H*W*C, then the dimension of the pseudo-visual feature map is also H*W*C.

[0103] After obtaining the pseudo-visual feature map, the structured information and the pseudo-visual feature map are combined into a multimodal input sample, which is then used as the input information for the second monitoring model.

[0104] It is understandable that the input information, in addition to structured information and pseudo-visual feature maps, may also include video segments of preset durations before and after the corresponding frame images.

[0105] It should be noted that if the edge nodes lack sufficient computing power, the raw sensor data (or intermediate spectra) can be directly sent to the cloud, where a projection module deployed in the cloud will generate a pseudo-visual feature map. However, to save bandwidth, it is preferable to complete the projection at the edge and only transmit the compressed pseudo-visual feature map, as its data volume is much smaller than that of the original point cloud stream.

[0106] This application embodiment uses multimodal data for deep semantic fusion, enabling the second monitoring model to perform joint reasoning using complementary information provided by vibration, Doppler and laser point cloud data, which greatly improves the target recognition accuracy and environmental robustness of the system in complex environments.

[0107] Based on the above embodiments, after the second monitoring model outputs the second analysis result, the backend processor can send the second analysis result to the frontend processor. Upon receiving the second analysis result, the frontend processor combines the second analysis result with its corresponding monitoring data into a training sample. This training sample is then used to optimize the first monitoring model, improving its performance. After the performance of the first monitoring model is improved, the number of calls to the second monitoring model will be reduced, achieving a positive loop.

[0108] The following section describes how the first and second monitoring models can be deployed on the same platform, for example, both on a front-end processor or a back-end processor. We will use deployment on a back-end processor as an example: Figure 3 This is a schematic flowchart of another intrusion object detection method provided in an embodiment of this application. The method includes: Step 301: Obtain monitoring data collected from the target monitoring area; Step 302: Analyze the monitoring data using the first monitoring model to obtain the first analysis result; the first analysis result includes the reasoning result; the reasoning result is used to characterize whether there is an intrusion object in the target monitoring area, and related information about the intrusion object; Step 303: If the review conditions are met based on the first analysis result, generate input information according to the monitoring data and reasoning results; Step 304: Analyze the input information again using the second monitoring model to obtain the second analysis result; wherein, the second analysis result is used to characterize whether there is an intrusion object in the target monitoring area, and the relevant information of the intrusion object; the number of parameters of the first monitoring model is less than the number of parameters of the second monitoring model, and the performance of the first monitoring model is lower than that of the second monitoring model.

[0109] It should be noted that steps 301-303 are the same as steps 201-203 in the above embodiments, except for the execution subject. Therefore, please refer to the specific implementation of steps 201-203, which will not be repeated here. The specific implementation of the backend processor performing further analysis through the second monitoring module in step 304 has also been described in the previous embodiments, and will not be repeated here either.

[0110] In another embodiment, when the first monitoring model and the second monitoring model are deployed on different devices, the first monitoring model can be deployed on an edge node of the intrusion detection system, and the edge node can be a probe host. The second monitoring model can be deployed on a cloud server.

[0111] Figure 4 This is a schematic diagram of an intrusion detection device provided in an embodiment of this application. The device can be a module, program segment, or code on an electronic device. It should be understood that this device is similar to the one described above. Figure 3 The method implementation corresponds to this and can be executed. Figure 3 The specific functions of the device involved in the method embodiment can be found in the description above; to avoid repetition, detailed descriptions are omitted here. The device includes: a data acquisition module 401, a first analysis module 402, an input information generation module 403, and a second analysis module 404, wherein: The data acquisition module 401 is used to acquire monitoring data collected from the target monitoring area; The first analysis module 402 is used to analyze the monitoring data using the first monitoring model to obtain the first analysis result; the first analysis result includes the reasoning result; the reasoning result is used to characterize whether there is an intrusion object in the target monitoring area, and related information about the intrusion object; The input information generation module 403 is used to generate input information for the second monitoring model based on monitoring data and inference results, when the review conditions are met based on the first analysis result. The second analysis module 404 is used to re-analyze the input information through the second monitoring model to obtain a second analysis result; wherein, the second analysis result is used to characterize whether there is an intrusion object in the target monitoring area, and related information of the intrusion object; the number of parameters of the first monitoring model is less than the number of parameters of the second monitoring model, and the performance of the first monitoring model is lower than that of the second monitoring model.

[0112] Figure 5 This is a schematic diagram of another intrusion detection device provided in an embodiment of this application. This device can be a module, program segment, or code on an electronic device. It should be understood that this device is similar to the one described above. Figure 2 The method implementation corresponds to this and can be executed. Figure 2The specific functions of the device involved in the method embodiment can be found in the description above; to avoid repetition, detailed descriptions are omitted here. The device includes: a monitoring data acquisition module 501, a third analysis module 502, an information generation module 503, and an information sending module 504, wherein: The monitoring data acquisition module 501 is used to acquire monitoring data collected from the target monitoring area; The third analysis module 502 is used to analyze the monitoring data using the first monitoring model to obtain the first analysis result; the first analysis result includes the reasoning result; the reasoning result is used to characterize whether there is an intrusion object in the target monitoring area, and related information about the intrusion object; The information generation module 503 is used to generate input information based on monitoring data and reasoning results when the review conditions are met based on the first analysis result. The information sending module 504 is used to send input information to the backend processor; wherein, a second monitoring model is deployed on the backend processor, and the second monitoring model is used to re-analyze the input information to obtain a second analysis result; wherein, the second analysis result is used to characterize whether there is an intrusion object in the target monitoring area, and related information of the intrusion object; the number of parameters of the first monitoring model is less than the number of parameters of the second monitoring model, and the performance of the first monitoring model is lower than that of the second monitoring model.

[0113] Based on the above embodiments, the monitoring data includes video data; the first monitoring model includes a classification recognition model and an action recognition model; the third analysis module 502 is specifically used for: The video data is decoded to obtain multiple frames of images; The classification and recognition model is used to classify and recognize multiple frames of images to obtain classification results; the classification results include whether the target object is contained, and if the target object is contained, the classification results also include the type of the target object; If the target object is a person, the action recognition model is used to perform action recognition on the image to obtain the position of the limb joints in the image, and the action of the person in the image is determined based on the position of the limb joints.

[0114] Based on the above embodiments, the information generation module 503 is specifically used for: The reasoning results are described in a structured manner to obtain structured information; The structured information and the corresponding image are used as the input information.

[0115] Based on the above embodiments, the monitoring data also includes vibration data, Doppler data, and laser point cloud data; the device also includes a data mapping module for: The vibration data, Doppler data, and laser point cloud data are projected onto the visual-language joint latent space of the second monitoring model through the projection module to obtain a pseudo-visual feature map. The information generation module 503 is specifically used for: The structured information, the pseudo-visual feature map, and the corresponding image are used as the input information.

[0116] Based on the above embodiments, the information generation module 503 is specifically used for: The reasoning results are described in a structured manner to obtain structured information; Extract video clips of preset duration before and after an image containing an intrusion target; The structured information and the video clip are used as the input information.

[0117] Based on the above embodiments, the first analysis result also includes the confidence level corresponding to the inference result; the verification conditions include the confidence level being less than a preset value, and / or the inference result indicating the existence of an intrusion object.

[0118] Based on the above embodiments, the first analysis result further includes the confidence level corresponding to the inference result; the device further includes a condition judgment module, used for: Obtain the feature map of the last layer of the first monitoring model; The feature map and the confidence level are input into the routing network module to obtain the result of whether the verification conditions are met, as output by the routing network module.

[0119] Based on the above embodiments, the device further includes an optimization module for: Receive the second analysis result sent by the backend processor; Training samples are generated based on the second analysis results and the corresponding monitoring data; The first monitoring model is optimized using the training samples.

[0120] Based on the above embodiments, the device further includes an alarm module, used for: The data quality of the monitoring data is checked, and if the data quality does not meet the analysis requirements, an alarm is output.

[0121] Figure 6 This is a schematic diagram of the physical structure of the electronic device provided in the embodiments of this application, such as... Figure 6 As shown, the electronic device includes: a processor 601, a memory 602, and a bus 603; wherein: The processor 601 and the memory 602 communicate with each other through the bus 603; The processor 601 is used to call program instructions in the memory 602 to execute the methods provided in the above-described method embodiments.

[0122] Processor 601 can be an integrated circuit chip with signal processing capabilities. The processor 601 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor.

[0123] The memory 602 may include, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0124] This embodiment discloses a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer can perform the methods provided in the above-described method embodiments.

[0125] This embodiment provides a non-transitory computer-readable storage medium that stores computer instructions that cause the computer to execute the methods provided in the above-described method embodiments.

[0126] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0127] Furthermore, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0128] Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0129] In this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, without necessarily requiring or implying any such actual relationship or order between these entities or operations.

[0130] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A machine vision-based intrusion object detection method, characterized in that, Applied to a front-end processor, the method includes: Acquire monitoring data collected from the target monitoring area; The monitoring data is analyzed using a first monitoring model to obtain a first analysis result; the first analysis result includes an inference result; the inference result is used to characterize whether there is an intrusion object in the target monitoring area, and related information about the intrusion object; If the review conditions are met based on the first analysis result, input information is generated according to the monitoring data and the reasoning result. The input information is sent to the backend processor; wherein, a second monitoring model is deployed on the backend processor, and the second monitoring model is used to re-analyze the input information to obtain a second analysis result; wherein, the second analysis result is used to characterize whether there is an intrusion object in the target monitoring area, and related information of the intrusion object; the number of parameters of the first monitoring model is less than the number of parameters of the second monitoring model, and the performance of the first monitoring model is lower than that of the second monitoring model.

2. The method according to claim 1, characterized in that, The monitoring data includes video data; the first monitoring model includes a classification recognition model and an action recognition model; the analysis of the monitoring data using the first monitoring model includes: The video data is decoded to obtain multiple frames of images; The classification and recognition model is used to classify and recognize multiple frames of images to obtain classification results; the classification results include whether the target object is contained, and if the target object is contained, the classification results also include the type of the target object; If the target object is a person, the action recognition model is used to perform action recognition on the image to obtain the position of the limb joints in the image, and the action of the person in the image is determined based on the position of the limb joints.

3. The method according to claim 2, characterized in that, The step of generating input information based on the monitoring data and the reasoning result includes: The reasoning results are described in a structured manner to obtain structured information; The structured information and the corresponding image are used as the input information.

4. The method according to claim 3, characterized in that, The monitoring data also includes vibration data, Doppler data, and laser point cloud data; the method further includes: The vibration data, Doppler data, and laser point cloud data are projected onto the visual-language joint latent space of the second monitoring model through the projection module to obtain a pseudo-visual feature map. The step of using the structured information and the corresponding image as the input information includes: The structured information, the pseudo-visual feature map, and the corresponding image are used as the input information.

5. The method according to claim 2, characterized in that, The step of generating input information based on the monitoring data and the reasoning result includes: The reasoning results are described in a structured manner to obtain structured information; Extract video clips of preset duration before and after an image containing an intrusion target; The structured information and the video clip are used as the input information.

6. The method according to claim 1, characterized in that, The first analysis result also includes the confidence level corresponding to the reasoning result; the verification conditions include the confidence level being less than a preset value, and / or the reasoning result indicating the existence of an intrusion object.

7. The method according to claim 2, characterized in that, The first analysis result also includes the confidence level corresponding to the reasoning result; After obtaining the first analysis result, the method further includes: Obtain the feature map of the last layer of the first monitoring model; The feature map and the confidence level are input into the routing network module to obtain the result of whether the verification conditions are met, as output by the routing network module.

8. The method according to claim 1, characterized in that, The method further includes: Receive the second analysis result sent by the backend processor; Training samples are generated based on the second analysis results and the corresponding monitoring data; The first monitoring model is optimized using the training samples.

9. The method according to any one of claims 1-8, characterized in that, After acquiring the monitoring data collected from the target monitoring area, the method further includes: The data quality of the monitoring data is checked, and if the data quality does not meet the analysis requirements, an alarm is output.

10. A machine vision-based intrusion object detection method, characterized in that, include: Acquire monitoring data collected from the target monitoring area; The monitoring data is analyzed using the first monitoring model to obtain a first analysis result; The first analysis result includes the reasoning result; The reasoning result is used to characterize whether there is an intrusion object in the target monitoring area, and related information about the intrusion object; If the review conditions are met based on the first analysis result, input information is generated according to the monitoring data and the reasoning result. The input information is analyzed again by the second monitoring model to obtain a second analysis result; wherein, the second analysis result is used to characterize whether there is an intrusion object in the target monitoring area, and related information about the intrusion object; The first monitoring model has fewer parameters than the second monitoring model, and the performance of the first monitoring model is lower than that of the second monitoring model.

11. The method according to claim 10, characterized in that, The first monitoring model is a neural network model, and the second monitoring model is a model based on the Transformer architecture.

12. The method according to claim 10 or 11, characterized in that, The first monitoring model is deployed on the front-end processor, and the second monitoring model is deployed on the back-end processor.

13. An intrusion object detection device based on machine vision, characterized in that, include: The data acquisition module is used to acquire monitoring data collected from the target monitoring area; The first analysis module is used to analyze the monitoring data using the first monitoring model to obtain a first analysis result; The first analysis result includes the reasoning result; The reasoning result is used to characterize whether there is an intrusion object in the target monitoring area, and related information about the intrusion object; The input information generation module is used to generate input information for the second monitoring model based on the monitoring data and the reasoning result, when the review conditions are met based on the first analysis result. The second analysis module is used to further analyze the input information through the second monitoring model to obtain a second analysis result; wherein, the second analysis result is used to characterize whether there is an intrusion object in the target monitoring area, and related information about the intrusion object; The first monitoring model has fewer parameters than the second monitoring model, and the performance of the first monitoring model is lower than that of the second monitoring model.

14. An electronic device, characterized in that, include: Processor, memory, and bus, among which: The processor and the memory communicate with each other via the bus; The memory stores program instructions that can be executed by the processor, and the processor can execute the method as described in any one of claims 1-9 by calling the program instructions.

15. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions that, when executed by a computer, cause the computer to perform the method as described in any one of claims 1-9.

16. A computer program product, characterized in that, It includes computer program instructions, which, when read and executed by a processor, perform the method as described in any one of claims 1-9.

17. An intrusion object detection system based on machine vision, characterized in that, It includes a front-end processor and a back-end processor; the front-end processor is used to acquire monitoring data collected from the target monitoring area; and to analyze the monitoring data using a first monitoring model to obtain a first analysis result; The first analysis result includes the reasoning result; The reasoning result is used to characterize whether there is an intrusion object in the target monitoring area, and related information about the intrusion object; If the verification conditions are met based on the first analysis result, input information is generated according to the monitoring data and the reasoning result; The backend processor is used to re-analyze the input information through the second monitoring model to obtain a second analysis result; wherein, the second analysis result is used to characterize whether there is an intrusion object in the target monitoring area, and related information about the intrusion object; The first monitoring model has fewer parameters than the second monitoring model, and the performance of the first monitoring model is lower than that of the second monitoring model.