Camera cooperative monitoring method and system based on multi-modal perception

Through the camera collaborative monitoring method based on multimodal perception, the accuracy and real-time problems of traditional monitoring systems under the influence of light changes, environmental noise, target occlusion and other factors are solved, and efficient, intelligent, safe and energy-saving monitoring effects are achieved.

CN120223845AActive Publication Date: 2025-06-27JIANGXI BOSHI INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510487611.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-06-27
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

Under the influence of light changes, environmental noise, target occlusion and other factors, traditional video surveillance systems are difficult to achieve high accuracy and real-timeness, and lack effective linkage mechanisms and spatial and temporal data fusion methods, resulting in discontinuous target tracking, insufficient privacy protection, and unoptimized energy consumption.

Method used

A camera collaborative monitoring method based on multimodal perception is adopted to detect abnormal event signals through non-visual sensors, trigger camera activation and visual capture; relay tracking is achieved using feature matching to generate continuous motion trajectory; a multi-view angle collaborative occlusion mechanism is used to protect privacy-sensitive information; during the non-event trigger period, the flow density is predicted based on historical data, and the camera energy consumption is optimized.

Benefits of technology

It improves the system response speed, tracking accuracy and security, effectively prevents privacy leakage, reduces energy consumption, and adapts to multi-scene applications in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223845A_ABST
    Figure CN120223845A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video monitoring, and discloses a camera cooperative monitoring method and system based on multi-modal perception, and the method comprises the steps: detecting an abnormal event signal in an environment through a non-visual sensor node, triggering the activation of a camera, and carrying out the visual capture of a target region; the camera performs relay tracking based on target feature matching to generate a continuous motion track; performing multi-view collaborative shielding on privacy sensitive information in the tracking target, generating desensitized monitoring data, and uploading the desensitized monitoring data; and optimizing the energy consumption of the camera in a non-event triggering period, and updating the behavior recognition model based on the desensitization data. According to the method, the contradiction between privacy protection and monitoring efficiency is solved, cross-regional model evolution and energy consumption reduction are realized, and the problems of response lag, data redundancy, privacy disclosure and high energy consumption of traditional monitoring are systematically avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video surveillance, and in particular to a camera collaborative surveillance method and system based on multimodal perception. Background Art

[0002] With the continuous maturity of Internet of Things, edge computing and artificial intelligence technologies, the application of multimodal perception in the field of intelligent surveillance has shown a rapid development trend. Traditional surveillance mostly relies on a single camera or video surveillance means, and realizes surveillance and event judgment through image recognition and target detection algorithms. However, due to many factors such as light changes, occlusion interference, and complex background, its accuracy and real-time performance often fail to meet the high requirements of modern security prevention and control.

[0003] There are still many limitations in the existing technologies in practical applications. First of all, single video surveillance is often affected by factors such as environmental noise, light fluctuations, and target occlusion, resulting in misjudgment or missed judgment of abnormal events; even if non-visual sensors are introduced, without effective linkage mechanisms and spatio-temporal data fusion means, it is difficult to achieve all-round and continuous accurate tracking of targets. Secondly, in the process of generating target trajectories in traditional surveillance systems, due to time delay and angle differences in data acquisition between cameras, the smoothness of continuous motion trajectories is insufficient, and it is difficult to perform reasonable interpolation processing when continuous coordinate data is lost, directly affecting the training effect of subsequent behavior recognition models. In addition, for surveillance data involving privacy-sensitive information, most existing solutions do not start from the perspective of multi-view collaborative masking to ensure data desensitization while maintaining the availability of information. At the same time, the surveillance system often fails to make full use of historical data to predict the population density in the area during non-event trigger periods, so as to achieve intelligent optimization of energy consumption and dynamic adjustment of cameras. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a camera collaborative surveillance method based on multimodal perception, which can solve problems such as light, angle limitation, data loss, and insufficient privacy protection in traditional surveillance, thereby improving the system response speed, tracking accuracy and security, effectively preventing privacy leakage and reducing energy consumption, and adapting to multi-scenario applications in complex environments.

[0006] To solve the above technical problems, the present invention provides the following technical solutions. A camera collaborative monitoring method based on multi-modal perception includes: detecting abnormal event signals in the environment through non-visual sensor nodes, triggering the activation of cameras and performing visual capture on the target area; the first camera that captures the target extracts the color block distribution and movement direction of the target, broadcasts them to adjacent cameras through a local area network, and realizes relay tracking based on feature matching to generate a continuous movement trajectory; different cameras perform perspective coordinate mapping on the same privacy-sensitive target, negotiate to generate complementary shielding areas, and generate desensitized monitoring data; during non-event trigger periods, predict and turn off the power of some cameras according to historical activity, and the remaining cameras switch to a wide-angle mode to cover the entire area; the cloud updates the global behavior recognition model using a federated learning framework based on the desensitized movement trajectory data.

[0007] As a preferred solution of the camera collaborative monitoring method based on multi-modal perception according to the present invention, wherein: the non-visual sensors include a sound sensor and a vibration sensor, wherein:

[0008] The sound sensor collects environmental sound wave signals, extracts spectral features and matches them with abnormal sound pattern templates in a preset sound pattern library;

[0009] The vibration sensor is deployed at the boundary of the monitoring area, and determines climbing or impact events by detecting the vibration frequency and duration through an accelerometer.

[0010] As a preferred solution of the camera collaborative monitoring method based on multi-modal perception according to the present invention, wherein: the performing visual capture includes, when the non-visual sensor detects an abnormal signal, synchronously activating the cameras within the sensor transmission range through a hardware trigger signal to obtain the video stream of the corresponding area;

[0011] Calculate the time difference between the abnormal signal and the movement of the object in the camera image based on the sound and light propagation speeds. When the time difference is less than the target value and the detected displacement of the object in the image exceeds the target displacement, it is determined as an effective event trigger.

[0012] As a preferred solution of the camera collaborative monitoring method based on multi-modal perception according to the present invention, wherein: the generating a continuous movement trajectory includes, the first camera that captures the target extracts the color block distribution and movement direction of the target, converts the pixels of the target area to the HSV color space, and extracts the continuous area with a set saturation value as the color block; calculates the movement direction angle through the centroid offset of the block between adjacent frames;

[0013] Broadcast the HSV histogram of the color block and the movement direction angle to adjacent cameras in the same local area network through the UDP protocol;

[0014] Each camera uploads the two-dimensional coordinates and timestamps of the target in its own coordinate system to the edge server;

[0015] The edge server aligns the data of each camera according to the timestamps, and uses the cubic spline interpolation algorithm to generate a smooth trajectory for the missing coordinate points between two consecutive frames. When the number of continuously missing coordinate points in the motion trajectory is greater than the target missing quantity, this section is marked as an occlusion and missing section.

[0016] As a preferred solution of the camera collaborative monitoring method based on multi-modal perception according to the present invention, wherein: the multi-view collaborative occlusion includes that each camera performs perspective coordinate mapping on the same privacy-sensitive target, and converts the target pixel coordinates to the same world coordinate system through calibration parameters; according to the three-dimensional position of the target in the world coordinate system, calculate the projection of the occlusion area under the perspective of each camera;

[0017] Among them, the generation rule of the complementary occlusion area is to evenly divide the target area into grids to generate sub-areas with the same number as the number of cameras; allocate at least one sub-area to each camera, and use the dynamic mosaic algorithm during occlusion, and the mosaic granularity is adaptively adjusted according to the moving speed of the target.

[0018] As a preferred solution of the camera collaborative monitoring method based on multi-modal perception according to the present invention, wherein: the non-event trigger period includes training a regional activity model based on historical monitoring data to predict the pedestrian flow density in each period within a preset future period;

[0019] When the predicted pedestrian flow density is lower than the preset activity threshold, turn off the power of the first preset proportion of cameras in the target area, and the first preset proportion is calculated according to the area of the area and the total number of cameras, and the remaining cameras are switched to the wide-angle mode to cover the entire area.

[0020] As a preferred solution of the camera collaborative monitoring method based on multi-modal perception according to the present invention, wherein: the updating of the behavior recognition model includes transmitting the desensitized motion trajectory data to the cloud server, and extracting the motion direction angle and speed change sequence in the motion trajectory as training features; using the federated learning framework to perform gradient descent optimization on the global behavior recognition model and update the weight parameters of the fully connected layer;

[0021] Encrypt the parameters of the updated global behavior recognition model and send them to each edge node through the HTTP / 2 protocol.

[0022] As a preferred solution of the camera collaborative monitoring system based on multi-modal perception according to the present invention, wherein: it includes an anomaly detection module, a target tracking module, a privacy protection module, and an energy-saving update module;

[0023] The abnormal detection module uses non-visual sensors to monitor the environment in the monitoring area in real time, generates event trigger signals, and synchronously activates the cameras within the transmission range of the sensors through hardware trigger instructions to start visual capture.

[0024] The target tracking module, after an event is triggered, realizes the collaborative operation between cameras, and performs continuous and seamless visual tracking and data acquisition on the target.

[0025] The privacy protection module performs masking processing on the parts of the targets captured during monitoring that contain privacy-sensitive information, so that the monitoring data can meet the security requirements without infringing on personal privacy.

[0026] The energy-saving update module, during the period when non-abnormal events are not triggered, realizes the optimization of camera energy consumption through intelligent control; uses the desensitized monitored motion trajectory data, and adopts an advanced federated learning framework to perform online optimization and update on the global behavior recognition model.

[0027] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the camera collaborative monitoring method based on multi-modal perception are realized.

[0028] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the steps of the camera collaborative monitoring method based on multi-modal perception are realized.

[0029] Advantages of the present invention: In the present invention, the triggering of abnormal events and the activation of cameras are completed in advance by non-visual sensors, which not only reduces the influence of light and viewing angle limitations on event detection, but also improves the response speed and monitoring coverage rate of the system. By extracting the color block distribution and motion direction of the target, and then using the UDP protocol to achieve data broadcast and edge server data alignment within the local area network, a smooth and continuous target motion trajectory can be generated, and the cubic spline interpolation is used to solve the problem of data loss, thereby ensuring the tracking accuracy. The multi-view collaborative masking mechanism adopts a dynamic mosaic algorithm and uniform grid division to achieve complementary masking of privacy-sensitive areas, effectively preventing the risk of privacy infringement caused by the leakage of monitoring data. At the same time, based on the prediction of regional activity based on historical monitoring data, it is possible to intelligently reduce the workload of some cameras during low-traffic periods, achieving the effect of energy conservation and consumption reduction. The present invention provides an efficient, intelligent, secure and energy-saving comprehensive solution for the monitoring system, with broad application prospects and obvious technical advantages. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0031] Figure 1 Schematic diagram of the process of the camera collaborative monitoring method based on multi-modal perception provided by an embodiment of the present invention. Specific embodiments

[0032] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the accompanying drawings of the specification. Obviously, the described embodiments are some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0033] Embodiment 1, referring to Figure 1 , which is the first embodiment of the present invention. This embodiment provides a camera collaborative monitoring method based on multi-modal perception, including:

[0034] S1: Detect abnormal event signals in the environment through non-visual sensor nodes, trigger the activation of the camera, and perform visual capture on the target area.

[0035] Furthermore, the non-visual sensor nodes include a sound sensor and a vibration sensor, where:

[0036] The sound sensor collects environmental sound wave signals, extracts spectral features, and matches them with abnormal sound pattern templates in a preset sound pattern library. The abnormal sound pattern templates include the frequency band characteristics of the spectral peak of broken glass sound in the frequency range of 2 kHz - 4 kHz;

[0037] Vibration sensors are deployed at intervals of 5 meters at the boundary of the monitoring area, using a three-axis MEMS accelerometer with a sampling frequency of 1 kHz;

[0038] Detect climbing or impact events where the vibration frequency detected by the accelerometer exceeds 50 Hz and the duration is greater than 0.3 seconds.

[0039] It should be noted that when the non-visual sensor detects an abnormal signal, the camera within 3 meters of the distance sensor is synchronously activated through a hardware trigger signal, the video stream of the corresponding area is obtained, the straight-line distance L (meters) between the sensor and the nearest camera is calculated, and the theoretical time difference is calculated based on the sound speed of 340 m / s and the light speed of 3×10^8 m / s:

[0040]

[0041] where Δt is the time difference, v s is the speed of sound, and v1 is the speed of light; if the time difference between the starting moment of the object's movement in the actual image and the trigger signal is within the interval [Δt - δ, Δt + δ] (δ is the error tolerance value), then the displacement verification is performed;

[0042] Furthermore, the background difference method is used to extract the moving area in the target area. In the initialization stage, when there is no target activity, 30 frames of images are continuously collected. By establishing the initial background, each pixel point is described by 3 Gaussian distributions, and the weight coefficients are initialized as [a, b, c]. (For example, the weight coefficients can take values of [0.3, 0.5, 0.2]); the background is updated every 5 seconds. For pixel points that have not changed for more than 10 frames, the weight of their corresponding Gaussian distribution is increased, and the weights of other distributions are attenuated to adapt to the gradually changing illumination scene;

[0043] After differentiating the current frame from the background model, the double-threshold method is used for processing: the area with a high threshold (greater than 50 gray levels) is directly determined as the moving area; the area with a low threshold (20 - 50 gray levels) needs to meet the condition that there are high-threshold points in the adjacent 3×3 window to be retained; the morphological closing operation (3×3 elliptical kernel) is performed on the segmented connected areas to eliminate small noises.

[0044] S2: The camera performs relay tracking based on target feature matching to generate a continuous motion trajectory.

[0045] Furthermore, the camera that first captures the target extracts the color block distribution and motion direction of the target, specifically including: converting the pixels in the target area to the HSV color space;

[0046] The saturation value range is [0, 1]. The higher the value, the more vivid the color, and it is less affected by changes in illumination; the low-saturation area (below 0.4) usually corresponds to shadows, reflections, or grayish-white backgrounds, which are prone to interference; the high-saturation area (above 0.8) usually corresponds to artificial markers (such as warning signs), which are prone to misjudgment. In the present invention, continuous areas with a saturation value higher than 0.6 are extracted as color blocks to filter out shadow interference, making the color block features more focused on the real target, and at the same time adapting to the hardware computing power constraints to meet the real-time monitoring requirements.

[0047] The motion direction angle is calculated through the centroid offset of the blocks between adjacent frames. Corner points with significant gradient changes are extracted within the target area, and at the same time, the minimum eigenvalue threshold is set; for two adjacent frames of images, an 8×8 pixel search window is established at the corner point positions, and the optical flow vector is solved by the iterative weighted least squares method, satisfying:

[0048] ΣW 2 ·(I x u + I y v + It ) 2 →Minimum value

[0049] where W is the Gaussian weight matrix, and I x is the horizontal gradient of the pixel point, and I y is the vertical gradient of the pixel point, u and v are the optical flow vectors, and I t is the time gradient; only the vectors with the matching error (SSD) between the front and rear frames < 5 pixels are retained;

[0050] Perform direction clustering on the retained displacement vectors: divide 0° - 360° into 12 intervals of 30°; count the number of vectors in each interval, and select the median value of the interval with a proportion exceeding 60% as the main direction; the remaining part is determined that the target has turned, and the average direction of the longest continuous trajectory segment is taken.

[0051] Broadcast the HSV histogram and the motion direction angle of the color block to adjacent cameras within the same local area network through the UDP protocol. Each camera uploads the two-dimensional coordinates and time stamps of the target in its own coordinate system to the edge server at a frequency of 10Hz, and the data format is (camera ID, time stamp, image coordinate system);

[0052] The edge server aligns the data of each camera according to the time stamp, uses the cubic spline interpolation algorithm to generate a smooth trajectory for the missing coordinate points between two consecutive frames, and takes the first 2 known trajectory points (P0, P1) and the last 2 points (P2, P3) before and after the missing segment as control points; if the missing segment is at the start or end position of the trajectory, virtual control points are generated by mirroring adjacent points;

[0053] Input the control points into the Catmull - Rom spline generator, set the parameterization step size, generate a smooth path, and adjust the speed consistency of the interpolation points, and calculate the average speed v of the known trajectory segment avg ; limit the distance between interpolation points not to exceed v avg × ±20% of the time interval.

[0054] If more than 5 trajectory points are continuously lost, mark the missing time period [t1, t2], and send a reshooting instruction to the backup camera deployed within a radius of 3 meters from the missing area.

[0055] It should be noted that according to the missing time period [t1, t2] and the nearest trajectory point, predict the target position and generate a spherical search area, with the center of the sphere being the last known coordinate and the radius being: the maximum speed of the trajectory point × (t2 - t1);

[0056] According to the time range and estimated position of the occluded missing section, adjust the pan-tilt angle of the pan-tilt head to 30° - 60°, and shoot and supplement the video in wide-angle mode; perform moving target detection on the supplemented video. If a target that meets the original features is detected, add its coordinates to the trajectory data and smoothly connect it with the original trajectory through a B-spline curve.

[0057] S3: Perform multi-view collaborative occlusion on the privacy-sensitive information in the tracking target, generate desensitized monitoring data and upload it.

[0058] Furthermore, multiple cameras perform perspective coordinate mapping on the same privacy-sensitive target. Each camera converts the target pixel coordinates to the same world coordinate system (X, Y, Z) through calibration parameters. The conversion formula is:

[0059]

[0060] where (c x , c y ) is the optical center coordinate of the camera, (x, y) is the target pixel coordinate, f x , f y are the focal lengths corresponding to the horizontal and vertical coordinates, and Z is obtained by calculating the binocular camera parallax;

[0061] According to the three-dimensional position of the tracking target in the world coordinate system, calculate the projection of the occlusion area under each camera's perspective; each camera occludes 50% of the area preset in the target projection area, and the spatial positions of the occlusion areas of each camera do not overlap;

[0062] Perform semantic segmentation on the target privacy area (face / license plate), obtain its minimum bounding rectangle, and divide the rectangle into m sub-regions (m = the number of cameras participating in the occlusion); assign a sub-region to each camera, and the occlusion algorithm uses dynamic mosaics. The size of the mosaic blocks is adaptively adjusted according to the target movement speed v d :

[0063]

[0064] Furthermore, the EXIF metadata is stripped from the occluded video stream, and the motion trajectory data is stored separately from the video. The motion trajectory data fields include (timestamp, world coordinates (X, Y), speed v d , motion direction); for the unoccluded privacy area part in any camera's video, check whether there is other camera to occlude this part. If so, allow it to be retained, otherwise, append an occlusion instruction to this camera;

[0065] For the occluded target, retain the aspect ratio of its bounding rectangle (for vehicle type / human body judgment) and the motion vector direction, and delete the color histogram features.

[0066] S4: Optimize the energy consumption of the camera during non-event trigger periods and update the behavior recognition model based on the desensitized data.

[0067] Furthermore, during non-event trigger periods, train a regional activity model based on historical monitoring data to predict the pedestrian flow density in each period within a future preset cycle (configurable as 12 / 24 / 48 hours); the activity model uses a prediction model based on time series trained with historical data, with the input being the date type (weekday / holiday) and weather data, and the output being the pedestrian flow in each region in the next 12 / 24 / 48 hours. The prediction model uses an LSTM network with the number of hidden layer units = 64.

[0068] When the predicted pedestrian flow density is lower than the preset activity threshold (the pedestrian flow density on weekdays is the 20th percentile of the historical pedestrian flow density; the pedestrian flow density on holidays is the 80th percentile of the historical pedestrian flow density), turn off the power of the first preset proportion of cameras in the target area. The first preset proportion R is calculated based on the area S of the region and the total number N of cameras:

[0069]

[0070] where α is the effective coverage area of a single camera; the remaining cameras are switched to the 4K ultra-wide-angle mode, the horizontal field of view is extended to 150°, and the frame rate is reduced to 10fps; the cameras with power off are woken up once every 30 minutes and work continuously for 2 minutes to update the background model.

[0071] Furthermore, after the cloud server receives the desensitized motion trajectory data, extract the motion direction angle and speed change sequence in the motion trajectory as training features, and construct a global behavior recognition model; use the federated learning framework to optimize the global behavior recognition model by gradient descent and update the weight parameters of the fully connected layer. Each edge node locally stores the desensitized motion trajectory dataset Di, initializes the global behavior recognition model, removes the original last fully connected layer (the original output dimension is 1000), and adds a new fully connected layer structure as:

[0072] Linear(512→128)→ReLU→Linear(128→3)→Softmax

[0073] where the output 3-class corresponding probability distributions are: normal walking (Class 0), running (Class 1), staying (Class 2); the initial model weights of each node are synchronized to download the initial weight file from the central server, and after loading, all parameters except the newly added fully connected layer are frozen.

[0074] Encrypt the updated global behavior recognition model parameters and send them to each edge node. Use the AES-256 encryption algorithm, adopt the CBC mode, the initialization vector (IV) length is 128 bits, and the padding method is PKCS7; serialize and encrypt the weight file to generate ciphertext data blocks (each group is 16 bytes); after the edge node receives the data packet, decrypt and verify CRC32. If the verification passes, load the new weight into the memory; adopt a double-buffer mechanism to retain a copy of the global behavior recognition model before the update until the updated global behavior recognition model has completed 10 consecutive inferences without anomalies, and then switch to the updated global behavior recognition model.

[0075] Example 2, the second example of the present invention, which is different from the previous example in that:

[0076] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0077] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device.

[0078] More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connections (electronic devices) having one or more wirings, portable computer disk cartridges (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber devices, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0079] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc.

[0080] Embodiment 3, an embodiment of the present invention, provides a camera collaborative monitoring system based on multi-modal perception, including an anomaly detection module, a target tracking module, a privacy protection module, and an energy-saving update module;

[0081] The anomaly detection module uses non-visual sensors to monitor the environment in the monitoring area in real time, generates an event trigger signal, and synchronously activates the cameras within the sensor transmission range and starts visual capture through a hardware trigger instruction;

[0082] The target tracking module, after the event is triggered, realizes collaborative operation between cameras, and performs continuous and seamless visual tracking and data collection on the target;

[0083] The privacy protection module performs masking processing on the part of the target captured during the monitoring process that contains privacy-sensitive information, so that the monitoring data meets the security requirements without infringing on personal privacy.

[0084] The energy-saving update module, during the non-anomaly event trigger period, realizes the optimization of camera energy consumption through intelligent control; uses the desensitized monitored motion trajectory data and adopts an advanced federated learning framework to perform online optimization and update on the global behavior recognition model.

[0085] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A camera collaborative monitoring method based on multimodal perception, characterized in that: include, Detect abnormal event signals in the environment through non-visual sensor nodes, trigger the camera to activate and visually capture the target area; The first camera that captures the target extracts the target's color block distribution and movement direction, broadcasts it to adjacent cameras via the local area network, and implements relay tracking based on feature matching to generate a continuous movement trajectory. Different cameras map the view coordinates of the same privacy-sensitive target, negotiate to generate complementary masking areas, and generate desensitized monitoring data; During non-event trigger periods, some cameras are powered off based on historical activity predictions, and the remaining cameras switch to wide-angle mode to cover the entire area; The cloud uses a federated learning framework to update the global behavior recognition model based on the desensitized motion trajectory data.

2. The camera collaborative monitoring method based on multimodal perception as claimed in claim 1, characterized in that: The non-visual sensors include sound sensors and vibration sensors, wherein: The sound sensor collects environmental sound wave signals, extracts spectrum features and matches them with abnormal voiceprint templates in a preset voiceprint library; The vibration sensor is deployed at the boundary of the monitoring area, and determines climbing or impact events by detecting the vibration frequency and duration through an accelerometer.

3. The camera collaborative monitoring method based on multimodal perception as claimed in claim 2, characterized in that: The visual capture includes, when the non-visual sensor detects an abnormal signal, synchronizing the camera within the sensor transmission range through a hardware trigger signal to obtain a video stream of the corresponding area; The time difference between the abnormal signal and the movement of the object in the camera image is calculated based on the speed of sound and light propagation. If the time difference is less than the target value and the displacement of the object detected in the image exceeds the target displacement, it is determined to be a valid event trigger.

4. The camera collaborative monitoring method based on multimodal perception as claimed in claim 3, characterized in that: Generating a continuous motion trajectory includes: extracting the color block distribution and motion direction of the target by a camera that first captures the target, converting the pixels of the target area into the HSV color space, extracting a continuous area with a set saturation value as a color block; and calculating the motion direction angle by the centroid offset of the block between adjacent frames; Broadcasting the HSV histogram and the motion direction angle of the color block to adjacent cameras in the same local area network through the UDP protocol; Each camera uploads the two-dimensional coordinates and timestamp of the target in its own coordinate system to the edge server; The edge server aligns the camera data according to the timestamp, and uses the cubic spline interpolation algorithm to generate a smooth trajectory for the missing coordinate points between two consecutive frames. When the number of consecutive missing coordinate points in the motion trajectory is greater than the number of target missing points, the segment is marked as an occluded missing segment.

5. The camera collaborative monitoring method based on multimodal perception as claimed in claim 4, characterized in that: The multi-view collaborative masking includes: each camera performs view coordinate mapping on the same privacy-sensitive target, converting the target pixel coordinates into the same world coordinate system through calibration parameters; and calculating the masking area projection under the view angle of each camera according to the three-dimensional position of the target in the world coordinate system; Among them, the generation rule of complementary masking areas is to divide the target area into uniform grids to generate sub-areas with the same number of cameras; at least one sub-area is allocated to each camera, and a dynamic mosaic algorithm is used for masking, and the mosaic granularity is adaptively adjusted with the target moving speed.

6. The camera collaborative monitoring method based on multimodal perception as claimed in claim 5, characterized in that: The non-event triggering period includes training the regional activity model based on historical monitoring data to predict the flow density of people in each period within a future preset period; When the predicted crowd density is lower than the preset activity threshold, a first preset proportion of cameras in the target area is turned off, the first preset proportion is calculated based on the area and the total number of cameras, and the remaining cameras are switched to wide-angle mode to cover the entire area.

7. The camera collaborative monitoring method based on multimodal perception as claimed in claim 6, characterized in that: The updating behavior recognition model includes transmitting the desensitized motion trajectory data to a cloud server, and extracting the motion direction angle and speed change sequence in the motion trajectory as training features; Adopting the federated learning framework, the global behavior recognition model is optimized by gradient descent to update the weight parameters of the fully connected layer; The updated model parameters are encrypted and sent to each edge node via the HTTP / 2 protocol.

8. A system using the camera collaborative monitoring method based on multimodal perception as claimed in any one of claims 1 to 7, characterized in that: It includes anomaly detection module, target tracking module, privacy protection module, and energy-saving update module; The anomaly detection module uses non-visual sensors to monitor the environment in the monitoring area in real time, generates event trigger signals, and synchronously activates cameras within the sensor transmission range and starts visual capture through hardware trigger instructions; The target tracking module realizes the collaborative operation between cameras after the event is triggered, and performs continuous and seamless visual tracking and data collection of the target; The privacy protection module masks the part of the target captured during the monitoring process that contains privacy-sensitive information, so that the monitoring data meets security requirements while not infringing on personal privacy; The energy-saving update module optimizes the camera energy consumption through intelligent control during non-abnormal event triggering periods; utilizes the motion trajectory data monitored after desensitization and adopts an advanced federated learning framework to perform online optimization and update of the global behavior recognition model.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • A surveillance system, a surveillance camera and an image processing method

    CN105898208A

  • Intelligent opening and closing apparatus and method of security and protection camera based on channel state information

    CN107277452A

  • Highway vehicle trajectory tracking method

    CN111667507A

  • Intelligent video identification linkage system and method based on multi-camera physical topology

    CN114422751A

  • Smart park multi-source data dynamic monitoring and real-time analysis system and method

    CN118072255A

Cited By

  • Security monitoring device based on human body perception

    CN120711151A

  • A security monitoring device based on human perception

    CN120711151B

  • Smart park monitoring system based on big data

    CN120747866A

  • A big data-based smart park monitoring system

    CN120747866B

  • Method and system for quickly identifying and tracking low-slow small target

    CN121190518A