Camera cooperative monitoring method and system based on multi-modal perception

By using a multimodal sensing-based camera collaborative monitoring method, the problems of lighting, occlusion, privacy, and energy consumption in traditional monitoring systems are solved, achieving efficient, safe, and energy-saving monitoring results.

CN120223845BActive Publication Date: 2025-12-09JIANGXI BOSHI INTELLIGENT TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510487611.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-12-09
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

Traditional monitoring systems are susceptible to changes in lighting and interference from obstructions, resulting in discontinuous target trajectory generation, insufficient privacy protection, unoptimized energy consumption, and a lack of linkage mechanisms among multiple visual sensors.

Method used

A multimodal perception-based camera collaborative monitoring method is adopted. Abnormal events are detected by non-visual sensors, cameras collaboratively track targets, viewpoint coordinate mapping and privacy masking are performed, historical data is combined to predict people flow density to optimize energy consumption, and federated learning is used to update the behavior recognition model.

Benefits of technology

It improves the response speed and tracking accuracy of the monitoring system, ensures privacy and security, reduces energy consumption, and adapts to multiple application scenarios in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223845B_ABST
    Figure CN120223845B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of video monitoring, and discloses a camera cooperative monitoring method and system based on multi-modal perception, which comprises detecting abnormal event signals in the environment through a non-vision sensor node, triggering camera activation and visual capture of a target area; the camera performs relay tracking based on target feature matching to generate a continuous motion trajectory; multi-view cooperative shielding is performed on privacy-sensitive information in the tracking target to generate desensitized monitoring data and upload the same; camera energy consumption is optimized during a non-event triggering period, and a behavior recognition model is updated based on the desensitized data. The method solves the contradiction between privacy protection and monitoring efficiency, realizes cross-region model evolution and energy consumption reduction, and systematically avoids the problems of traditional monitoring response lag, data redundancy, privacy leakage and high energy consumption.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of video monitoring, and in particular to a camera cooperative monitoring method and system based on multi-modal perception. BACKGROUND

[0002] With the continuous maturity of Internet of Things, edge computing and artificial intelligence technology, the application of multi-modal perception in the field of intelligent monitoring has shown a rapid development trend. Traditional monitoring relies on single camera or video monitoring means, and realizes monitoring and event judgment through image recognition and target detection algorithm, but is subject to light changes, shielding interference and complex background and many other factors, so its accuracy and real-time performance often cannot meet the high requirements of modern security and prevention.

[0003] The prior art still has many limitations in practical application. First, single video monitoring is often affected by environmental noise, light fluctuation, target shielding and other factors, resulting in misjudgment or missed judgment of abnormal events; even if non-visual sensors are introduced, lack of effective linkage mechanism and spatio-temporal data fusion means, it is also difficult to realize the accurate tracking of the target in all directions and continuously. Second, in the process of generating the target trajectory, the traditional monitoring system has a time delay and angle difference in data acquisition between each camera, which makes the continuous motion trajectory lack of smoothness, and it is difficult to perform reasonable interpolation processing when the coordinate data is continuously lost, which directly affects the training effect of the subsequent behavior recognition model. In addition, for monitoring data involving sensitive information of privacy, most existing schemes fail to start from the perspective of multi-angle cooperative shielding, ensuring data desensitization while maintaining the usability of information. At the same time, the monitoring system often fails to make full use of historical data to predict regional crowd density during non-event triggering period, so as to realize intelligent optimization of energy consumption and dynamic adjustment of camera. SUMMARY

[0004] In view of the above existing problems, the present application is proposed.

[0005] Therefore, the present application provides a camera cooperative monitoring method based on multi-modal perception, which can solve the problems of insufficient privacy protection, light, angle limitation, data loss and other problems in traditional monitoring, thereby improving the system response speed, tracking accuracy and security, effectively preventing privacy leakage and reducing energy consumption, and adapting to multi-scene applications in complex environments.

[0006] To solve the above technical problems, the application provides the following technical scheme, a camera cooperative monitoring method based on multi-modal perception, comprising: detecting an abnormal event signal in an environment through a non-vision sensor node, triggering camera activation and visual capture of a target area; a first camera that captures the target extracts the color block distribution and motion direction of the target, broadcasts to adjacent cameras through a local area network, realizes relay tracking based on feature matching, and generates a continuous motion trajectory; different cameras perform perspective coordinate mapping on the same privacy-sensitive target, negotiate to generate complementary shielding areas, and generate desensitized monitoring data; during a non-event triggering period, part of the camera power is turned off according to historical activity prediction, and the remaining cameras switch to wide-angle mode to cover the entire area; the cloud updates the global behavior recognition model based on the desensitized motion trajectory data using a federated learning framework.

[0007] As a preferred scheme of the camera cooperative monitoring method based on multi-modal perception, wherein: the non-vision sensor comprises a sound sensor and a vibration sensor, wherein:

[0008] The sound sensor extracts frequency spectrum features and matches them with abnormal voiceprint templates in a preset voiceprint library by collecting environmental sound wave signals.

[0009] The vibration sensor is deployed at the boundary of the monitoring area, and the vibration frequency and duration are detected by an accelerometer to determine climbing or impact events.

[0010] As a preferred scheme of the camera cooperative monitoring method based on multi-modal perception, wherein: the visual capture comprises synchronously activating the cameras within the sensor transmission range through a hardware trigger signal when the non-vision sensor detects an abnormal signal, and obtaining the video stream of the corresponding area.

[0011] The time difference between the abnormal signal and the object motion in the camera picture is calculated based on the speed of sound and light propagation, and when the time difference is less than a target value and the detected object displacement in the picture exceeds a target displacement, it is determined as an effective event trigger.

[0012] As a preferred scheme of the camera cooperative monitoring method based on multi-modal perception, wherein: the continuous motion trajectory generation comprises: the first camera that captures the target extracts the color block distribution and motion direction of the target, converts the target area pixels to HSV color space, and extracts the continuous area with a set saturation value as the color block; the motion direction angle is calculated through the block centroid offset between adjacent frames.

[0013] The HSV histogram of the color block and the motion direction angle are broadcast to adjacent cameras in the same local area network through the UDP protocol.

[0014] The two-dimensional coordinates of the target on each camera in the self-coordinate system and the time stamp are transmitted to an edge server;

[0015] The edge server aligns the data of each camera according to the time stamp, generates a smooth trajectory by using a cubic spline interpolation algorithm for the missing coordinate points between two continuous frames, and marks the section as an occlusion missing section when the number of continuously missing coordinate points in the motion trajectory is greater than the missing number of the target.

[0016] As a preferred scheme of the camera collaborative monitoring method based on multi-modal perception, the multi-view collaborative shielding includes: each camera performs view angle coordinate mapping on the same privacy-sensitive target, and converts the target pixel coordinates to the same world coordinate system through calibration parameters; and the shielding area projection under the view angle of each camera is calculated according to the three-dimensional position of the target in the world coordinate system.

[0017] The generation rule of the complementary shielding area is that the target area is uniformly meshed to generate sub-areas same in number as the cameras; at least one sub-area is allocated to each camera, and a dynamic mosaic algorithm is used for shielding, and the mosaic granularity is adaptively adjusted according to the moving speed of the target.

[0018] As a preferred scheme of the camera collaborative monitoring method based on multi-modal perception, the non-event triggering period includes: training a region activity model according to historical monitoring data to predict the pedestrian density of each period in a future preset period.

[0019] When the predicted pedestrian density is lower than a preset activity threshold, the power supply of a first preset proportion of cameras in the target area is turned off, the first preset proportion is calculated according to the area of the region and the total number of cameras, and the remaining cameras are switched to a wide-angle mode to cover the whole region.

[0020] As a preferred scheme of the camera collaborative monitoring method based on multi-modal perception, the updating of the behavior recognition model includes: transmitting the desensitized motion trajectory data to a cloud server, extracting the motion direction angle and speed change sequence in the motion trajectory as training features; using a federated learning framework to perform gradient descent optimization on the global behavior recognition model, and updating the weight parameters of the full connection layer.

[0021] The updated global behavior recognition model parameters are encrypted and distributed to each edge node through the HTTP / 2 protocol.

[0022] As a preferred scheme of the camera collaborative monitoring system based on multi-modal perception, the system includes an anomaly detection module, a target tracking module, a privacy protection module, and an energy-saving updating module.

[0023] The abnormality detection module utilizes a non-vision sensor to monitor the environment in the monitoring area in real time, generates an event trigger signal, and synchronously activates the camera in the sensor transmission range and starts visual capture through a hardware trigger instruction;

[0024] The target tracking module realizes the cooperative operation between the cameras after the event trigger, and continuously and seamlessly tracks and collects data of the target;

[0025] The privacy protection module performs shielding processing on the part containing privacy sensitive information in the captured target in the monitoring process, so that the monitoring data meets the safety requirement while not infringing on personal privacy.

[0026] The energy-saving updating module realizes the energy consumption optimization of the camera through intelligent control during the non-abnormal event trigger period, and uses the motion trajectory data obtained after desensitization to update and optimize the global behavior recognition model online by using an advanced federated learning framework.

[0027] A computer device comprises a memory and a processor, the memory stores a computer program, and the processor realizes the steps of the camera cooperative monitoring method based on multi-modal perception when executing the computer program.

[0028] A computer readable storage medium stores a computer program, and the computer program realizes the steps of the camera cooperative monitoring method based on multi-modal perception when executed by a processor.

[0029] The present application has the following advantages: in the present application, the triggering of abnormal events and the activation of cameras are completed in advance by non-vision sensors, which reduces the influence of illumination and viewing angle limitations on event detection, and improves the response speed and monitoring coverage of the system. By extracting the color block distribution and motion direction of the target, and using the UDP protocol to realize data broadcasting and edge server data alignment in the local area network, a smooth and continuous target motion trajectory can be generated, and the data loss problem can be solved by using cubic spline interpolation, thereby ensuring the tracking accuracy. The multi-view cooperative shielding mechanism adopts a dynamic mosaic algorithm and uniform grid division, realizes complementary shielding of the privacy sensitive area, and effectively prevents the risk of privacy infringement caused by monitoring data leakage. At the same time, based on the area activity prediction of historical monitoring data, the working load of part of the cameras can be intelligently reduced during the low passenger flow period, achieving the effect of energy saving and consumption reduction. The present application provides an efficient, intelligent, safe and energy-saving comprehensive solution for the monitoring system, and has wide application prospect and obvious technical advantages. BRIEF DESCRIPTION OF DRAWINGS

[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.

[0031] Figure 1 The flow chart of the camera cooperative monitoring method based on multi-modal perception provided by an embodiment of the present application is shown. DETAILED DESCRIPTION

[0032] In order to make the above-mentioned objects, features and advantages of the present application more apparent and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort should be within the scope of protection of the present application.

[0033] Embodiment 1, refer to Figure 1 For the first embodiment of the present application, the embodiment provides a camera cooperative monitoring method based on multi-modal perception, which comprises:

[0034] S1: detecting an abnormal event signal in the environment by a non-vision sensor node, triggering camera activation and visual capture of the target area.

[0035] Further, the non-vision sensor node comprises a sound sensor and a vibration sensor, wherein:

[0036] The sound sensor extracts the frequency spectrum features by collecting the environmental sound wave signal and matches them with the abnormal voiceprint templates in the preset voiceprint library, and the abnormal voiceprint templates include the frequency band features of the glass breaking sound spectrum peak value in the range of 2kHz-4kHz;

[0037] The vibration sensor is deployed at an interval of 5 meters on the boundary of the monitoring area, and a three-axis MEMS accelerometer is used with a sampling frequency of 1kHz.

[0038] The climbing or impact event with vibration frequency exceeding 50Hz and duration greater than 0.3 seconds is detected by the accelerometer.

[0039] It should be noted that when the non-vision sensor detects an abnormal signal, the camera within 3 meters of the distance sensor is synchronously activated by a hardware trigger signal to obtain the video stream of the corresponding area, the straight line distance L(m) between the sensor and the nearest camera is calculated, and the theoretical time difference is calculated based on the sound speed 340m / s and the light speed 3x10^8m / s:

[0040]

[0041] wherein, Δt is time difference, v s is sound speed, v1 is light speed; if the time difference between the starting moment of the object movement in the actual picture and the trigger signal is within the interval [Δt-δ, Δt+δ] (δ is an error tolerance value), then the displacement amount verification is performed;

[0042] Further, the target region is extracted by background difference method, in the initialization stage, 30 frames of images are continuously collected when there is no target activity, an initial background is established, each pixel point is described by three Gaussian distributions, and the weight coefficient is initialized as [a, b, c] (for example, the weight coefficient can be [0.3, 0.5, 0.2]); the background is updated once every 5 seconds, and the weight of the corresponding Gaussian distribution of the pixel point which does not change for more than 10 frames is increased, and the weight of other distributions is attenuated to adapt to the gradual change of light.

[0043] After the difference between the current frame and the background model, a double threshold method is used for processing: the high threshold (greater than 50 gray level) region is directly determined as a motion region; the low threshold (20-50 gray level) region needs to satisfy the high threshold point in the adjacent 3*3 window to be retained; the morphological closing operation (3*3 elliptical kernel) is performed on the segmented connected region to eliminate small noise.

[0044] S2: the camera performs relay tracking based on target feature matching to generate a continuous motion trajectory.

[0045] Further, the first camera that captures the target extracts the color block distribution and the motion direction of the target, specifically including: converting the target region pixels to HSV color space;

[0046] The saturation value range is [0, 1], the higher the value, the brighter the color, and the less affected by light changes; the low saturation region (lower than 0.4) usually corresponds to shadows, reflections or gray backgrounds, which is easy to be disturbed; the high saturation region (higher than 0.8) usually corresponds to artificial markers (such as warning signs), which is easy to be misjudged. In the present application, the continuous region with a saturation value higher than 0.6 is extracted as a color block, the shadow interference is filtered, the color block feature is more focused on the real target, and the hardware computing power constraint is adapted to meet the real-time monitoring demand.

[0047] The motion direction angle is calculated by the block centroid offset amount between adjacent frames, the gradient change significant corner points in the target region are extracted, and a minimum feature value threshold is set; for the adjacent two frames of images, an 8*8 pixel search window is established at the corner point position, the optical flow vector is solved by iterative weighted least squares method, and the following formula is satisfied:

[0048] ΣW 2 ·(I x u+I y v+It ) 2 → min

[0049] where W is the Gaussian weight matrix, I x is the horizontal gradient of the pixel, I y is the vertical gradient of the pixel, u, v are the optical flow vector, I t is the temporal gradient; only keep the vectors whose SSD (Sum of Squared Difference) is less than 5 pixels;

[0050] Direction clustering of the retained displacement vectors: divide 0°-360° into 12 30° intervals; count the number of vectors in each interval, and select the median value in the interval with a proportion of more than 60% as the main direction; the remaining part is determined as the target turning, and the average direction of the longest continuous trajectory segment is taken.

[0051] Broadcast the HSV histogram and motion direction angle of the color block to the adjacent cameras in the same LAN through the UDP protocol. Each camera uploads the two-dimensional coordinates of the target in its own coordinate system and the timestamp to the edge server at a frequency of 10Hz, and the data format is (camera ID, timestamp, image coordinate system);

[0052] The edge server aligns the data of each camera according to the timestamp, generates a smooth trajectory for the missing coordinate points between two consecutive frames using a cubic spline interpolation algorithm, and takes the first two known trajectory points (P0, P1) and the last two points (P2, P3) as control points; if the missing segment is located at the beginning or end of the trajectory, virtual control points are generated by mirroring adjacent points;

[0053] Input the control points into the Catmull-Rom spline generator, set the parameterization step, generate a smooth path, and adjust the speed consistency of the interpolation points. Calculate the average speed v avg of the known trajectory segment; limit the distance between interpolation points to be no more than ±20% of v avg × time interval.

[0054] If more than 5 consecutive trajectory points are missing, mark the missing time period [t1, t2] and send a retake instruction to the backup camera deployed within a radius of 3 meters in the missing area.

[0055] It should be noted that the target position is predicted according to the missing time period [t1, t2] and the latest trajectory point, and a spherical search area is generated, with the center of the sphere being the last known coordinate and the radius being: the maximum speed of the trajectory point × (t2-t1);

[0056] According to the time range and estimated position of the occlusion missing segment, the gimbal pitch angle is adjusted to 30°-60°, and the missing shot is taken in wide-angle mode; the motion target detection is performed on the missing shot, and if a target meeting the original characteristics is detected, its coordinates are added to the trajectory data and connected to the original trajectory through a B-spline curve.

[0057] S3: Multi-view collaborative masking of privacy sensitive information in the tracked target is performed to generate desensitized monitoring data and upload.

[0058] Further, multiple cameras perform view angle coordinate mapping on the same privacy sensitive target, and each camera converts the target pixel coordinates to the same world coordinate system (X, Y, Z) through calibration parameters, and the conversion formula is:

[0059]

[0060] where (c x ,c y ) is the camera optical center coordinates, (x, y) is the target pixel coordinates, f x , f y is the focal length corresponding to the horizontal and vertical coordinates, and Z is obtained through binocular camera parallax calculation.

[0061] According to the three-dimensional position of the tracked target in the world coordinate system, the projection of the masking area under the view angle of each camera is calculated; each camera masks 50% of the preset area of the target projection area, and the spatial positions of the masking areas of the cameras do not overlap with each other;

[0062] The target privacy area (face / vehicle license plate) is subjected to semantic segmentation to obtain its minimum bounding rectangle, and the rectangle is divided into m sub-areas (m = the number of cameras participating in masking); each camera is assigned a sub-area, and the masking algorithm uses dynamic mosaic, and the size of the mosaic block is adjusted according to the target moving speed v d Adaptive adjustment:

[0063]

[0064] Further, the masked video stream is stripped of EXIF metadata, and the motion trajectory data is stored separately from the video, and the motion trajectory data field includes (timestamp, world coordinates (X, Y), speed v d , motion direction); for the part of the privacy area in the picture of any camera that is not masked, it is checked whether there is any other camera that masks this part, and if there is, it is allowed to remain, otherwise, masking instructions are added to the camera;

[0065] For the masked target, the aspect ratio of its bounding rectangle (used for vehicle type / human body judgment) and the motion vector direction are retained, and the color histogram features are deleted.

[0066] S4: Optimize camera energy consumption during non-event triggering period, and update behavior recognition model based on desensitized data.

[0067] Further, during the non-event triggering period, a region activity model is trained according to historical monitoring data to predict the people flow density in each period in the future preset period (which can be configured as 12 / 24 / 48 hours); the activity model adopts a prediction model based on historical data training time series, the input is date type (workday / holiday), weather data, and the output is people flow in each region in the future 12 / 24 / 48 hours; the prediction model adopts LSTM network, and the number of hidden layer units is 64.

[0068] When the predicted people flow density is lower than the preset activity threshold (the people flow density on weekdays is the 20th percentile of the historical people flow density; the people flow density on holidays is the 80th percentile of the historical people flow density), the power of the first preset proportion of cameras in the target region is turned off, and the first preset proportion R is calculated according to the area S and the total number N of cameras:

[0069]

[0070] Wherein, α is the effective coverage area of a single camera; the remaining cameras switch to 4K ultra-wide angle mode, the horizontal field of view is expanded to 150°, and the frame rate is reduced to 10fps; the cameras with power off wake up once every 30 minutes and work for 2 minutes to update the background model.

[0071] Further, after the cloud server receives the desensitized motion trajectory data, the motion direction angle and speed change sequence in the motion trajectory are extracted as training features to construct a global behavior recognition model; using a federated learning framework, the global behavior recognition model is optimized by gradient descent to update the weight parameters of the fully connected layer, each edge node locally stores the desensitized motion trajectory dataset Di, initializes the global behavior recognition model, removes the original last fully connected layer (original output dimension 1000), and adds a new fully connected layer structure:

[0072] Linear(512→128)→ReLU→Linear(128→3)→Softmax

[0073] Wherein, the output 3-class probability distribution is: normal walking (Class 0), running (Class 1), and staying (Class 2); the initial model weight of each node is synchronized by downloading the initial weight file from the central server, and after loading, all parameters except the newly added fully connected layer are frozen.

[0074] The updated global behavior recognition model parameters are encrypted and then distributed to each edge node, using the AES-256 encryption algorithm, in CBC mode, with an initialization vector (IV) length of 128 bits and a PKCS7 padding mode; the weight file is serialized and then encrypted to generate a ciphertext data block (each 16 bytes as a group); after receiving the data packet, the edge node decrypts and checks CRC32, and if the check passes, the new weight is loaded into the memory; a double buffering mechanism is used, and the copy of the global behavior recognition model before the update is retained until the updated global behavior recognition model completes 10 consecutive inferences without any abnormalities, and then the updated global behavior recognition model is switched to.

[0075] Embodiment 2, the second embodiment of the present application, which is different from the previous embodiment is:

[0076] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts of the present application that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device) to execute all or part of the steps of the method described in the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0077] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered a list of executable instructions for implementing logic functions, and can be specifically embodied in any computer-readable medium for use by an instruction execution system, apparatus or device, such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, apparatus or device and execute the instructions, or in conjunction with these instruction execution systems, apparatus or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate or transport programs for use by an instruction execution system, apparatus or device, or in conjunction with these instruction execution systems, apparatus or devices.

[0078] More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection (electronic) having one or more wires, a portable computer diskette (magnetic), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, and a portable compact disc read-only memory (CDROM). Additionally, the computer readable medium can be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example via an optical scanner, then compiled, interpreted or otherwise processed in a suitable manner, if necessary, to generate an electronically readable version of the program, which can then be stored in the computer memory.

[0079] It should be understood that aspects of the application can be implemented in hardware, software, firmware or combinations thereof. In the above embodiments, various steps or methods can be implemented in software or firmware that is stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any of the following technologies, known in the art, or their combinations can be used: discrete logic circuitry having logic gates for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc.

[0080] Embodiment 3, as an embodiment of the present application, provides a camera cooperative monitoring system based on multi-modal perception, including an anomaly detection module, a target tracking module, a privacy protection module, and an energy saving update module.

[0081] The anomaly detection module uses non-vision sensors to monitor the environment in the monitoring area in real time, generates an event trigger signal, and synchronously activates the cameras within the sensor transmission range and starts visual capture through a hardware trigger instruction.

[0082] The target tracking module, after the event trigger, realizes the cooperative work between the cameras, and continuously and seamlessly tracks and collects data of the target.

[0083] The privacy protection module performs masking processing on the part of the target captured in the monitoring process that contains privacy sensitive information, so that the monitoring data meets the security requirements while not infringing on personal privacy.

[0084] The energy saving update module, during the non-anomaly event triggering period, realizes the optimal energy consumption of the cameras through intelligent control; uses the motion trajectory data obtained after desensitization, and uses an advanced federated learning framework to optimize and update the global behavior recognition model online.

[0085] It should be noted that the above examples are only used to illustrate the technical solutions of the present application but not to limit the present application. Although the present application is described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or equivalently replaced, without departing from the spirit and scope of the technical solutions of the present application, and all of them should be covered in the scope of the claims of the present application.

Claims

1. A method for camera collaborative monitoring based on multi-modal perception, characterized in that: Comprising, Detecting an abnormal event signal in the environment through a non-vision sensor node, triggering camera activation and visual capture of the target area; The first camera to capture the target extracts the color block distribution and motion direction of the target, broadcasts it to adjacent cameras through the local area network, and realizes relay tracking based on feature matching to generate a continuous motion trajectory; The generation of the continuous motion trajectory includes: the first camera to capture the target extracts the color block distribution and motion direction of the target, converts the target area pixels to the HSV color space, extracts the continuous area with a set saturation value as the color block; calculate the motion direction angle through the block centroid offset between adjacent frames, extract the corner points with significant gradient changes in the target area, and set a minimum feature value threshold; for the adjacent two frames of images, an 8x8 pixel search window is established at the corner point position, and the optical flow vector is solved by iterative weighted least squares method, which satisfies: wherein, is a Gaussian weight matrix, is a horizontal gradient of the pixel point, is a vertical gradient of the pixel point, is an optical flow vector, is a temporal gradient; only vectors with a sum of squared differences (SSD) of less than 5 pixels between the previous and the current frame are kept; Direction clustering of the retained displacement vector: divide 0°-360° into 12 30° intervals; count the number of vectors in each interval, and select the median value in the interval with more than 60% of the proportion as the main direction; the remaining part is determined as the target turning, and the average direction of the longest continuous trajectory segment is taken; Broadcast the HSV histogram and motion direction angle of the color block to adjacent cameras in the same local area network through the UDP protocol, and upload the two-dimensional coordinates and time stamps of the target in the coordinate system of each camera to the edge server at a frequency of 10Hz, with the data format being camera ID, timestamp, image coordinate system; The edge server aligns the data of each camera according to the timestamp, generates a smooth trajectory using the cubic spline interpolation algorithm for the missing coordinate points between two consecutive frames, and takes the first two known trajectory points (P0, P1) and the last two points (P2, P3) as control points; if the missing segment is located at the beginning or end of the trajectory, virtual control points are generated by mirroring adjacent points; The control points are input into a Catmull-Rom spline generator, a parameterized step length is set, a smooth path is generated, and the interpolation points are adjusted for speed consistency, and the average speed of the known trajectory segment is calculated ; the distance between interpolation points is limited to no more than ±20% of the time interval If more than 5 consecutive track points are lost, mark the missing time period Send a re-shoot instruction to the backup camera deployed within 3 meters of the missing area radius According to the missing time period And the latest trajectory point predicts the target position, generates a spherical search area, the spherical center is the last known coordinate, and the radius is: the maximum speed of the trajectory point × (t - t0) ); According to the time range and estimated position of the occluded missing segment, adjust the pan-tilt pitch angle to 30°-60°, and shoot the supplementary footage in wide-angle mode; perform motion target detection on the supplementary footage, and if a target that meets the original characteristics is detected, add its coordinates to the trajectory data and connect it to the original trajectory through a B-spline curve; Different cameras map the same privacy-sensitive target to different perspectives, negotiate to generate complementary masking areas, and generate desensitized monitoring data; Multi-perspective collaborative masking includes mapping the same privacy-sensitive target to different perspectives by each camera, and converting the target pixel coordinates to the same world coordinate system through calibration parameters; According to the three-dimensional position of the tracked target in the world coordinate system, calculate the projection of the masking area under the perspective of each camera; each camera masks 50% of the pre-set area in the target projection area, and the spatial positions of the masking areas of each camera do not overlap; Semantic segmentation is performed on the target privacy area to obtain a minimum bounding rectangle, and the current rectangle is divided into m sub-areas, m=number of cameras participating in shielding; each camera is assigned a sub-area, and a dynamic mosaic is used in the shielding algorithm, and the size of the mosaic block is adjusted according to the moving speed of the target Adaptive adjustment: The video stream after being shielded is stripped of EXIF metadata, and motion trajectory data is stored separately from the video. The motion trajectory data field includes timestamp, world coordinate , speed , and motion direction ; for the part of the privacy area in any camera picture that is not shielded, it is checked whether there is another camera shielding this part, and if so, the part is allowed to be retained, otherwise, shielding instructions are added to this camera; For the masked target, retain its aspect ratio and motion vector direction, and delete the color histogram features; The generation rule of the complementary shielding area is that the target area is uniformly meshed to generate the same number of sub-areas as the number of cameras; at least one sub-area is allocated to each camera, and a dynamic mosaic algorithm is used for shielding, and the mosaic granularity is adaptively adjusted according to the moving speed of the target; During the non-event triggering period, the power of part of the cameras is turned off according to the historical activity prediction, and the remaining cameras switch to the wide-angle mode to cover the whole area; the cloud updates the global behavior recognition model based on the desensitized motion trajectory data using the federated learning framework. 2.The multi-modal perception based camera collaborative monitoring method of claim 1, wherein: The non-vision sensor includes a sound sensor and a vibration sensor, wherein: The sound sensor collects environmental sound signals, extracts frequency spectrum features, and matches them with abnormal voiceprint templates in a preset voiceprint library; The vibration sensor is deployed at the boundary of the monitoring area, and the accelerometer detects the vibration frequency and duration to determine the climbing or impact event. 3.The multi-modal perception based camera collaborative monitoring method of claim 2, wherein: The visual capture includes synchronously activating the cameras within the transmission range of the sensor through a hardware trigger signal when the non-vision sensor detects an abnormal signal, and obtaining the video stream of the corresponding area; The time difference between the abnormal signal and the object motion in the camera picture is calculated based on the sound and light propagation speed, and when the time difference is less than a target value and the displacement of the object in the picture exceeds a target displacement, it is determined as an effective event trigger. 4.The multi-modal perception based camera collaborative monitoring method of claim 3, wherein: The non-event triggering period includes training a regional activity model based on historical monitoring data to predict the pedestrian density of each period in the future within a preset period; When the predicted pedestrian density is lower than a preset activity threshold, the power of a first preset proportion of cameras in the target area is turned off, and the remaining cameras switch to the wide-angle mode to cover the whole area. 5.The multi-modal perception based camera collaborative monitoring method of claim 4, wherein: The updated behavior recognition model includes transmitting the desensitized motion trajectory data to the cloud server, extracting the motion direction angle and speed change sequence in the motion trajectory as training features; Using the federated learning framework, the global behavior recognition model is optimized by gradient descent, and the full connection layer weight parameters are updated; The updated model parameters are encrypted and distributed to each edge node through the HTTP / 2 protocol.

6. A system employing the multi-modal perception based camera collaborative monitoring method according to any one of claims 1-5, characterized in that: It includes an abnormality detection module, a target tracking module, a privacy protection module, and an energy-saving update module. The abnormality detection module uses the non-vision sensor to monitor the environment in the monitoring area in real time, generates an event trigger signal, and synchronously activates the cameras within the transmission range of the sensor through a hardware trigger instruction and starts visual capture; The target tracking module cooperates with the cameras to continuously and seamlessly track and collect data of the target after the event is triggered; The privacy protection module masks the part of the target containing private sensitive information captured during monitoring, so that the monitoring data meets the security requirements while not infringing on personal privacy; The energy-saving update module optimizes the energy consumption of the cameras through intelligent control during the non-abnormal event triggering period, and uses the desensitized motion trajectory data to update the global behavior recognition model using the advanced federated learning framework. 7.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-6 when the computer program is executed by the processor. The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 5.

8. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 5.

Citation Information

Patent Citations

  • A surveillance system, a surveillance camera and an image processing method

    CN105898208A

  • Intelligent opening and closing apparatus and method of security and protection camera based on channel state information

    CN107277452A

  • Image processing method and system for intelligent security and protection monitoring

    CN118887622A

  • Intelligent security monitoring and responding method for power generation enterprise

    CN119783043A