Multi-camera behavior correlation analysis method and system based on adaptive illumination enhancement

By constructing a pixel-level illumination degradation model and adaptive illumination enhancement technology, the problems of target recognition and identity consistency caused by illumination changes and perspective changes in multi-camera video surveillance systems are solved, achieving high-precision behavior analysis and anomaly detection, and enhancing the practicality and reliability of the system.

CN122116228APending Publication Date: 2026-05-29HANYOU TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANYOU TECH CO LTD
Filing Date
2026-02-02
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing multi-camera video surveillance systems have limitations in handling changes in lighting, target recognition, and cross-view identity unification, resulting in insufficient target detection and tracking accuracy, and difficulties in unifying identities across cameras, which affects the accuracy of behavior analysis and anomaly detection performance.

Method used

By constructing a pixel-level illumination degradation model to generate illumination-invariant reflective substrate images, combined with adaptive illumination enhancement and multi-scale gradient noise suppression, spatial-motion joint features are extracted to construct multi-view spatiotemporal maps, achieving cross-camera identity unification. Furthermore, structured anomaly reports are generated through three rounds of weighted graph propagation and anomaly detection, driving camera view scheduling and audio-visual warnings.

Benefits of technology

It improves the accuracy of target recognition and cross-camera individual movement trajectory tracking, significantly enhancing the practicality and reliability of the multi-camera behavior correlation analysis system, enabling timely detection of abnormal behavior and providing reliable early warning support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122116228A_ABST
    Figure CN122116228A_ABST
Patent Text Reader

Abstract

The application discloses a multi-camera behavior correlation analysis method and system based on adaptive light enhancement, constructs a pixel-level light degradation model and generates a light-invariant reflection base image, effectively compensates for the problem of increased target recognition difficulty caused by light changes, improves the target recognition accuracy, extracts spatial-motion joint features through a shared encoder, combines the motion trajectory and spatial projection relationship of the target instance to construct a multi-view space-time graph, realizes cross-camera identity unification, significantly improves the accuracy of individual moving track tracking between different cameras, and through analysis of the reconstructed three-dimensional centroid track and mask deformation energy, can accurately identify stable cooperative clusters and extract dominant behavior patterns, effectively solves the technical problems caused by light changes and view transformation, and greatly enhances the practicability and reliability of the multi-camera behavior correlation analysis system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video surveillance technology, specifically to a method and system for multi-camera behavior correlation analysis based on adaptive illumination enhancement. Background Technology

[0002] In modern video surveillance systems, multi-camera networks are widely used to monitor the security of public spaces and critical infrastructure. However, existing multi-camera behavior analysis methods have limitations in handling changes in lighting, target recognition, and cross-view identity verification. Specifically, the image quality varies significantly under different lighting conditions, which not only affects the accurate identification of targets but also hinders subsequent behavior analysis. Traditional methods often rely on manually adjusting parameters or using fixed enhancement algorithms to address lighting issues, but these methods are ill-suited to complex and ever-changing real-world environments, resulting in insufficient accuracy in target detection and tracking.

[0003] Furthermore, existing technologies typically rely on physical feature matching when handling identity verification across multiple cameras. However, due to factors such as changing viewpoints and occlusion, using physical features alone for matching is not ideal and can easily lead to identity confusion. In such cases, it becomes difficult to accurately track an individual's movement trajectory across multiple cameras, thus affecting the performance of behavior pattern-based anomaly detection and early warning systems. Summary of the Invention

[0004] This invention aims to provide a multi-camera behavior association analysis method and system based on adaptive illumination enhancement, which solves the problems of inaccurate target recognition caused by changes in illumination conditions and difficulties in unifying cross-camera identities due to factors such as perspective changes and occlusion in the existing technology, thereby improving the overall performance of the multi-camera behavior association analysis system.

[0005] To achieve the above objectives, the technical solution adopted in this invention is as follows: a multi-camera behavior association analysis method based on adaptive illumination enhancement, comprising: acquiring the original grayscale image and constructing a pixel-level illumination degradation model to generate an illumination-invariant reflective substrate image; performing adaptive illumination enhancement using the illumination-invariant reflective substrate image, and generating an enhanced image by combining multi-scale gradient noise suppression; extracting spatial-motion joint features from the enhanced image through a shared encoder, and simultaneously generating a pixel-level segmentation mask and motion vector field to form a target instance set with motion attributes; constructing a multi-view spatiotemporal map based on the motion trajectory and spatial projection relationship of the target instances, and performing three rounds of weighted summation. Graph propagation enables cross-camera identity unification and outputs cross-view identity allocation results; it aggregates observation data from various viewpoints according to identity labels, reconstructs the 3D centroid trajectory and analyzes mask deformation energy, generating a spatiotemporal action primitive sequence; it calculates the spatiotemporal proximity and motion coupling between action primitives of different identities, identifies stable cooperative clusters and extracts dominant behavior patterns, and outputs cooperative behavior labels; it constructs a spatiotemporal reference envelope for cooperative behavior labels, performs anomaly discrimination through three-dimensional deviation measurement of space, rate and rhythm, and generates structured anomaly reports; it drives camera viewpoint scheduling and audio-visual warning units according to the type and severity of anomaly reports, and executes multi-granularity linkage response.

[0006] On the other hand, this invention proposes a multi-camera behavior association analysis system based on adaptive illumination enhancement, comprising: an illumination modeling unit for acquiring the original grayscale image and constructing a pixel-level illumination degradation model to generate an illumination-invariant reflective substrate image; an image enhancement unit for performing adaptive illumination enhancement using the illumination-invariant reflective substrate image and generating an enhanced image by combining multi-scale gradient noise suppression; a target analysis unit for extracting spatial-motion joint features from the enhanced image through a shared encoder and simultaneously generating a pixel-level segmentation mask and motion vector field; an identity association unit for constructing a multi-view spatiotemporal map based on the motion trajectory and spatial projection relationship of the target instance, and achieving cross-camera identity unification through weighted graph propagation; a behavior parsing unit for aggregating observation data from each viewpoint according to identity labels, reconstructing the three-dimensional centroid trajectory and analyzing mask deformation energy; an interaction recognition unit for calculating the spatiotemporal proximity and motion coupling between different identity action primitives and identifying stable collaborative behavior patterns; an anomaly discrimination unit for constructing a spatiotemporal reference envelope for collaborative behavior and performing abnormal behavior recognition through multi-dimensional deviation measurement; and a response execution unit for driving camera viewpoint scheduling and audio-visual warning units according to the severity of the anomaly, and executing multi-granularity linkage response.

[0007] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0008] This invention effectively compensates for the increased difficulty in target recognition caused by illumination changes by constructing a pixel-level illumination degradation model and generating illumination-invariant reflective substrate images, thereby improving the accuracy of target recognition. By extracting spatial-motion joint features through a shared encoder and combining the motion trajectory of target instances with the spatial projection relationship to construct multi-view spatiotemporal maps, it achieves cross-camera identity unification, significantly improving the accuracy of tracking individual movement trajectories across different cameras. Through analysis of the reconstructed 3D centroid trajectory and mask deformation energy, it can accurately identify stable cooperative clusters and extract dominant behavioral patterns, providing reliable technical support for the timely detection and early warning of abnormal behavior. It effectively solves the technical challenges caused by illumination changes and perspective transformations, greatly enhancing the practicality and reliability of the multi-camera behavior correlation analysis system. Attached Figure Description

[0009] Figure 1 This is a flowchart of the multi-camera behavior correlation analysis method based on adaptive illumination enhancement of the present invention; Figure 2 This is a block diagram of the multi-camera behavior correlation analysis system based on adaptive illumination enhancement according to the present invention. Detailed Implementation

[0010] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.

[0011] like Figure 1 As shown, this invention proposes a multi-camera behavior correlation analysis method based on adaptive illumination enhancement, aiming to embed the physical laws of light propagation into the image representation generation process, and drive the structured alignment of multi-view observations with spatiotemporal consistency as a constraint. Specifically, it includes the following steps: The process involves acquiring the original grayscale image and constructing a pixel-level illumination degradation model to generate an illumination-invariant reflective substrate image. The construction of the pixel-level illumination degradation model includes: performing sensor response linearization correction on the original grayscale image to generate a linear domain image; calculating the maximum illumination intensity distribution within a local window based on the linear domain image to construct a spatially varying illumination field; using the spatially varying illumination field as a constraint, jointly solving the Gauss-Poisson equations to obtain a reflectance map and a shadow mask; and then applying gamma correction and dynamic range compression to the reflectance map to generate an illumination-invariant reflective substrate image.

[0012] Adaptive illumination enhancement is performed using an illumination-invariant reflective substrate image, combined with multi-scale gradient noise suppression to generate an enhanced image. The adaptive illumination enhancement process includes: calculating an enhancement objective function with reference to the brightness distribution of the illumination-invariant reflective substrate image; combining the enhancement objective function with the local gradient magnitude of the linear domain image to construct a brightness transfer map field; applying wavelet domain threshold filtering to the linear domain image to preserve high-frequency details and suppress noise; and applying the brightness transfer map field to the filtered image to generate an enhanced image with preserved edge details.

[0013] A shared encoder is used to extract spatial-motion joint features from the enhanced image, and pixel-level segmentation masks and motion vector fields are generated simultaneously to form a set of target instances with motion attributes. Specifically, the enhanced image is input into a shared encoder with a pyramid structure to extract four levels of spatial-motion joint feature maps; the joint feature maps are processed by an instance segmentation head to generate pixel-level instance masks and instance confidence scores; the same joint feature maps are processed by a motion regression head to generate horizontal and vertical motion vector components; the instance masks and motion vector components are fused, and the centroid coordinates and average motion vector of each instance are calculated to form a set of target instances with motion attributes.

[0014] A multi-view spatiotemporal graph is constructed based on the motion trajectory and spatial projection relationship of target instances. Cross-camera identity unification is achieved through three rounds of weighted graph propagation, and cross-view identity allocation results are output. Specifically, this includes: extracting motion trajectory segments and appearance feature vectors of each instance from the set of target instances with motion attributes; constructing edge connections of the multi-view spatiotemporal graph based on the similarity of motion trajectory segments and the overlapping area of ​​camera fields of view; assigning weights to each edge, with the weights determined by both motion similarity and spatial projection distance; performing three rounds of weighted message passing on the multi-view spatiotemporal graph, aggregating the identity labels of neighboring nodes, and outputting cross-view identity allocation results.

[0015] Aggregating observation data from various perspectives based on identity tags, reconstructing the 3D centroid trajectory, and analyzing mask deformation energy to generate a spatiotemporal action primitive sequence; specifically including: collecting observation sequences of the same identity across all cameras based on cross-view identity allocation results; using the short-time window line-of-sight intersection estimation method to fuse the centroid positions from multiple perspectives to reconstruct the 3D centroid trajectory; calculating the derivatives of mask contour length and area with respect to time to generate deformation energy curves and detect energy abrupt change points; aligning the extreme points of 3D trajectory curvature with deformation energy abrupt change points to segment action primitives with durations between 400 milliseconds and 6 seconds.

[0016] The study calculates the spatiotemporal proximity and motion coupling between different identity action primitives, identifies stable cooperative clusters, extracts dominant behavior patterns, and outputs cooperative behavior labels. Specifically, it includes: calculating the three-dimensional centroid distance integral of different identity action primitives in the time overlap region to obtain the spatiotemporal proximity; analyzing the stability of the cosine average of the angle between the motion direction vectors and the velocity ratio during the overlap period to calculate the motion coupling; constructing an identity interaction graph based on the spatiotemporal proximity and motion coupling, and identifying cooperative clusters with high internal coupling through hierarchical clustering; calculating the average motion speed within the cooperative cluster, marking a stationary cooperative when the speed is below 0.3 m / s, marking a co-directional cooperative when the speed is above 0.3 m / s and the motion direction is consistent, and marking a close formation when the variance of the distance between members is less than 2.0 square meters.

[0017] A spatiotemporal reference envelope is constructed for the labels of collaborative behaviors. Anomaly detection is performed by measuring deviations in three dimensions: space, rate, and rhythm, and a structured anomaly report is generated. Specifically, this includes: calculating the three-dimensional convex hull of the member position for each time offset point of the collaborative cluster to form a spatiotemporal reference envelope; measuring the maximum spatial distance between the new action primitive and the reference envelope, the rate outside the envelope ratio, and the relative deviation in duration; dynamically updating the judgment threshold based on the cluster's historical deviation data, and marking it as an anomaly when all three deviations exceed the adaptive threshold simultaneously; and classifying it into detached individual anomalies, individual detachment anomalies, or collaborative structure disintegration based on the anomaly identity, the relationship with the collaborative cluster, and the deviation pattern.

[0018] Based on the type and severity of the anomaly report, the camera view scheduling and audio-visual warning units are driven to execute multi-granular linkage responses. Specifically, this includes: calculating the response level based on the anomaly type and spatial location; setting a level 1 response for anomalies involving stray individuals in non-key areas, a level 2 response for anomalies involving individuals detaching from the passageway, and a level 3 response for anomalies involving the collapse of a collaborative structure; selecting the first two available cameras that can cover the area and have a viewing angle greater than 0.65 for the spatial coverage area of ​​the calculation action primitives for level 2 and 3 anomalies, and generating zoom and tracking commands; activating the audio-visual warning units within a 15-meter range at the geometric center of the spatial coverage area for level 3 anomalies, setting an 85-decibel directional sound field and a 5-second visual warning; and associating all commands with the original anomaly data, execution time window, and spatial coordinates, and writing them into an auditable execution log containing an immutable timestamp and data hash value.

[0019] On the other hand, this embodiment also proposes a multi-camera behavior correlation analysis system based on adaptive illumination enhancement, such as... Figure 2 As shown, it includes: The system comprises the following units: a lighting modeling unit, which acquires the original grayscale image and constructs a pixel-level lighting degradation model to generate an illumination-invariant reflective substrate image; an image enhancement unit, which performs adaptive lighting enhancement using the illumination-invariant reflective substrate image and combines it with multi-scale gradient noise suppression to generate an enhanced image; a target analysis unit, which extracts spatial-motion joint features from the enhanced image using a shared encoder and simultaneously generates a pixel-level segmentation mask and motion vector field; an identity association unit, which constructs a multi-view spatiotemporal map based on the motion trajectory and spatial projection relationship of the target instance and achieves cross-camera identity unification through weighted graph propagation; a behavior parsing unit, which aggregates observation data from various perspectives according to identity tags, reconstructs the 3D centroid trajectory, and analyzes the mask deformation energy; an interaction recognition unit, which calculates the spatiotemporal proximity and motion coupling between different identity action primitives and identifies stable collaborative behavior patterns; an anomaly detection unit, which constructs a spatiotemporal reference envelope for collaborative behavior and performs abnormal behavior recognition through multi-dimensional deviation measurement; and a response execution unit, which drives camera view scheduling and audio-visual warning units based on the severity of the anomaly and executes multi-granularity linkage responses.

[0020] Furthermore, during execution, the aforementioned units are also used to implement other steps of a multi-camera behavior correlation analysis method based on adaptive illumination enhancement, specifically including: Step 1: Construct pixel-level illumination degradation modeling and initial reflection component estimation Step 1.1: Acquire the original grayscale image and perform sensor response linearization correction. Let the input image be ,in These represent the image height and width, respectively, with each pixel value output as an 8-bit integer quantization. Because CMOS sensors exhibit non-linear response characteristics in low-light conditions, the inverse response function obtained from device calibration is first used... right By performing mapping, we obtain the linear domain image. Here It is a monotonically increasing piecewise polynomial whose coefficients are obtained by fitting the uniform surface light source response curves of the factory calibration board at different exposure times; this mapping makes The magnitude of the value is approximately proportional to the incident photon flux per unit area, thus providing a physical dimensionally consistent input basis for subsequent lighting modeling.

[0021] Step 1.2: Establish the prior distribution of local illumination intensity and solve for the spatially varying illumination field. The result obtained in step 1.1 Based on this, define a sliding window. The center is located at the pixel The size is This is used to capture local lighting smoothness features. Assume any point in the scene... Light intensity at the location Satisfies Gaussian random field priors: ; among which the mean Standard deviation Therefore, for each pixel We can construct a maximum posterior estimate of its illumination intensity: ;in Represents the Laplace operator. This is an empirical adjustment factor used to suppress noise-induced noise. High-frequency oscillation, making It more closely resembles the continuous spatial variation trend of a real lighting field. The spatially varying illumination field required for subsequent degradation modeling is constructed, with each location value derived from the linearized image provided in step 1.1. It is deduced that... A deterministic space function.

[0022] Step 1.3: Jointly solve the reflectivity map and shadow mask to generate a dual-channel degradation decomposition result. Based on the results obtained in step 1.2 Introduce the classic illumination-reflection decoupling model: ;in It represents the reflectivity of an object's surface in the sense of wavelength integral, and its value is determined only by the material and is independent of the observation conditions; This is a shadow mask, and its value reflects whether the point is in an area blocked by direct sunlight. The additive noise term follows a zero-mean Gaussian distribution with a standard deviation of . It is proportional to the square root. The goal of this step is to simultaneously estimate... Therefore, we define the energy functional: ; in As regularization weights, constrain the reflectivity maps respectively. With shadow mask Spatial smoothness; This represents the gradient operator. The alternating direction optimization method is used to solve it: [Fixed] renew Then fix renew The algorithm converges after three iterations. The output consists of two images at the same resolution: a reflectance image and a reflection image. With shadow mask This result directly depends on step 1.2. Spatial distribution pattern -- if a certain area If the value is extremely low, then It will be forcibly raised to compensate for the missing product term, thereby avoiding reflectivity collapse; conversely, if higher and Still low, then Pressed down to zero, it is marked with a strong shadow. Therefore, yes and The unique definite solution under the combined effect constitutes the dual-channel degenerate decomposition result.

[0023] Step 1.4: Generate an illumination-invariant reflective substrate image and perform dynamic range normalization. Using the results obtained in step 1.3 Using it as an illumination-invariant substrate, a reflection substrate image is defined. However, due to limitations in solution accuracy and residual noise, The dynamic range is often concentrated in This range is unfavorable for subsequent network processing. Therefore, dynamic range normalization is performed on it: ;in After normalization This is a standard 8-bit integer image. It is worth noting that... Each pixel value is composed of It comes from a linear mapping, and It is uniquely determined entirely by the energy functional minimization solution in step 1.3, which is rooted in step 1.2. The spatial structure can ultimately be traced back to the linearized correction in step 1.1. .

[0024] Step 2: Perform adaptive illumination enhancement and noise suppression based on a reflective substrate. Step 2.1: Generate a pixel-wise brightness objective function based on the reflective substrate image. The result obtained from step 1.4 For input, define each pixel Expected brightness value This value should not be simply set to a fixed threshold, but should reflect the material's reflectivity and the human eye's adaptive response in low-light environments. Referring to the CIE luminance perception model, in low-light conditions, the human eye is most sensitive to mid-gray areas (reflectivity approximately 0.18), and less sensitive to pure black and pure white. Therefore, the following setting should be made: ; in For shape control parameters, Guarantee output limiting The interval corresponds to the brightness band that the human eye can most easily distinguish in a dark scene. This formula indicates that when... Approximately 128 (medium gray). Approaching 128, the highest recognition weight is given; when Significantly deviated from 128, Gradually saturate towards 64 or 192 to avoid applying excessive enhancement pressure to extremely dark or bright materials. This objective function is entirely derived from... Decision, and because Source Therefore Essentially, it's a mapping of material reflectivity into the perceptual domain, rather than a blind increase in original brightness. Therefore, yes The nonlinear order-preserving transformation forms the pixel-by-pixel brightness reference system for the subsequent enhancement process.

[0025] Step 2.2: Construct a luminance migration map field under local contrast preservation constraints The result obtained in step 2.1 Based on this, a mapping needs to be designed to make the original linear image... After transformation, it approximates Without disrupting the local structure, define a mapping function. Acting on Centered Neighborhood , requires for any ,have: This equation forces the relative gray-level difference within the mapped neighborhood to be strictly equal to the corresponding difference in the reflective substrate image, thereby locking in the local contrast topology. Combined with step 2.1... should approach The target can be deduced to be an affine form: ; where the coefficient satisfy: ; For all The median is used to suppress outlier interference. This design ensures that each pixel has its own dedicated mapping parameters, and these parameters are determined by... and The local differences are uniquely determined. Therefore, the mapped field The deterministic output, constrained by all three factors, has no free hyperparameters and is entirely inherited from the physical modeling chain of step one.

[0026] Step 2.3: Integrate multi-scale gradient-guided noise suppression paths If the mapping obtained in step 2.2 is directly applied... This significantly amplifies readout noise in dark areas. To suppress this effect, gradient domain collaborative filtering is introduced. The gradient magnitudes of the original image in the horizontal and vertical directions are defined as follows: ; ;in For unit displacement, calculate its weighted sum. This serves as an edge strength indicator. Subsequently, a spatial adaptive filtering kernel is constructed: ;in The first three terms are standard bilateral filter weights, and the last term... It is a key adjustment term: when When the factor is large (i.e., located at a strong edge), it approaches zero, significantly weakening the filtering strength and protecting edge sharpness; when When the factor approaches zero (i.e., in the smooth region), it approaches 1, enabling full-intensity filtering. The weight field is entirely determined by... Local gradient Drive, and Furthermore, stemming from the linearized image in step 1.1, the filter intensity distribution is naturally coupled with the original image quality. Ultimately, for each... Calculate the filtered value: ;Should Become the new input for the mapping in step 2.2, replacing the original. This allows noise to be suppressed before enhancement, and the degree of suppression is automatically adjusted according to the complexity of the local structure.

[0027] Step 2.4: Perform reflective substrate-guided brightness transfer and output the enhanced image. The result obtained in step 2.3 Substituting the mapping function defined in step 2.2, we obtain the final enhanced image: In this expression, As already done in step 2.2 and Solve together, The noise has already been reduced by gradient-guided filtering in step 2.3. Therefore, It is a deterministic combination of three inputs: it satisfies both Point brightness approximation And maintain the relative difference within the neighborhood. The input is consistent, and most noise components have already been removed from the input itself. To adapt to the input specifications of subsequent segmentation networks, [the following is done / adjusted / adjusted]. Perform truncation and normalization: ;in Values ​​outside the range are clamped to the boundary. At this point, This becomes the sole image input for target segmentation in step three. Each pixel value of this image is determined by... Guided generation, and This also stems from the physical degradation and decomposition of the original image in step one, therefore It is the first visual output of the entire physical guidance chain, carrying the dual guarantee of material invariance and perception rationality, laying a solid foundation for subsequent pixel-level target separation.

[0028] Step 3: Achieve illumination-robust target instance segmentation and motion vector field extraction Step 3.1: Construct a shared encoder and extract multi-level spatial-motion joint feature maps The result obtained from step 2, step 2.4 The input is fed into a seven-layer convolutional encoder. The kernel size of the first layer of this encoder is... The step size is 2, and the subsequent six layers are all... Convolutional layers with a stride of 1 are used, followed by batch normalization and ReLU activation at each layer. The encoding process does not introduce any skip connections or attention modules, ensuring that the feature evolution path is unique and parsable. Let the th... The layer output feature map is ,in Specifically, the seventh layer output... Size is , recorded as The feature map is simultaneously fed into two parallel decoding branches: one for instance segmentation and the other for motion vector regression. Since the two branches share all encoding parameters, The 256-dimensional vector at each spatial location naturally integrates the appearance texture, boundary orientation, and potential movement trend information of the target at that location.

[0029] Step 3.2: Generate an instance-aware pixel-level segmentation mask set based on the joint feature map The result obtained in step 3.1 Based on this, a segmentation and decoding branch is added. This branch consists of four sets of upsampling and convolution, with each upsampling factor being 2, and the final output is... Feature maps of the same size ,in The maximum number of instances is preset (32). For each position... This represents the probability distribution of a pixel belonging to each instance. To generate hard segmentation results, [the following is done]: Perform softmax normalization and take the index corresponding to the highest probability. ,Right now: ; Subsequently, for each Collect all that satisfy pixel set A morphological closing operation (structural element radius 3 pixels) is applied to fill the internal holes, and then the connected regions are calculated. Only regions with an area greater than 200 pixels are retained as valid instances. The final output is an instance mask set. ,in This represents the actual number of instances detected. This process is entirely dependent on... The ability to discriminate, and And from Generated through deterministic encoding, therefore each mask The spatial boundary is The geometric accuracy of the projection of the material boundary highlighted by physical enhancement into the depth feature space is directly constrained by the fidelity of the entire physical modeling chain from step one to step two.

[0030] Step 3.3: Regress pixel-wise two-dimensional motion vector field from the same joint feature map In parallel with step 3.2, the motion decoding branch receives the same... After four sets of upsampling and convolution, a two-dimensional vector field is output. For each pixel This represents the normalized displacement (in pixels) that the pixel is expected to move in the next frame. To ensure alignment between the vector field and the segmentation result space, the last convolutional kernel of the motion branch is initialized to zero, and only residual corrections are learned; simultaneously, during the training phase, a constraint is enforced: for any ,like ,but The vector mean deviation from that of other pixels within the same instance must be less than a threshold. This constraint is achieved through a consistency term in the loss function: ; This loss term forces all pixels within the same instance to output highly consistent motion vectors, thus ensuring that the motion field is smooth inside the instance mask and abruptly changes at the mask boundary, consistent with the assumption of rigid body motion of the target. Because and shared ,and Source Therefore The direction and amplitude are essentially The reliability of the differentiable extension of the target appearance in the temporal dimension is positively correlated with the enhanced physical fidelity in step two.

[0031] Step 3.4: Fuse instance masks and motion vectors to generate a set of target instances with motion attributes The mask set obtained in step 3.2 The vector field obtained in step 3.3 Combined, for each instance Calculate its overall motion properties. For each Define its centroid location: And calculate its average motion vector: Furthermore, calculate the motion confidence score for this instance: The closer this value is to 1, the higher the consistency of motion within the instance, and the more likely it is to be a real rigid body target. Finally, the output is a set of target instances with motion attributes: Each element in this set contains a quadruple of spatial mask, position, direction of motion, and confidence level, which is the result of step two. A complete instantiation representation in the spacetime dimension. Because All originate from the same And from Deterministic encoding, therefore There is no modal separation between the various attributes, which provides a geometric and motion-based dual alignment basis for cross-view association in step four.

[0032] Step 4: Construct a multi-view spatiotemporal graph and perform cross-camera target identity propagation. This step will come from Different cameras Abstracted into a set of nodes, a spatiotemporal graph structure is constructed based on physical space constraints and motion continuity, and cross-perspective migration of identity tags is achieved through graph propagation mechanism.

[0033] Step 4.1: Extract the spatiotemporal feature vectors of each viewpoint instance and construct the initial node set. For each camera Take the output of step three. For each of these instances Construct its spatiotemporal feature vector: , , , , , ; in For the number of mask pixels, The distance from the centroid to the top-left corner of the image is the Euclidean distance. For the speed of motion, The angle between the direction of motion and the horizontal axis (in radians). For motion consistency confidence. This five-dimensional vector is entirely composed of... The calculations for each element are derived without any external parameters. All... corresponding Considered as graph nodes, a total of 10 nodes, forming the initial node set Each node Related Unique Perspective Index With instance index Its characteristics It is a compact numerical mapping of the target instantiation result in step three, carrying the observable spatiotemporal attributes of the target from the current perspective.

[0034] Step 4.2: Construct edge sets based on motion trajectory similarity and spatial visibility. Connections between nodes are not necessarily fully connected; they must satisfy physical feasibility. For any two nodes... Add an undirected edge to the graph if all of the following conditions are met. : (1) That is, from different cameras; (2) A pixel indicates that the two centroids are projected close enough on their respective image planes that they may be the same physical target. (3) The angle between the directions of motion is less than the threshold: Radius (approximately 40 degrees), and the speed ratio is between 0.5 and 2.0: ; (4) The motion confidence of both instances is higher than the threshold: .

[0035] This four-fold screening mechanism ensures that every edge Each represents a cross-perspective candidate match that is highly compatible in terms of spatial projection, motion trend, and credibility. (Boundary set) Completely by Features of each node The calculation and determination require no prior camera parameters; its existence directly reflects the inherent consistency of the target attributes extracted in step three. Therefore, the figure... Topological manifestation of the contained spatiotemporal information.

[0036] Step 4.3: Define node identity propagation weights and initialize the identity label matrix. To achieve identity propagation, each edge needs to be... Assign propagation weights This reflects the reliability of the connection in transmitting identity information. Definition: ; This weighting incorporates four factors: the confidence level of the two nodes' own motion (product term), the spatial projection distance decay (exponential term), and the cosine similarity of the motion directions (cosine term). All factors are set to values ​​from... And only if all the conditions in step 4.2 are met. Then, the identity tag matrix is ​​initialized. ,in The maximum number of identity categories is preset (128). For each node... If it appears for the first time in the current frame (i.e., it has not been assigned an identity before), then... The rest are 0, among which A unique serial number is temporarily assigned to it; if the node already has an associated historical identity, the corresponding column is inherited. This initialization depends only on whether the node is a new detection, and the determination of a new detection is based on step three. The set of instances is updated, therefore The initial state is the continuation of the time series output in step three.

[0037] Step 4.4: Perform multi-hop graph propagation and output consistent identity assignment results across viewpoints. In the figure Above, define the propagation operator. The normalized form of the weighted adjacency matrix: The matrix satisfies that the row sum is 1, which can be considered as Markov transition probabilities. Perform three rounds of propagation: ;in This is the initial label matrix. In each round of propagation, node labels are smoothly diffused to their neighbors, with high-weight edges dominating the information flow. After three rounds, for each node... ,Pick Index of column corresponding to the maximum value This serves as their final identity label. The final output is the cross-perspective identity assignment result: In this result, the same All the triples corresponding to the values ​​represent observation instances of the same physical target viewed from different cameras. Since the propagation process is entirely determined by... Topology and weights drive, and And the output from step three Construction, therefore This is a natural extension of the single-view target understanding obtained in step three into the multi-view space, marking the completion of cross-camera behavior association.

[0038] Step 5: Construct a target-level behavioral semantic graph and extract spatiotemporal action primitive sequences. Step 5.1: Aggregate cross-perspective spatiotemporal observation sequences by identity tag Based on the results obtained in step 4.4 For each unique identity Collect all associated triples, denoted as set. ,in For each frame, the timestamp is provided (in milliseconds, obtained from system clock synchronization). Extract its quaternary attribute calculated in step 3.4: mask Center of mass Mean motion vector Motion confidence Arrange these attributes in ascending order of time to form the identity. Complete spatiotemporal observation sequence: ; Each item in this sequence comes from the deterministic output of step three, and after propagation in step four, it has achieved cross-perspective consistency; timestamp To ensure absolute physical time, data from different cameras is aligned on the same timeline. Therefore, It is identity Complete, observable records from the real world are the only original material for subsequent behavioral analysis.

[0039] Step 5.2: Reconstruct the 3D centroid trajectory and calculate the local motion curvature sequence For each It needs to be distributed across the two-dimensional centroids of each image plane. The trajectory is merged into a unified three-dimensional spatial trajectory. Instead of using calibration parameters, a motion-constrained inversion method is employed: assuming the target moves at approximately a constant velocity within a short time window, then any two adjacent observation points... The line-of-sight intersections should be located near each other in the same spatial location. Define the time window length. milliseconds, for each ,Pick China satisfies All observations are denoted as the local observation set. For each pair of observations Construct its line-of-sight vector: ;in For homogeneous image coordinates, For camera Approximate principal point and focal length matrix (take) (Within the tolerance range, it does not affect the relative geometry). For camera A rough orientation rotation matrix is ​​provided (preset by the installation azimuth and elevation angles, with an error within ±5°). Although not precisely calibrated, this approximation is sufficient since it is only used for relative distance sorting. Subsequently, the shortest distance pair between the two lines of sight is solved. And take the midpoint as the three-dimensional estimate of the pair: For all Yes, calculate the weighted average of its three-dimensional estimates: ; in To correspond to the motion confidence level, Same as before. That is, identity At any moment The reconstructed three-dimensional centroid location. For All , obtain the trajectory Subsequently, cubic spline interpolation was performed on the trajectory (nodes are...). ), to obtain continuous functions And calculate its curvature: ; in Indicates to Differentiate. Curvature sequence It characterizes the local bending intensity of the trajectory and is a key indicator of the transition in motion, entirely determined by... The centroid and confidence level are derived from the geometric extension of the identity unification result in step four in three-dimensional space.

[0040] Step 5.3: Extract the start and end boundaries of the action based on mask deformation energy. Besides trajectory curvature, changes in the target's appearance also imply behavioral semantics. Every moment Calculate its mask contour length With area And define deformation energy: ; This energy reflects the rate of change of the target's scale and shape on the imaging plane. If The value remains above the threshold within a certain range. This indicates that the target is undergoing significant pose adjustment or occlusion interaction. To locate the start and end points of the action, [the following is used:] Sequence adaptive thresholding: First calculate the sliding window Mean within 5 frames (length) with standard deviation Set dynamic threshold ;when And both the preceding and following frames are lower than Then mark Let each be a potential starting point; similarly define the ending point. All starting and ending points constitute the candidate boundary set. Sort by time. This set is entirely composed of Middle mask Driven by geometric evolution, and This also originates from the deterministic decoding of the joint feature map in step three, therefore It is an objective representation of the deformation law of the target itself, without relying on manual annotation or prior action templates.

[0041] Step 5.4: Generate a sequence of action primitives by fusing curvature extrema and deformation boundaries. The curvature sequence obtained in step 5.2 The boundary set obtained in step 5.3 Perform alignment and blending. For each In its neighborhood Inner search for local maxima of curvature If it exists and Then Included in the modified boundary set Otherwise, discard. Then, for Sort and take adjacent boundary pairs ,Require milliseconds and Milliseconds, excluding excessively short jitter periods and excessively long periods of stillness. Ultimately, each valid interval... Define an action primitive: ; in This primitive contains complete three-dimensional trajectory segments, curvature response, appearance evolution, and confidence decay information, and is an identity... The smallest identifiable unit of behavior performed in physical space. All Constituting identity Action primitive sequence This is the first semantic assignment of the identity established in step four in the behavioral dimension, providing atomic-level input for cross-identity behavioral association in step six.

[0042] Step Six: Construct a multi-identity spatiotemporal interaction graph and identify collaborative and conflict behavior patterns. Step 6.1: Calculate the spatiotemporal proximity of identity pairs within the action primitive time period. For any two different identities traverse its action primitive sequence All primitive pairs For each pair, calculate its time overlap: ; in Indicates the length of the time interval (milliseconds). A larger value indicates that the two actions are more synchronized in time. Only when... At that time, it was believed that there was a potential interaction period. Subsequently, Within, for each sampling time (Step size 100ms), calculate the distance between the three-dimensional centroids of the two individuals: ; Define the average proximity of this primitive pair: The higher the score, the more frequently the two identities are within a safe distance during the interaction period. The three-dimensional trajectory reconstructed entirely from step five The calculation shows that it is and Direct measurements in physical space form the basis of subsequent interactive modeling.

[0043] Step 6.2: Define motion coupling degree and construct identity pair weighted edges Proximity only reflects spatial co-occurrence; motion coordination also needs to be characterized. For the same primitive pair... ,exist Calculate the mean cosine of the angle between the motion directions of the two individuals: Meanwhile, the computational speed is better than the stability: The higher this value, the more constant the rate ratio between the two identities. Combining these three indicators, an identity pair is defined. Coupling strength under this primitive pair: ; It can be negative, take Only cases of unidirectional or weakly opposite directions are retained (conflicts often manifest as strong opposite directions, so coupling is not considered here). All The primitive pairs constitute identity pairs. The weighted edges, with weights of The maximum value (taking the strongest interaction). This set of edges constitutes the initial interaction graph. , where the node set For all identities, edge It exists if and only if it exists This diagram is entirely composed of the action primitive sequence output from step five. It is derived from its three-dimensional trajectory and is the first structured organization of single-objective behavior in the group dimension.

[0044] Step 6.3: Identify stable cooperative clusters based on interaction intensity distribution exist Above, hierarchical clustering is performed to discover stable, cooperating groups. Nodes are defined. Weighted degree centrality: ;in for The strongest coupling value between them. Select... The node is used as a high-activity identity seed. For each seed... Collect from their neighbors The nodes form a preliminary cluster. Subsequently, iterative expansion: for Join all that meet new node Until no new cases are added. Ultimately, for all... Merge overlapping clusters while preserving size And internal average coupling degree Clusters of these clusters are denoted as cooperative clusters. Each cluster represents a set of identities that are highly coupled in space, time, and motion, and its formation depends entirely on... The edge weight distribution.

[0045] Step 6.4: Extract dominant behavior patterns within the cluster and output cooperative behavior labels. For each cooperative cluster Extract the action primitives of all its members during the period the cluster exists. Define the cluster behavior pattern as the average motion direction and speed of all primitives in three-dimensional space: ;like Marked as "Resident Collaboration"; if Marked as "moving in the same direction"; if Furthermore, the variance of the relative distance between the centroids within the cluster This is labeled "Closely Formed". Finally, the output is a set of cooperative behavior labels: Each element in this set is a behavioral semantic annotation of the cooperative clusters identified in step 6.3, and its label is directly calculated from the average motion attribute within the cluster.

[0046] Step 7: Perform spatiotemporal consistency verification and abnormal behavior detection. Step 7.1: Construct the spatiotemporal reference envelope of the cooperative cluster For each cooperative cluster obtained in step 6.4 Using the set of trajectories of all its member action primitives in three-dimensional space as samples, a spatiotemporal reference envelope is constructed for this cluster. For each time offset... Milliseconds (step size 200ms) are used to calculate the identities of all members within a cluster at time 10:00. set of centroids convex hull The convex hull characterizes the spatial extent of the cluster at that time offset. Simultaneously, it calculates the identities of all members within the cluster. rate set at 90th percentile With 10th percentile This forms the rate envelope. Together, they constitute a cluster. spatiotemporal reference envelope The envelope is entirely composed of The action primitive trajectories of each identity are derived from the geometric representation of the collaborative behavior in step six, providing a rigid reference for subsequent consistency verification.

[0047] Step 7.2: Calculate the three-term deviation measure between the new action primitive and the reference envelope. When new identity Action primitives After generation, perform three verifications on it: (1) Spatial deviation: for Every moment Calculate its three-dimensional centroid To the nearest co-cluster Reference envelope The Euclidean distance, where The initial time of this primitive; take all Maximum distance on .

[0048] (2) Rate deviation: for the same ,examine Has it fallen into exist Within the rate envelope; defined as not falling into the proportion .

[0049] (3) Rhythm deviation: Calculation Duration with cluster Average duration of similar action primitives The relative difference: ; All three biases are for the best-matching cluster. (even though The smallest one is calculated. This process completely reuses the envelope constructed in step six. No new parameters were introduced; all deviation values ​​were derived from... The geometric comparison directly produces results.

[0050] Step 7.3: Perform adaptive multi-threshold joint discrimination An anomaly is triggered only if all three deviations exceed the limit simultaneously. The threshold is not fixed but adaptively updated based on cluster historical fluctuations: for each cluster... Maintain its historical deviation mean and standard deviation The current discrimination threshold is set as follows: ;like and and Then determine This is an anomalous primitive. Subsequently, historical statistics are updated: The same applies to the rest. This adaptive mechanism ensures that the threshold always matches the dynamic characteristics of the current cooperative cluster, avoiding false alarms caused by slow environmental changes.

[0051] Step 7.4: Classify anomaly types and output structured anomaly reports For each action primitive that is judged to be abnormal Further categorization: like Not belonging to any existing cooperative cluster (i.e. (No match found), and its If it deviates significantly from all clusters in terms of space, rate, and rhythm, it is marked as "free individual anomaly"; if Originally belonged to a certain cluster However, if the deviations of the primitive elements exceed the limits and no similar deviations are observed in other members of the cluster during the same period, it is marked as "individual detachment anomaly"; if multiple members originally belonging to the same cluster detach from the cluster during the same period... If all identities are deemed abnormal and their deviation patterns show a divergent trend, then it is marked as "collaborative structure collapse".

[0052] Finally, a structured anomaly report is output: Each entry in this report is a semantic classification of the discrimination results in step 7.3, and the classification is based entirely on the combination pattern of the three biases.

[0053] Step 8: Generate multi-granularity linkage response commands and drive physical execution units. Step 8.1: Generate a basic response level based on the anomaly type and location. right Each anomaly First, determine its response level. :like "Free individual abnormality", and If it occurs in a non-priority area (such as the periphery of the park), then (Information recording level); if "Individual detachment from abnormality", and Occurring at passageways or entrances / exits, or Rice, then (Enhanced monitoring level); If "Disintegration of the collaborative structure", or ,but (Proactive Intervention Level). This level classification is based solely on... The existing type labels and three deviation values ​​are used; no new measurements are introduced. (Level) It becomes the top-level constraint for generating all subsequent instructions, determining the strength and coverage of the response.

[0054] Step 8.2: Generate view scheduling instructions and bind them to the spatiotemporal execution window. For level An anomaly generates a viewpoint scheduling command. Let the anomaly primitive be... The three-dimensional spatial coverage area is Select a field of view that can cover And the current set of cameras not participating in collaborative tracking For each Calculate its center line of view and minimum distance and angular coverage (by Project to The image plane is obtained by calculating the proportion of the bounding rectangle. (Select...) The first two cameras generate scheduling instructions: ; ; All instructions come with an execution time window. This ensures that the actions and abnormal behaviors of physical equipment are strictly contained in time. This instruction set is entirely composed of... three-dimensional trajectory It is derived from its spatiotemporal scope.

[0055] Step 8.3: Generate audio-visual warnings and path guidance instructions For level In the event of an anomaly, in addition to viewpoint adjustment, on-site intervention commands are generated. Based on geometric center Locate the nearest physical warning unit (such as a sound column or LED screen) For each Calculate its to distance Select The unit is meters, and the generation instructions include precise geographic coordinates. The execution time should be matched with the sound and light coverage and the range of abnormal spaces.

[0056] Step 8.4: Encapsulate the instruction set and write it to an auditable execution log. All instructions (viewpoint scheduling, audio-visual warnings, path guidance) generated in steps 8.1–8.3 are encapsulated into logs according to anomaly entries. These logs ensure that each physical response can be traced back to the pixel-level image input and verified backward to the device execution result, forming an end-to-end auditable response loop.

[0057] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.

Claims

1. A multi-camera behavior correlation analysis method based on adaptive illumination enhancement, characterized in that, include: The original grayscale image is acquired and a pixel-level illumination degradation model is constructed to generate an illumination-invariant reflective substrate image. Adaptive illumination enhancement is performed using an illumination-invariant reflective substrate image, combined with multi-scale gradient noise suppression to generate an enhanced image; By extracting spatial-motion joint features from the enhanced image through a shared encoder, pixel-level segmentation masks and motion vector fields are generated simultaneously to form a set of target instances with motion attributes. Based on the motion trajectory and spatial projection relationship of the target instance, a multi-view spatiotemporal map is constructed. Cross-camera identity unification is achieved through three rounds of weighted graph propagation, and cross-view identity allocation results are output. Aggregate observation data from various perspectives according to identity tags, reconstruct the three-dimensional centroid trajectory, analyze mask deformation energy, and generate a spatiotemporal action primitive sequence. Calculate the spatiotemporal proximity and motion coupling between action primitives of different identities, identify stable cooperative clusters and extract dominant behavior patterns, and output cooperative behavior labels; A spatiotemporal reference envelope is constructed for collaborative behavior labels, and anomaly detection is performed by measuring the deviation in three dimensions: space, rate, and rhythm, generating structured anomaly reports. Based on the type and severity of the anomaly report, the camera view scheduling and audio-visual warning units are driven to execute multi-granular linkage responses.

2. The multi-camera behavior correlation analysis method based on adaptive illumination enhancement according to claim 1, characterized in that, The construction of the pixel-level illumination degradation model includes: The original grayscale image is linearized by sensor response correction to generate a linear domain image; Based on the linear domain image, the distribution of maximum illumination intensity within a local window is calculated to construct a spatially varying illumination field; By using the spatially varying illumination field as a constraint, the Gauss-Poisson equations are solved jointly to obtain reflectance maps and shadow masks. The reflectance map is then subjected to gamma correction and dynamic range compression to generate an illumination-invariant reflectance substrate image.

3. The multi-camera behavior correlation analysis method based on adaptive illumination enhancement according to claim 1, characterized in that, The implementation of adaptive lighting enhancement includes: The enhancement objective function is calculated with reference to the brightness distribution of the illumination-invariant reflective substrate image; By combining the enhancement objective function with the local gradient magnitude of the linear domain image, a brightness transfer mapping field is constructed. Wavelet domain thresholding filtering is applied to linear domain images to preserve high-frequency details and suppress noise. Applying a brightness transfer map field to the filtered image generates an enhanced image that preserves edge details.

4. The multi-camera behavior correlation analysis method based on adaptive illumination enhancement according to claim 1, characterized in that, The formation of the target instance set with motion attributes includes: The enhanced image is input into a shared encoder with a pyramid structure to extract a four-level space-motion joint feature map. The joint feature map is processed by the instance segmentation head to generate pixel-level instance masks and instance confidence scores. By processing the same joint feature map using a motion regression head, horizontal and vertical motion vector components are generated. The instance mask is fused with the motion vector components, and the centroid coordinates and average motion vector of each instance are calculated to form a set of target instances with motion attributes.

5. The multi-camera behavior correlation analysis method based on adaptive illumination enhancement according to claim 1, characterized in that, The achievement of unified identity across cameras includes: Extract motion trajectory fragments and appearance feature vectors for each instance from a set of target instances with motion attributes; Based on the similarity of motion trajectory segments and the overlapping area of ​​the camera's field of view, edge connections of a multi-view spatiotemporal graph are constructed; Each edge is assigned a weight, which is determined by both motion similarity and spatial projection distance. Perform three rounds of weighted message passing on a multi-view spatiotemporal graph, aggregate the identity labels of neighboring nodes, and output the cross-view identity allocation results.

6. The multi-camera behavior correlation analysis method based on adaptive illumination enhancement according to claim 1, characterized in that, The generated spatiotemporal action primitive sequence includes: Collect observation sequences of the same identity across all cameras based on cross-view identity assignment results; A three-dimensional centroid trajectory is reconstructed by using the line-of-sight intersection estimation method within a short time window and integrating the centroid positions from multiple perspectives. Calculate the derivatives of the mask contour length and area with respect to time, generate deformation energy curves, and detect energy abrupt change points; Aligning the extreme points of curvature of the three-dimensional trajectory with the abrupt change points of deformation energy, action primitives with durations ranging from 400 milliseconds to 6 seconds are segmented.

7. The multi-camera behavior correlation analysis method based on adaptive illumination enhancement according to claim 1, characterized in that, The output cooperative behavior labels include: Calculate the three-dimensional centroid distance integral of different identity action primitives in the temporal overlap region to obtain the spatiotemporal proximity. Analyze the stability of the cosine average of the angle between the motion direction vectors during the overlapping period and the velocity ratio, and calculate the motion coupling degree; An identity interaction graph is constructed based on spatiotemporal proximity and motion coupling, and collaborative clusters with high internal coupling are identified through hierarchical clustering. Calculate the average movement speed within the cooperative cluster. When the speed is below 0.3 m / s, it is marked as stationary cooperative. When the speed is above 0.3 m / s and the movement direction is consistent, it is marked as moving in the same direction. When the variance of the distance between members is less than 2.0 square meters, it is marked as close formation.

8. The multi-camera behavior correlation analysis method based on adaptive illumination enhancement according to claim 1, characterized in that, The generation of structured anomaly reports includes: Calculate the three-dimensional convex hull of the member positions for each time offset point of the cooperative cluster to form a spatiotemporal reference envelope; Measure the maximum spatial distance between the new motion primitive and the reference envelope, the out-of-envelope ratio of the velocity envelope, and the relative deviation of the duration; The judgment threshold is dynamically updated based on the cluster's historical deviation data. When all three deviations exceed the adaptive threshold simultaneously, the data is marked as abnormal. Based on the relationship between anomalous identities and synergistic clusters and the deviation patterns, they are classified into anomalous free individuals, anomalous individual detachment, or synergistic structure disintegration.

9. The multi-camera behavior correlation analysis method based on adaptive illumination enhancement according to claim 1, characterized in that, The execution of multi-granularity linkage response includes: The response level is calculated based on the anomaly type and spatial location. Anomalies of free individuals in non-key areas are set as Level 1, anomalies of individuals detached from the passageway are set as Level 2, and anomalies of collaborative structure disintegration are set as Level 3. For the spatial coverage area of ​​the computational action primitives for Level 2 and Level 3 anomalies, select the first two available cameras that can cover the area and have a viewing angle greater than 0.65, and generate zoom and tracking commands. At the geometric center of the computational space coverage area of ​​Level 3 anomalies, activate the audio-visual warning unit within a 15-meter radius, and set an 85-decibel directional sound field and a 5-second visual warning. Associate all instructions with the original exception data, execution time window, and spatial coordinates, and write them to an auditable execution log containing an immutable timestamp and data hash value.

10. A multi-camera behavior correlation analysis system based on adaptive illumination enhancement for implementing the method as described in any one of claims 1-9, characterized in that, include: The illumination modeling unit is used to acquire the original grayscale image and construct a pixel-level illumination degradation model to generate an illumination-invariant reflective substrate image. An image enhancement unit is used to perform adaptive illumination enhancement using an illumination-invariant reflective substrate image and generate an enhanced image by combining multi-scale gradient noise suppression. The target analysis unit is used to extract spatial-motion joint features from the enhanced image through a shared encoder, and simultaneously generate pixel-level segmentation masks and motion vector fields. The identity association unit is used to construct a multi-view spatiotemporal map based on the motion trajectory and spatial projection relationship of the target instance, and to achieve cross-camera identity unification through weighted graph propagation. The behavior analysis unit is used to aggregate observation data from various perspectives by identity tags, reconstruct the three-dimensional centroid trajectory, and analyze the mask deformation energy. The interaction recognition unit is used to calculate the spatiotemporal proximity and motion coupling between different identity action primitives and to identify stable collaborative behavior patterns. The anomaly detection unit is used to construct a spatiotemporal reference envelope for collaborative behavior and perform anomaly behavior identification through multi-dimensional deviation measurement. The response execution unit is used to drive the camera view scheduling and sound and light warning unit according to the severity of the anomaly, and to execute multi-granular linkage response.