Camera layout position automatic determination method based on video analysis

By generating reflection mask data and cross-camera field-of-view correlation maps, and combining homography matrix and dynamic time warping calculations, the problem of optical image interference in the deployment of surveillance cameras was solved, enabling precise deployment and resource optimization in complex environments.

CN122391362APending Publication Date: 2026-07-14SICHUAN LANGJI INNOVATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN LANGJI INNOVATION TECHNOLOGY CO LTD
Filing Date
2026-04-23
Publication Date
2026-07-14

Smart Images

  • Figure CN122391362A_ABST
    Figure CN122391362A_ABST
Patent Text Reader

Abstract

The application discloses a camera layout position automatic determination method based on video analysis, relates to the technical field of confirmation analysis, and comprises the following steps: calculating conditional information entropy of a mirror area relative to a real object area in a real object-mirror pair; if the conditional information entropy is lower than a preset redundancy threshold, performing zero processing on the layout value weight corresponding to the mirror area; if the conditional information entropy is higher than a preset gain threshold, marking the mirror area as a virtual gain field of view and improving the layout value weight corresponding to the mirror area; mapping the changed reflective medium area to a three-dimensional space map to generate a layout space potential field containing repulsive force fields and attractive force fields; taking the layout space potential field as an environmental constraint boundary, iteratively calculating by using a layout optimization algorithm, and outputting camera physical coordinate parameters and a holder posture parameter which avoid the repulsive force fields and tend to the attractive force fields. The application has the effect of improving the camera layout position determination efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of verification analysis technology, and in particular to a method for automatically determining the location of cameras based on video analysis. Background Technology

[0002] Current technologies for automated joint deployment of surveillance cameras typically rely on static geometric distance calculations based on 3D scene models or 2D video feature density analysis based on a single viewpoint. They often assume that the effective field of view of each camera is absolutely independent and only affected by physical obstacles. This idealized assumption of independent fields of view severely deviates from the actual optical characteristics of complex physical environments. Consequently, deployment algorithms completely lack mechanisms to prevent optical image interference between cameras when facing complex spaces with large areas of smooth reflective media.

[0003] Specifically, traditional placement optimization algorithms, when processing video streams, only extract texture richness and target motion trajectories as criteria for evaluating region coverage value. When the fields of view of multiple cameras converge on a smooth physical medium, light reflection projects the real dynamic scene from afar onto the surface, creating a mirrored ghosting region. The single-field-of-view analysis logic of traditional algorithms mechanically identifies this ghosting region as a legitimate activity space with high-frequency motion characteristics and incorrectly assigns it an extremely high placement coverage weight. This deep-seated logical flaw directly induces the system to repeatedly deploy devices at extremely close physical distances in an attempt to cover the reflected image. This not only disrupts the balanced distribution of the global placement topology but also, due to the lack of a spatiotemporal collaborative verification mechanism between multiple perspectives, completely masks the physical mapping conflict between the real scene and the false mirror image. Consequently, the overall placement scheme suffers from insurmountable technical bottlenecks in spatial perception performance and environmental adaptability, requiring further improvement. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this application provides a method for automatically determining camera placement locations based on video analysis.

[0005] The automatic camera placement method based on video analysis provided in this application includes the following steps: Acquire synchronous video streams from multiple candidate placement locations, process pixel distribution information in the synchronous video streams, generate reflection mask data corresponding to the reflective medium region, and delineate the non-masked effective region in the synchronous video stream based on the reflection mask data; extract the set of static feature points and the sequence of dynamic target trajectories within the non-masked effective region, calculate the matching correlation degree between the static feature point set and the images corresponding to different candidate placement locations, construct a cross-camera field of view correlation map based on the matching correlation degree, and extract node pairs with matching correlation degrees higher than a preset correlation threshold in the cross-camera field of view correlation map as sub-region pairs to be verified; Trigger a mirror verification loop for the sub-region pair to be verified. The mirror verification loop includes: calculating the homography matrix between the sub-region pair to be verified, extracting the determinant sign of the homography matrix to determine the geometric symmetry features; performing dynamic time warping calculation on the dynamic target trajectory sequence within the sub-region pair to be verified, and obtaining trajectory mirror synchronization data; extracting the optical feature parameters of the sub-region pair to be verified, and calculating the physical characteristic deviation including brightness attenuation and contrast loss values. When the geometric symmetry features, trajectory mirror synchronization data, and physical characteristic deviations all meet the preset mirror judgment conditions, the sub-region pair to be verified is marked as a physical-mirror pair, and the value reassessment logic is initiated. The value reassessment logic includes: calculating the conditional information entropy of the mirror region relative to the physical region in the physical-mirror pair; if the conditional information entropy is lower than the preset redundancy threshold, the placement value weight corresponding to the mirror region is reset to zero; if the conditional information entropy is higher than the preset gain threshold, the mirror region is marked as a virtual gain field of view and the placement value weight corresponding to the mirror region is increased. The reflective medium region after the change of the placement value weight is mapped onto the three-dimensional spatial map to generate a placement spatial potential energy field containing repulsive force field and attractive force field. Among them, the mirror region after zeroing generates a repulsive force field, and the virtual gain field of view generates an attractive force field. Using the spatial potential energy field of the deployment point as the environmental constraint boundary, the deployment point optimization algorithm is used for iterative calculation to output the camera physical coordinate parameters and gimbal attitude parameters that avoid the repulsive force field and tend to the attractive force field.

[0006] In summary, this application includes at least one of the following beneficial technical effects: 1. This application provides a method for automatically determining camera placement based on video analysis. It generates reflection mask data by synchronizing pixel distribution information in the video stream and delineates the effective non-mask area in reverse, realizing active perception of the reflective medium area. By calculating the cross-camera field of view correlation map and extracting the matching sub-region pairs to be verified with a correlation degree higher than the threshold, it triggers a mirror verification loop that includes a triple judgment of "geometric symmetry features, trajectory mirror synchronization data, and physical characteristic deviation". Thus, it can accurately distinguish between real scene areas and mirror ghost areas from multiple dimensions such as spatial flip relationship, dynamic target spatiotemporal synchronization, and optical physical characteristic attenuation. 2. By calculating the conditional information entropy of the mirror region relative to the physical region in a physical-mirror pair, the incremental information contained in the mirror region, independent of the physical region, is quantified. When the conditional information entropy is lower than a preset redundancy threshold, its deployment value weight is set to zero, thereby effectively preventing the system from repeatedly deploying devices at physically close locations to cover the same scene as a false mirror image; when the conditional information entropy is higher than a gain threshold, it is marked as a virtual gain field of view and its weight is increased. This differentiated processing breaks the rigid balance of traditional deployment topology and realizes adaptive optimization configuration of deployment resources in complex optical environments. 3. The mirrored region after value reassessment is mapped onto a 3D spatial map, generating a spatial potential energy field containing repulsive and attractive force fields. Using this potential energy field as the environmental constraint boundary, a point optimization algorithm is used for iterative calculation. This allows the optimization process of camera physical coordinates and gimbal attitude to naturally avoid spatial positions that generate mirror redundancy interference, while actively tending towards positions that can provide additional effective field of view information. This mechanism not only ensures a balanced distribution of the global point topology, but also upgrades the original static point placement problem based purely on geometric distance or texture density into a dynamic collaborative optimization problem that integrates the physical laws of optical reflection and information theory value criteria, significantly improving the system's spatial perception performance in complex environments. Attached Figure Description

[0007] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0008] Figure 1 This is a flowchart of a method for automatically determining the location of cameras based on video analysis, according to an embodiment of this application.

[0009] Figure 2 This is a flowchart of a method for creating a cross-camera field-of-view association diagram according to an embodiment of this application. Detailed Implementation

[0010] The following description, in conjunction with the implementation of this invention, is merely an example and illustration of the concept of this invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the inventive concept or exceed the scope defined in these claims, all of which should fall within the protection scope of this invention.

[0011] Application Overview: In existing technologies, automated joint deployment of surveillance cameras largely relies on static geometric distance calculations based on 3D scene models or 2D video feature density analysis based on a single viewpoint. This makes it difficult to balance the uniformity of global distribution with the actual optical characteristics of complex physical environments. When traditional methods encounter large areas of smooth reflective media in a scene, the single-viewpoint analysis logic is susceptible to optical image interference, causing the algorithm to misclassify dynamic false features within the reflective surface as high-value independent activity areas. This leads to severe clustering of deployment locations and wasted equipment resources. Existing mechanisms cannot simultaneously perceive the impact of spatiotemporal coordination between multiple viewpoints on physical mapping conflicts. Especially in spaces with dense reflective sources, such as glass curtain walls or calm water surfaces, the independent viewpoint assumption will exhibit systematic coverage deviations, making it difficult to meet the precise configuration requirements of surveillance networks in complex optical environments.

[0012] To address the aforementioned issues, the inventors discovered a strict law of geometric spatial reversal and kinematic mirror co-operation between the physical reality and the optical reflective surface. By establishing a multi-dimensional spatiotemporal and physical optics coupling verification model, they achieved accurate image ghosting removal and error compensation. During the research, it was found that pure geometric judgment is easily affected by perspective distortion, while a comprehensive judgment integrating homography mapping, dynamic time warping verification, and physical illumination energy loss characteristics can significantly improve the confidence level of image recognition. Therefore, they proposed a potential energy guidance approach that transforms redundant images and virtual gain fields of view into spatial repulsion and attraction forces. Further experimental verification involved introducing the spatial potential energy field of the placement points as a nonlinear constraint boundary into the particle swarm optimization algorithm. Combined with a closed-loop dynamic verification mechanism after physical execution by the device, an automated placement correction closed-loop system was formed.

[0013] Specifically, the detection system first synchronously acquires video streams from each candidate location. It then extracts pixel-level reflectance confidence scores through semantic segmentation and Fresnel variation fitting. After stripping high-frequency reflective areas, it constructs a cross-camera field-of-view association map using local descriptor matching. Simultaneously, the system uses the determinant sign of the homography matrix to determine the geometric symmetry of region pairs, combines a dynamic time warping algorithm to calculate the mirror synchronization degree of spatiotemporal trajectories, and extracts physical characteristic deviations such as brightness attenuation and contrast loss to accurately establish object-mirror pairs. Once a mirror image is identified, the system uses conditional information entropy to evaluate its gain value, defining purely redundant mirror images as repulsive force sources and high-value virtual gain fields of view as attractive force sources, generating a potential energy field in the location space. During iterative calculations, the system uses the comprehensive potential energy field as a reward / penalty factor superimposed into the fitness function to dynamically guide the camera particle swarm optimization. After outputting device parameters, the system reacquires secondary video streams through hardware feedback, performs secondary verification using a dynamic exponential penalty association threshold, and forms an optimization closed loop to dynamically eliminate mirror deadlock.

[0014] Compared to existing technologies, traditional methods rely on static 3D projection with a single, independent field of view and lack compensation mechanisms for morphological and optical laws. This makes them prone to systematic errors due to layout imbalances when reflective media are present in the environment. This solution innovatively integrates multi-view spatiotemporal dynamic characteristics with physical optics principles, achieving automatic isolation between the real scene and the false image by establishing a geometry-motion-photometry coupled judgment model. Unlike existing static and crude strategies that directly shield interference, this solution intelligently reassesses the potential value of the reflective surface to supplement the viewing angle based on conditional information entropy. Furthermore, it continuously optimizes the device's pose parameters by establishing a physically feedback-driven attractive and repulsive potential energy field, significantly improving monitoring and sensing efficiency and algorithm robustness under complex optical interference conditions inside and outside buildings.

[0015] Through the above technical solutions, this application effectively overcomes the problems of optical image misjudgment and equipment aggregation waste caused by smooth media in the physical environment, and improves the scientific nature of hardware resource distribution while ensuring the three-dimensional spatial coverage of the camera. The multi-dimensional spatiotemporal and optical fusion verification mechanism takes into account both the extremely high filtering ability of artifact features and the keen ability to capture potential gain fields of view. The particle swarm optimization based on potential energy field and the dynamic threshold anti-deadlock self-correction function ensure the accuracy of the final deployment scheme.

[0016] After introducing the basic concept of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0017] Example 1: This application discloses an automatic method for determining camera placement based on video analysis.

[0018] Reference Figure 1 The method for automatically determining camera placement based on video analytics includes the following steps: Acquire synchronous video streams from multiple candidate placement locations, process pixel distribution information in the synchronous video streams, generate reflection mask data corresponding to the reflective medium region, and delineate the non-masked effective region in the synchronous video stream based on the reflection mask data; extract the set of static feature points and the sequence of dynamic target trajectories within the non-masked effective region, calculate the matching correlation degree between the static feature point set and the images corresponding to different candidate placement locations, construct a cross-camera field of view correlation map based on the matching correlation degree, and extract node pairs with matching correlation degrees higher than a preset correlation threshold in the cross-camera field of view correlation map as sub-region pairs to be verified; Trigger a mirror verification loop for the sub-region pair to be verified. The mirror verification loop includes: calculating the homography matrix between the sub-region pair to be verified, extracting the determinant sign of the homography matrix to determine the geometric symmetry features; performing dynamic time warping calculation on the dynamic target trajectory sequence within the sub-region pair to be verified, and obtaining trajectory mirror synchronization data; extracting the optical feature parameters of the sub-region pair to be verified, and calculating the physical characteristic deviation including brightness attenuation and contrast loss values. When the geometric symmetry features, trajectory mirror synchronization data, and physical characteristic deviations all meet the preset mirror judgment conditions, the sub-region pair to be verified is marked as a physical-mirror pair, and the value reassessment logic is initiated. The value reassessment logic includes: calculating the conditional information entropy of the mirror region relative to the physical region in the physical-mirror pair; if the conditional information entropy is lower than the preset redundancy threshold, the placement value weight corresponding to the mirror region is reset to zero; if the conditional information entropy is higher than the preset gain threshold, the mirror region is marked as a virtual gain field of view and the placement value weight corresponding to the mirror region is increased. The reflective medium region after the change of the placement value weight is mapped onto the three-dimensional spatial map to generate a placement spatial potential energy field containing repulsive force field and attractive force field. Among them, the mirror region after zeroing generates a repulsive force field, and the virtual gain field of view generates an attractive force field. Using the spatial potential energy field of the deployment point as the environmental constraint boundary, the deployment point optimization algorithm is used for iterative calculation to output the camera physical coordinate parameters and gimbal attitude parameters that avoid the repulsive force field and tend to the attractive force field.

[0019] In a specific embodiment, the synchronized video stream refers to a sequence of real-time monitoring images captured by multi-view cameras at the same time reference. Specifically, it can be implemented using Network Time Protocol (NTP) alignment or hardware-triggered synchronization technology to provide a spatially consistent image data base in the scene. The reflection mask data refers to the image mask of potential highly reflective media such as glass and water surfaces identified through pixel distribution information. Specifically, it can be implemented using semantic segmentation networks or high brightness threshold detection algorithms to isolate interference areas that may cause optical distortion. The non-masked effective area refers to the real physical scene image portion remaining after removing reflective media from the video image. Specifically, it can be implemented by performing Boolean difference operations on the original image and the reflection mask data to ensure that the initial feature extraction and matching process is not contaminated by false images.

[0020] Among them, the cross-camera field-of-view association map refers to a network graph that represents the physical topological relationship of overlapping areas between different camera viewpoints. Specifically, it can be achieved by calculating the distribution density and weighted matching degree of static feature points in the grid area. It is used to locate regional nodes that may have spatial symmetry or overlapping real field of view. The trajectory mirror synchronization data refers to an indicator that quantifies the degree of spatiotemporal coordination of the motion trajectory of a dynamic target in the real object and the suspected mirror area. Specifically, it can be achieved by using the dynamic time warping (DTW) algorithm to calculate the cumulative distance matrix of the spatiotemporal state vector sequence. It is used to assist in confirming the mirror relationship from the kinematic dimension. The physical characteristic deviation refers to a quantitative parameter that reflects the difference in photometric attenuation between direct sunlight and specular reflection. Specifically, it can be achieved by extracting the mean brightness and the variance of the grayscale histogram in the color space, and calculating the weighted fusion of brightness attenuation and contrast loss value. It is used to confirm the existence of mirrors from the optical physics level.

[0021] Among them, conditional information entropy refers to the amount of additional effective visual information that a mirrored region can provide against the background of a known physical area. Specifically, it can be calculated by combining the difference between the joint information entropy of the gray-level co-occurrence matrix and the information entropy of a single edge. It is used to evaluate whether the reflective surface generates visual redundancy or has monitoring blind spot gain value. The deployment spatial potential energy field refers to transforming the three-dimensional monitoring environment into a dynamic distribution model controlled by a virtual force field. Specifically, it can be implemented by using a distance gradient decay function to generate a repulsive force field and a curvature logarithmic distribution function to generate an attractive force field. It serves as a digital spatial constraint boundary to guide the camera to avoid interference and tend towards a high-value field of view. The deployment optimization algorithm refers to the iterative calculation process of finding the optimal three-dimensional spatial coordinates of the camera and the attitude of the gimbal under the constraint of the potential energy field. Specifically, it can be implemented by using a particle swarm optimization algorithm or a genetic algorithm to optimize the position and attitude states encoded as particles. It is used to output the globally optimal monitoring network deployment parameters.

[0022] The working process and principle of this application are as follows: First, acquire synchronous video streams from multiple candidate placement locations, process the pixel distribution information in the synchronous video streams, generate reflection mask data corresponding to the reflective medium region, and delineate the non-masked effective region in the synchronous video stream based on the reflection mask data; then, extract the set of static feature points and the dynamic target trajectory sequence within the non-masked effective region, calculate the matching correlation degree between the static feature point set and the images corresponding to different candidate placement locations, construct a cross-camera field-of-view correlation map based on the matching correlation degree, and extract node pairs with matching correlation degrees higher than a preset correlation threshold in the cross-camera field-of-view correlation map as sub-region pairs to be verified; then, trigger a mirror verification loop for the sub-region pairs to be verified. The mirror verification loop includes: calculating the homography matrix between the sub-region pairs to be verified, extracting the determinant sign of the homography matrix to determine the geometric symmetry features, performing dynamic time warping calculation on the dynamic target trajectory sequence within the sub-region pairs to obtain trajectory mirror synchronization data, and extracting the optical feature parameters of the sub-region pairs to be verified to calculate the physical characteristic deviation including brightness attenuation and contrast loss values; subsequently, when the geometric pair When the symmetry characteristics, trajectory mirror synchronization data, and physical characteristic deviations all meet the preset mirror judgment conditions, the sub-regions to be verified are marked as object-mirror pairs, and the value reassessment logic is initiated. This involves calculating the conditional information entropy of the mirror region relative to the object region in the object-mirror pair. If the conditional information entropy is lower than a preset redundancy threshold, the placement value weight corresponding to the mirror region is zeroed out. If the conditional information entropy is higher than a preset gain threshold, the mirror region is marked as a virtual gain field of view, and the placement value weight corresponding to the mirror region is increased. Then, the reflective medium region after the placement value weight change is mapped onto a 3D spatial map to generate a placement spatial potential energy field containing repulsive and attractive fields. The zeroed-out mirror region generates a repulsive field, and the virtual gain field of view generates an attractive field. Finally, using the placement spatial potential energy field as the environmental constraint boundary, the placement optimization algorithm is used for iterative calculation to output the camera physical coordinate parameters and pan-tilt attitude parameters that avoid the repulsive field and tend towards the attractive field. In this way, adaptive optimization deployment of the monitoring network in complex optical environments is achieved, ensuring that the camera configuration achieves efficient coverage with no global redundancy.

[0023] Furthermore, reflection masking data corresponding to the reflective medium region is generated, and the non-masked effective region is delineated in the synchronous video stream based on the reflection masking data. This includes: extracting semantic features from the synchronous video stream to obtain material texture feature data containing material attribute tags and spatial distribution feature data containing planar normal vector information; constructing a preliminary reflection probability map for each frame using the material texture feature data, mapping the spatial distribution feature data to the preliminary reflection probability map for weight correction, and calculating the reflection confidence score of each pixel in the synchronous video stream; extracting the spatial distribution displacement vector of the reflection confidence score in multiple consecutive frames, identifying the set of pixels whose spatial distribution displacement vector is lower than a preset drift threshold within a preset time period, and using the set of pixels as the high-frequency reflection region; obtaining the intrinsic and extrinsic parameters of the camera at the candidate placement location, performing inverse projection transformation in combination with the pixel coordinates of the high-frequency reflection region, determining the geometric boundary of the high-frequency reflection region in the three-dimensional spatial map, generating reflection masking data based on the geometric boundary, and delineating the non-masked effective region by removing the pixel data corresponding to the reflection masking data in the synchronous video stream.

[0024] In one specific embodiment, the synchronous video stream is first subjected to in-depth pixel-level analysis and spatiotemporal mapping to achieve accurate localization and stripping of reflective media in the physical environment. The input synchronous video stream is first fed frame by frame into a pre-trained lightweight multi-task semantic segmentation network for feature extraction. The network decoder is divided into two parallel branches, which output material texture feature data containing physical material attribute labels (such as glass curtain wall, calm water surface, marble floor category) and spatial distribution feature data containing depth values ​​and planar normal vector information of scene geometry.

[0025] Next, the inherent base reflectivity of each pixel in the material texture feature data is extracted to generate a preliminary reflectivity probability map. Then, the planar normal vectors from the spatial distribution feature data are mapped onto this preliminary reflectivity probability map for weight correction. Finally, the reflectivity confidence score for each pixel in the synchronized video stream is calculated using the following formula: .in, Image coordinates The final reflection confidence score for each pixel is normalized and truncated to... interval; This is a preliminary baseline value for the reflection probability obtained by looking up a table based on the material texture classification; This is the normalized normal vector of the pixel in three-dimensional space extracted from spatial distribution feature data; Let be the normalized line-of-sight vector from the camera's optical center to the surface of the entity; the absolute value of their dot product. The cosine of the angle between the line of sight and the reflecting surface; The preset grazing angle gain coefficient has a value range of 1.5 to 3.0, preferably 2.0, to amplify the weight of "the more flush the line of sight is with the surface, the more obvious the reflection phenomenon" in physical optics; The normal sensitivity decay constant ranges from 0.3 to 0.7, preferably 0.5. The necessity of this variant formula lies in the fact that, compared to directly using the massive, complete ray tracing physics model, this variant introduces an exponential decay term. By relying solely on monocular vision to estimate the normal vector, the optical attenuation characteristic of increasing reflection with a tilted viewing angle is accurately fitted with extremely low computational overhead.

[0026] Subsequently, to eliminate false reflection noise caused by water ripple jitter or sudden changes in lighting, a temporal stability check was introduced. This involved extracting multiple consecutive frames (e.g., continuous...). In a frame, pixel blocks with a reflection confidence score greater than a constant threshold are traced using optical flow to determine their centroid coordinates and calculate their spatial displacement vectors. The specific formula for calculating the displacement vector magnitude is as follows: ,in, The average displacement vector magnitude over the time period is expressed in pixels per frame. For the first The two-dimensional coordinates of the corresponding pixel in the frame. This allows for the identification of the spatially distributed displacement vector within a preset time period. Pixels falling below a preset drift threshold are designated as high-frequency reflection regions. The preset time period is set to a continuous observation window of 30 to 60 frames (approximately 1 to 2 seconds), and the preset drift threshold is set to an adjustable range of 0.5 pixels / frame to 2.0 pixels / frame. To tolerate minor structural vibrations of the building's glass curtain wall caused by wind loads in the natural environment, and to completely filter out transient glare interference from fast-moving vehicle windows, a baseline drift threshold of 1.5 pixels / frame is preferred in actual engineering deployments.

[0027] Finally, the system reads the intrinsic parameter matrix of the cameras at the current candidate placement locations from the underlying hardware interface. (Including focal length and optical center parameters) and the extrinsic rotation matrix characterizing the camera's spatial attitude. With translation vector Combining depth map information attached to spatial distribution feature data For the boundary pixel coordinates within the high-frequency reflection region Perform the inverse projection transformation, the calculation formula is as follows: ,in, The distance values ​​provided for the depth map, in meters, are compared with the translation vector. The units are consistent. Based on the derived three-dimensional boundary coordinates... A smooth polygonal geometric boundary is constructed and reprojected forward onto the 2D pixel plane of the current viewpoint. A polygonal region filling algorithm is used to generate reflection mask data consisting of a binary pixel matrix of 0s and 1s. The system performs a bitwise logical NOT operation to obtain the complement matrix of the mask, and performs a Hadamard product (element-wise multiplication) operation on the complement matrix with each frame of the original synchronous video stream. This completely removes the pixel data corresponding to the reflection mask data, and outputs only the non-masked effective area of ​​the real non-reflective physical environment. This avoids data source pollution caused by mirroring for subsequent static feature extraction and cross-camera collaboration.

[0028] Furthermore, the matching correlation degree between the static feature point set and the corresponding images at different candidate point locations is calculated, and a cross-camera field-of-view correlation map is constructed based on the matching correlation degree, including: Timestamp alignment is performed on synchronized video streams with different candidate placement locations, and local descriptors of static feature point sets are extracted between frames with consistent timestamps; Calculate the Euclidean distance between local descriptors of different frames and establish initial feature point matching pairs; The epipolar geometric constraint algorithm is used to remove mismatched points from the initial feature point matching pairs to obtain the set of valid matching points; The density distribution values ​​of the effective matching point set within the preset grid area are statistically analyzed. Grid areas with density distribution values ​​greater than the preset density threshold are identified as field-of-view association nodes. Field-of-view association nodes are connected to generate a cross-camera field-of-view association map.

[0029] In a specific embodiment, the synchronous video streams acquired from different candidate locations are first aligned with nanosecond-level timestamps based on the network time protocol to eliminate network latency differences when multiple terminals are streaming. Then, between reference frames with strictly consistent timestamps, a scale-invariant feature transformation algorithm is used to extract local descriptors of static feature point sets for the non-masked effective area. Each feature point will be encoded as a 128-dimensional vector sequence composed of the gradients of the surrounding local images and its scale space level.

[0030] Next, based on the extracted local descriptors, the feature distance between different frames is calculated to establish initial feature point matching pairs. To overcome the scale scaling interference caused by changes in object distance in traditional multi-view scenarios, this embodiment uses a modified weighted Euclidean distance formula for matching measurement. The calculation formula is as follows: ,in, The weighted matching distance is the feature point between the feature points in the image of camera A and the feature points in the image of camera B. and These are the first two feature points in the 128-dimensional local descriptor. Each component value; and These are the scale level numbers where the feature points were extracted in their respective Gaussian difference pyramids; The preset scale sensitivity penalty coefficient ranges from 0.2 to 0.8, with 0.5 being preferred. The necessity of this variant formula lies in the fact that traditional pure Euclidean distance does not consider the abrupt changes in feature scale caused by differences in physical depth of field across cameras. By introducing an exponential penalty term based on scale hierarchy differences, it can significantly suppress pseudo-similar points with similar textures but significantly mismatched scales in 3D physical space; the system will... Feature point pairs below the nearest neighbor ratio threshold are retained, and initial feature point matching pairs are established.

[0031] Subsequently, to eliminate mismatches caused by periodically repeating textures in the environment from the initial feature point matching pairs, the system uses a random sampling consensus algorithm combined with epipolar geometric constraints for spatial verification. Specifically, the pixel coordinates of all feature points are first transformed into dimensionless normalized coordinates with values ​​ranging from [-1, 1] based on their respective camera resolution matrices. and Then, the fundamental matrix between the two views is solved. And calculate the symmetric epipolar geometric error cost for each matching pair, using the following formula: ,in, This represents the vertical geometric deviation of a spatial point projected onto the epipolar line. Since the coordinates have been pre-normalized and the fundamental matrix elements are dimensionless scaling factors, this epipolar geometric error is also a dimensionless parameter. The system will... Matching pairs that exceed the preset tolerance range are discarded, thereby obtaining a set of valid matching points that fully comply with the rigid body transformation constraints in three-dimensional space. The preset tolerance range is set to 0.01 to 0.05, preferably 0.02.

[0032] Finally, the system divides the 2D image into multiple uniform grids of fixed size and calculates the spatial density of valid matching points within these predefined grid regions. To prevent absolute quantity deception caused by large-area grids and the cumulative error of low-quality matching points, the algorithm calculates a weighted density distribution value, using the following formula: ,in, This represents the density distribution value for this grid region; This represents the total number of valid matching points falling within this grid. The physical pixel area of ​​a single grid; This represents the average polar geometric error of all valid matching points within the grid. The maximum allowable error is fixed at 1.0. The system will calculate... The density distribution values ​​are compared with a preset density threshold, and grid regions with density distribution values ​​greater than the preset density threshold are established as field-of-view association nodes. The preset density threshold is set to an adjustable range of 0.005 to 0.05 effective feature points per square pixel. To avoid interference from accidental noise in feature-poor areas such as white walls in the image, and to ensure that structured regions that can truly represent overlapping fields of view are extracted, a parameter threshold of 0.02 per square pixel is preferred. The system connects the different field-of-view association nodes extracted across cameras as topological edges, ultimately generating a cross-camera field-of-view association map representing the spatial overlap relationship of multiple cameras, providing a high-confidence data topology foundation for subsequent mirror ghost verification.

[0033] Furthermore, the homography matrix between pairs of sub-regions to be verified is calculated, and the determinant sign of the homography matrix is ​​extracted to determine the geometric symmetry features, including: Obtain the set of edge contour coordinate points within the two regions of the sub-region to be verified; The spatial mapping relationship between two edge contour coordinate point sets is calculated using the singular value decomposition method, and the third-order homography matrix is ​​obtained by solving. Calculate the eigenvalues ​​of the homography matrix and solve for the product of the eigenvalues ​​to obtain the determinant value; When the determinant value is negative and the absolute value of the determinant value is within the preset deformation tolerance range, it is determined that the sub-region to be verified satisfies the spatial flip relationship and generates geometric symmetry features.

[0034] In a specific embodiment, for the aforementioned extracted pair of sub-regions to be verified, a spatial transformation mapping verification is first initiated to determine whether the two image regions constitute a mirror image symmetry between the physical reality and the optical reflective surface. This process first extracts continuous structured contours in both regions of the pair using the Canny subpixel edge detection operator, obtaining the edge contour coordinate point sets within each region of the pair. To eliminate pixel size differences caused by different camera intrinsic parameters, the extracted coordinate point sets are multiplied by the corresponding camera's intrinsic parameter inverse matrix, transforming them into homogeneous coordinate vectors on the normalized projection plane. and .

[0035] Next, the spatial mapping relationship between the coordinate point sets of the two edge contours is calculated using the singular value decomposition method, and the third-order homography matrix is ​​obtained. Therefore, this embodiment adopts a spatial mapping variant function weighted by edge gradient information entropy, and the fitting formula is as follows: .in, Let be the third-order homography matrix to be determined. Since its function is to map the normalized space plane to another normalized plane, all elements in its matrix are dimensionless projection scaling coefficients. and These are the normalized homogeneous edge coordinate vectors of the real region and the suspected mirror region, respectively. Their dimensions are transformed into dimensionless pure numerical values ​​due to the elimination effect of the intrinsic parameter matrix. The entropy weight of local edge information, calculated based on the gray-level co-occurrence matrix of the image, represents the clarity and credibility of the contour at this coordinate point and is a dimensionless normalized scalar. Let be the weighted mapping error cost function. Expand the optimization problem and transform it into a homogeneous system of linear equations. Then, the singular value decomposition method was used to analyze the coefficient matrix. The process involves decomposition, extraction of the right singular vector corresponding to the minimum singular value, and reassembly to obtain the desired third-order homography matrix. The necessity of this variant formula lies in introducing... The weighting ensures that sharp physical edges dominate the matrix solution, significantly suppressing edge distortion interference caused by water stains or dirt on the glass surface.

[0036] Subsequently, the system calculates the homography matrix. Perform spectral decomposition and calculate the eigenvalues ​​of the homography matrix. The determinant is obtained by multiplying these three eigenvalues ​​together, and the formula is as follows: From the fundamental theory of physical geometric transformations, the absolute value of the determinant... It represents the area perspective scaling ratio between two regions, while the sign of the determinant value directly reflects the spatial topological rotation: if it is positive, it means that only translation, rotation or normal perspective scaling has occurred; if it is negative, it proves that the two-dimensional plane data has undergone a strict spatial flip phenomenon, which is the inherent optical characteristic formed by light reflecting through a plane mirror in the physical world.

[0037] Finally, the system performs a comprehensive verification based on the calculation results, and when the determinant value... When the determinant is negative and its absolute value is within the preset deformation tolerance range, the sub-region pair to be verified is determined to satisfy the spatial flip relationship. This sub-region pair is then marked as "true" in the memory stack, generating geometric symmetry feature data as a prerequisite for subsequent verification. The preset deformation tolerance range is set to an adjustable ratio range of [0.4, 2.5]. To avoid algorithmic misjudgments caused by extreme stretching of the perspective area due to extreme camera tilt, and to ensure that it can encompass normal optical deformation caused by reflections from typical building glass curtain walls (slight distortion caused by non-ideal planes), a ratio range of [0.75, 1.35] is recommended as the deformation judgment benchmark, allowing a maximum perspective scaling deviation of approximately 35% between the mirrored area and the actual area. Thus, the system accurately identifies and confirms optical mirroring phenomena in the physical scene through pure mathematical singular value decomposition and algebraic feature extraction.

[0038] Furthermore, dynamic time warping calculations are performed on the dynamic target trajectory sequence within the sub-region to be verified to obtain trajectory mirror synchronization data, including: The dynamic target trajectory sequence is decomposed into a spatiotemporal state vector containing velocity and direction components; Perform a mirror flip transformation on the spatiotemporal state vector of one region in the sub-region pair to be verified to generate a reference state vector; The cumulative distance matrix between the spatiotemporal state vector of another region and the reference state vector is calculated using the dynamic time warping algorithm. Extract the cost value of the shortest twisted path in the cumulative distance matrix, and normalize the cost value to obtain trajectory mirror synchronization data.

[0039] In one specific embodiment, to further confirm whether the geometrically symmetric regions initially screened by the homography matrix possess rigorous physical coherence in temporal dynamics, the system initiates a dynamic time warping verification process based on kinematic features for the dynamic target trajectory sequence within the sub-region to be verified. This process first utilizes a target tracking algorithm to obtain the pixel coordinates of each frame in the trajectory sequence and transforms the discrete two-dimensional coordinates into a spatiotemporal state vector containing velocity and orientation components. Specifically, for the first... Frame coordinates Calculate its lateral velocity component. With longitudinal velocity component This generates a spatiotemporal state vector. Since the time difference between two adjacent frames (usually one frame) is used as the differential base, this spatiotemporal state vector... The physical dimension is strictly defined as pixels per frame.

[0040] Next, the system extracts the spatiotemporal state vector sequence of the potential physical source region (set as region A) from the sub-region pair to be verified. Perform a mirror flip transformation on it to match the physical reflective surface to generate a reference state vector sequence. This transformation is achieved by left-multiplying by a two-dimensional reflection-rotation matrix. To achieve this, that is ,in, It is an orthogonal matrix extracted from the reflection axis of the homography matrix calculated in the previous step, and its internal elements are all dimensionless geometric transformation coefficients.

[0041] Subsequently, in order to overcome the frame rate instability and slight timestamp offset problems caused by network jitter during multi-camera network video streaming, a modified dynamic time warping algorithm was used to calculate the spatiotemporal state vector of another region (suspected mirror region B). With the generated reference state vector The cumulative distance matrix between them. Traditional DTW algorithms rely solely on absolute distance for measurement, which can easily lead to mismatches when trajectories intersect. This embodiment introduces a variant node cost function based on direction cosine with exponential penalty, and the fitting formula is: . In the formula, Twisting the mesh in time for two state vectors The cost of a single-step matching at a given point; The second norm of the difference between the two velocity vectors (i.e., the Euclidean distance); inside the parentheses of the exponential function on the right side of the formula, the numerator is the dot product of the two vectors, and the denominator is the product of the magnitudes of the two vectors. The quotient of the two is the dimensionless scalar representing the cosine of the velocity angle. To prevent extremely small positive numbers from being divided by zero; This is a preset penalty coefficient for directional deviation. The necessity of this variant formula lies in the fact that the reverse coordination of mirror motion in the physical world is extremely strict. However, due to different camera distances, the absolute magnitude of velocity may have scaling deviations. Introducing this exponential penalty term can impose an exponentially increasing penalty on matching points with inconsistent directions, while allowing for a certain perspective error in velocity magnitude, thus significantly improving the anti-interference capability.

[0042] Finally, the system uses the cumulative distance matrix By backtracking to the endpoint, the total value on the shortest twisted path is extracted. The value of this time factor is then normalized to obtain the trajectory mirror synchronization data. The normalization calculation formula is as follows: .in, The final output trajectory mirror synchronization data has a value range that strictly converges to 1. Between these, there is a dimensionless scoring index; The total number of aligned nodes in the shortest twisted path is denoted by , which is a dimensionless counter variable. This is a preset baseline speed tolerance constant, measured in pixels per frame. The preset directional deviation penalty coefficient is also included. The adjustable range is set to [1.0, 5.0], with a preferred value of 2.5 to balance the sensitivity of directional constraints; while the preset reference speed tolerance constant... Based on the typical object movement speed in the monitored scene, the frame rate is set to an adjustable range of 2.0 to 10.0 pixels per frame. To avoid excessive amplification of minute noise when the target is stationary, and to ensure tolerance for edge speed attenuation caused by camera lens distortion, 5.0 pixels per frame is preferred in practical engineering applications. When the output trajectory mirror synchronization data... When the value approaches 1, it provides irrefutable kinematic evidence that the dynamic target in region B is merely a physical optical projection of the target in region A, thus completing the closed-loop verification of the mirror ghost image from geometric to spatiotemporal dimensions.

[0043] Furthermore, the optical characteristic parameters of the sub-regions to be verified are extracted, and the physical characteristic deviations, including brightness attenuation and contrast loss values, are calculated, including: The image format of the sub-region pair to be verified is converted into color space data, and the luminance channel component and chrominance channel component in the color space data are separated. Calculate the global mean of the two regions in the luminance channel component of the sub-region to be verified respectively, and subtract the two global means to obtain the luminance attenuation. The variance data of the two regions on the grayscale histogram is calculated, and the contrast loss value is obtained by quotienting the two variance data. Then, the overall physical characteristic deviation is calculated based on the brightness attenuation and the contrast loss value.

[0044] In one specific embodiment, in order to perform the final verification of the physical authenticity of the region pairs that have undergone initial screening based on geometry and kinematics at the optical level, the system initiates a process of extracting optical feature parameters and calculating physical property deviations for the sub-region pairs to be verified. This process first extracts the original RGB color image matrix of the sub-region pairs to be verified (including region A, which serves as the physical source, and region B, which is suspected to be an optical mirror image), and then uses a color space conversion matrix to convert the original image format into YUV color space data.

[0045] Subsequently, in order to eliminate the interference caused by color temperature changes and color shift of the imaging sensor in complex environments, the system stripped and discarded the U and V chromaticity channel components that are not sensitive to changes in light intensity, and extracted only the Y luminance channel component data that can intuitively represent the intensity of light energy reflection. The value of each pixel in this component is precisely defined as gray level data with a value in the range of 0 to 255, and its physical dimension is defined as gray level.

[0046] Next, the system iterates through all valid pixels in regions A and B respectively, sums their Y-channel pixel values ​​and divides them by the total number of pixels in each region, thereby calculating the global mean of region A in the luminance channel component. and the global mean of region B Based on the objective law in the physical world that light waves inevitably undergo energy transmission, absorption, and scattering when penetrating and colliding with interfaces of different media, the brightness of a reflected image will inevitably be lower than that of a directly incident image. Therefore, the system will use the global average value of region A. Subtract the global mean of region B This allows us to obtain the absolute brightness decay, which characterizes energy loss. The calculation formula is: Meanwhile, the microscopic roughness of the surface of physical reflective media (such as glass and water surfaces) and the diffuse reflection effect of internal impurities inevitably lead to the blurring of high-frequency edge details within the reflected image. Based on this, the system separately statistically analyzes the brightness variance data of region A and region B on the Y-channel grayscale histogram. and Then the variance data for region B will be... Divide by the variance data of region A Perform a quotient calculation to obtain the contrast loss value. Finally, because surveillance cameras from different manufacturers have built-in automatic exposure compensation (AE) and wide dynamic range (WDR) algorithms, the global absolute brightness of the image can fluctuate dramatically throughout the day. Therefore, the system introduces a weighted Euclidean deviation variant model based on a relative attenuation benchmark to calculate the overall physical characteristic deviation. The fitting formula is . In the formula, Represents the relative brightness decay rate; The preset empirical relative brightness attenuation benchmark; Preserve the baseline for the preset empirical contrast. and These are the preset normalized importance weight coefficients for the corresponding features, and satisfy the following conditions: The necessity and significant advantage of this variant formula lies in the fact that, compared to traditional methods for measuring absolute Euclidean distance, this formula divides by the source mean. By converting absolute loss into a relative loss ratio, the interference of drastic changes in light intensity between day and night on the judgment results is completely eliminated. Simultaneously, prior anchor points conforming to the typical optical attenuation laws of glass curtain walls are incorporated, greatly improving the model's robustness under complex lighting conditions. During this process, a pre-set empirical relative brightness attenuation baseline is used. The adjustable range of 0.15 to 0.45 is used to avoid baseline failure caused by overall darkening of the image in low-light nighttime scenes, while ensuring full conformity with the Fresnel reflection characteristics of modern architectural coated glass; the preset empirical contrast retention baseline is set to 0.30. The preferred value is 0.65; weighting coefficient and The optimal values ​​are 0.4 and 0.6, respectively, to appropriately emphasize the contrast characteristics with stronger noise resistance. The system finally outputs the calculated dimensionless index. The deviation is compared with the preset physical loss tolerance threshold. If the deviation is at an extremely low level, it confirms from the fundamental level of physical optics that the region fully conforms to the energy loss characteristics of physical specular reflection, thus providing irrefutable data for the subsequent degradation of camera placement value or utilization of implicit gains.

[0047] Furthermore, the conditional information entropy of the mirror region relative to the real region in the object-mirror pair is calculated, including: Extract the pixel gray-level co-occurrence matrix of the real area and the mirror area respectively, and calculate the edge information entropy of the real area and the joint information entropy of the mirror area; Subtracting the edge information entropy of the physical region from the joint information entropy yields the independent information increment in the mirror region, which is independent of the physical region. The conditional information entropy is calculated by combining independent information increments with the resolution parameters of cameras at candidate locations.

[0048] In one specific embodiment, after the system successfully marks a sub-region to be verified as a true object-mirror pair through prior geometric transformation and optical verification, in order to determine whether the mirror image is purely visual redundancy or an effective supplementary monitoring feature that reflects blind spots behind the entity, the system then initiates a value reassessment logic. This logic aims to calculate the conditional information entropy of the mirror region relative to the object region in the object-mirror pair. This process first performs spatial inverse projection of the mirror region based on the calculated homography matrix, aligning it pixel-by-pixel with the object region in the coordinate system.

[0049] Subsequently, the system extracts the image texture features of the real object region and the aligned mirror region respectively. By statistically analyzing the occurrence frequency of gray values ​​of adjacent pixels in the local space, it constructs a one-dimensional normalized gray-level co-occurrence probability distribution sequence for the real object region. and the two-dimensional joint gray-level co-occurrence probability matrix of the object-mirror aligned region. ,in, and These represent the gray levels of the corresponding pixels. The system calculates the edge information entropy of the object region based on Shannon's information theory. and the joint information entropy of object-mirror pairs .

[0050] Next, the system will include the joint information entropy containing global field of view information. Subtract the entropy of the physical region edge information containing only direct source information. Obtain independent information increments in the mirrored region, separate from the physical region. However, in real-world physical video surveillance deployments, theoretical information that cannot be clearly resolved due to distance attenuation of optical sensors renders it practically useless for deployment. Therefore, the system does not directly use this increment, but instead incorporates a conditional information entropy variation formula based on the hardware physical sampling limit. The specific fitting calculation formula is as follows: . In the formula, The conditional information entropy of the final output; This represents the independent information increment obtained from the previous step; The surface projection resolution is obtained by dividing the total resolution parameters of the cameras at the candidate locations by the actual physical area of ​​the mirror surface estimated through the depth map. This is a preset basic recognition resolution threshold. The necessity and significant technical advantage of this variant formula lies in the fact that traditional information entropy algorithms are freed from the constraints of physical hardware and are prone to misclassifying extremely distant, blurry reflection artifacts with subtle texture differences as high-information-value areas; while this variant, by introducing... The exponential soft truncation function with kernel, utilizing the actual sampling resolution Compared with the benchmark identification threshold The ratio is used as a control variable; when the distance between the mirror surfaces is extremely large, it leads to... When the value approaches 0, the exponent term approaches 1, and the entire value within the parentheses decays to 0, thus forcibly erasing information increments that cannot be effectively identified by the monitored hardware; when the sampling rate is sufficient, it approaches 1 to fully retain its information value. In this specific operation, the preset basic identification resolution threshold... The adjustable range of 150 pixels / square meter to 500 pixels / square meter is set. To avoid low-resolution mosaic data being incorrectly included in the effective information gain, and to ensure that the minimum requirement for recognizing the approximate outline of a human body is met in standard security monitoring, a setting of 250 pixels / square meter is preferred for actual engineering deployments. The system then calculates the conditional information entropy. Then, based on this value, it can be accurately determined whether the mirrored area belongs to visual redundancy that can be eliminated, or whether it should be included in the global topology map and transformed into a virtual gain field of view that attracts camera placement resources to approach, providing a solid and reliable data-driven foundation for the entire automated placement algorithm.

[0051] Furthermore, a spatial potential energy field containing both repulsive and attractive force fields is generated, including: Extract the global 3D coordinates of each candidate point location in the 3D spatial map, and establish a 3D influence sphere with the global 3D coordinates as the origin; The geometric center of the mirrored region after zeroing is defined in the 3D spatial map as the repulsive force source point. The reciprocal of the distance from the repulsive force source point to each coordinate point in the 3D influence sphere is calculated to generate a repulsive force field that decays with a distance gradient. The intersection of the extensions of the normals of the virtual gain field of view in the 3D spatial map is defined as the attraction source point. The spatial curvature corresponding to the attraction source point is calculated, and an attraction field with a logarithmic distribution is generated based on the spatial curvature.

[0052] In one specific embodiment, after revaluing the object-mirror pair, the system transforms the extracted spatial properties of the reflective medium into dynamic constraints to guide optimal placement. By generating a spatial potential energy field containing both repulsive and attractive force fields, it achieves automated logical guidance for camera positioning. This process first extracts the global 3D coordinates of each candidate placement location from the 3D spatial map. With each global 3D coordinate as the center of a sphere, establish a sphere with a radius of... A three-dimensional influence sphere is defined, which defines the spatial threshold at which potential energy exerts a substantial pulling effect on the current position of the camera; wherein, the preset radius of the sphere... The adjustable range is set from 3 meters to 15 meters. To ensure that the force field can cover the local layout adjustment space while avoiding global calculation overload, it is preferred to set it to 8 meters.

[0053] Next, the system defines the geometric center of the mirrored region that underwent zeroing processing in the previous step as the repulsive force source point in the three-dimensional spatial map. To quantify the degree of repulsion of redundant mirror images on the placement points, the system calculates the reciprocal of the distance from the repulsive force source point to each coordinate point within the three-dimensional influence sphere, generating a repulsive force field that decays with a distance gradient. Its potential energy distribution formula is fitted as follows: . In the formula, Spatial coordinates The repulsive potential energy at the point; The preset exclusion constraint coefficient is used to adjust the squeezing intensity of the redundant area on the equipment distribution; The preset reference length is 1.0 meter. The necessity of introducing this parameter is to standardize and align the physical distance dimension of the denominator, thereby ensuring that the potential energy value output is a dimensionless scalar. Position the camera to the 1st The Euclidean distance between the repulsive force source points; This is a preset distance smoothing factor used to prevent the algorithm from crashing when the potential energy tends to infinity when the camera is extremely close to the source point.

[0054] Subsequently, for the region marked as the virtual gain field of view, the system defines the intersection of the extensions of the normals of this field of view in the 3D spatial map (i.e., the convergence center of reflected rays in virtual space) as the attraction source point. The spatial curvature corresponding to the source point is calculated, and a logarithmically distributed attractive field is generated based on the spatial curvature. The formula for calculating the attractive potential energy is as follows: . In the formula, for The attractive potential energy value at point; The preset gain guide weight; For the first The spatial curvature index corresponding to each attraction source point is a dimensionless constant, which is obtained by calculating the ratio of the angle subtended by the boundary of the reflecting surface in three-dimensional space to the radius of curvature, and characterizes the breadth contribution of the virtual field of view to the coverage of the blind zone. It is the natural logarithm function, whose internal parameter is 1 plus the reference length. The ratio to the distance term (m) is a dimensionless pure number. The advantage of this formula through logarithmic variation is that it simulates the nonlinear traction characteristics of "gradual increase in attraction at long distances and rapid lock-up at close distances," forcing the optimization algorithm to prioritize the point coordinates that can maximize the use of the virtual gain angle.

[0055] During this operation, a preset distance smoothing factor is used. It is recommended to set the distance to 0.1 meters to 0.5 meters, preferably 0.2 meters, to ensure the continuity of the force field gradient during fine-tuning of the point layout; the preset repulsion constraint coefficient. The density of redundant reflective surfaces within the scene is adjusted within the range of [5.0, 20.0], with a preferred value of 10.0. Finally, the system generates a comprehensive spatial potential field for the distributed points through superposition calculations. The integrated potential energy field forms a virtual energy manifold with high and low undulations. The high potential energy region attracts the camera to land, while the low potential energy region (the region dominated by repulsion) drives the camera away. This seamlessly transforms the complex image recognition results into environmental boundary constraints that can be directly called by the optimization algorithm.

[0056] Furthermore, using the spatial potential energy field of the point layout as the environmental constraint boundary, iterative calculations are performed using a point layout optimization algorithm, including: The camera's physical coordinate parameters and the gimbal's attitude parameters are encoded into a population of particles, and the position and velocity states of the population of particles are randomly initialized in a 3D spatial map. The repulsive field is added as a penalty coefficient to the fitness function of the population particles, and the attractive field is added as a reward coefficient to the fitness function. The current fitness value of each particle in the population is calculated using the fitness function, and the global optimal position and local optimal position of the particles in the population are updated based on the current fitness value. Adjust the position and velocity states based on the global and local optimal positions until the current fitness value of the population particles converges, and output the camera physical coordinate parameters and gimbal attitude parameters corresponding to the population particles in the converged state.

[0057] In one specific embodiment, the system uses the constructed spatial potential energy field of the deployment points as a nonlinear environmental constraint boundary, and employs an improved particle swarm optimization algorithm to perform iterative calculations across multiple spatial dimensions to find the globally optimal device deployment configuration. This process first involves determining the physical coordinate parameters of the cameras to be solved (including their three-dimensional spatial positions). ) and gimbal attitude parameters (including horizontal yaw and pitch angles) The joint encoding is the generalized position-state vector of the i-th population particle. Simultaneously, the initial position and state vectors of the population particles are randomly initialized within the physically feasible region of the three-dimensional spatial map. and the corresponding generalized velocity state vector (where position velocity) The unit of measurement is meters per revolution, and the attitude velocity is... The dimension is radians / time, where "time" represents the number of iterations, which is a dimensionless quantity.

[0058] Next, the system constructs a target fitness function for each particle that integrates multi-dimensional environmental feedback. Traditional point optimization algorithms often only use spatial frustum coverage as a single evaluation metric, which can easily lead to the algorithm getting stuck in local false optima in the mirror reflection region. Therefore, this embodiment constructs a comprehensive potential energy field by superimposing the repulsive force field and the attractive force field obtained in the previous step. An evaluation system is introduced, and an exponential multiplication variation is used to fit the comprehensive fitness function. The calculation formula is as follows: In this formula, Let be the current overall fitness value of the i-th particle, and be the dimensionless evaluation score. The total volume of non-occluded target voxels that can be effectively covered by the camera's view frustum in the 3D spatial map corresponding to the particle decoding is calculated by the system after performing voxel collision detection on the map using a ray casting algorithm. The total required voxel volume for the core area of ​​interest pre-defined in this monitoring scenario; This is a preset potential energy sensitivity adjustment factor; This represents the combined potential energy value at the particle's coordinates. The extreme technical necessity of this variant formula lies in: using division operations... Obtain the basic physical space coverage ratio (a dimensionless number in the [0,1] interval), and then use the natural exponential function. The potential energy field is transformed into a nonlinear reward / penalty coefficient—when a particle mistakenly enters the repulsive region of high-frequency reflection ( The exponential term exhibits a precipitous decline, forcibly reducing its fitness and compelling it to migrate out; when the particle approaches the virtual gain attraction region of the reflective blind zone ( If the index is greater than 1, an excess reward is given for the basic coverage rate.

[0059] Subsequently, the system uses the fitness function to calculate the fitness value of each particle in the current iteration, compares the calculation result with the historical extreme values, and updates the individual optimal state extreme value experienced by the particle in real time based on the current fitness value. and the globally optimal state extreme value shared by the entire population. Based on extreme value guidance, the system dynamically adjusts the position and velocity states of the next generation of particles using variations of Newtonian kinematics. Taking position and velocity updates as an example, the recursive calculation formula is as follows: The position update formula is: . in the formula This is the inertia weighting coefficient. and These are the individual cognitive constant and the social learning constant, respectively. and The coefficients are random distribution coefficients between [0,1], and all of the above are standard dimensionless parameters; the position coordinate differences within parentheses are... The dimensionless coefficient multiplied by the superposition of the units is meters. The unit of measurement is maintained at meters per iteration, and this is strictly corresponded in the position update addition, achieving stable displacement search in three-dimensional space. This iterative process continues to loop until either of the following termination criteria is met: the variance of the current global optimal fitness value of the population particle over N consecutive generations is lower than the convergence threshold, or the maximum allowed number of iterations is reached. At this point, the system determines that the algorithm has reached steady-state convergence, and then extracts the global optimal particle in the converged state. It performs a reverse decoupling operation on the camera, outputting the camera's physical coordinate parameters and corresponding gimbal attitude parameters that avoid the repulsive force field and accurately approach the attractive force field. In a specific application embodiment, a preset potential energy sensitivity adjustment factor is used. The value is set to an adjustable range of 0.5 to 2.0. To avoid the algorithm ignoring the basic blind zone coverage due to excessive potential energy field weight, the preferred value is 0.8. The maximum number of iterations is recommended to be set to 200 generations.

[0060] Furthermore, after outputting the camera's physical coordinate parameters and the gimbal's attitude parameters, the process further includes: Update the candidate placement positions according to the camera physical coordinate parameters and gimbal attitude parameters in the converged state, and obtain the secondary synchronous video stream after updating the candidate placement positions; Reconstruct the secondary cross-camera field-of-view correlation graph using the secondary synchronous video stream; Determine whether there are any pairs of secondary sub-regions to be verified in the secondary cross-camera field-of-view association graph that are higher than a preset association threshold; When there are secondary sub-region pairs to be verified, the mirror verification loop is retried until there are no physical-mirror pairs that meet the preset mirror judgment conditions in the secondary cross-camera field of view association graph.

[0061] In one specific embodiment, after the system outputs the globally optimal camera physical coordinate parameters and gimbal attitude parameters using the spatial potential energy field of the deployment points, in order to prevent the exposure of previously obscured reflective media due to nonlinear changes in the viewing angle after the device moves in the real physical space, or to prevent new image distortion caused by slight deviations between the actual and calculated placement due to mechanical backlash in physical actuators (such as motor stepper gears), the system further constructs a closed-loop self-correction verification mechanism based on physical execution feedback. This process first sends the converged camera physical coordinate parameters and gimbal attitude parameters to the motor drive modules of each front-end monitoring device through the IoT underlying control protocol. After the system delays and waits for the motor to complete its mechanical action and for the image stabilization algorithm to stabilize, it captures the secondary synchronized video stream updated with the candidate deployment point positions according to a preset frame synchronization strategy.

[0062] Next, the system re-executes the aforementioned feature extraction and spatial grid density statistics on the acquired secondary synchronous video stream to reconstruct the secondary cross-camera field-of-view association map using the underlying pixel-level data. During this process, traditional static association threshold determination methods, after multiple rounds of position fine-tuning, are prone to getting stuck in an infinite loop of "Zeno's paradox" deadlock iterations due to extremely small areas of weak reflection in the image (such as reflections from metal doorknobs). Therefore, this embodiment creatively introduces a dynamic exponential penalty association threshold variation formula based on iteration depth when determining whether there are secondary sub-region pairs in the secondary cross-camera field-of-view association map that exceed a preset association threshold. The specific calculation formula is as follows: . In the formula, is the secondary dynamic correlation threshold at the k-th closed-loop correction iteration, and is the dimensionless normalized score value; k is the iteration loop number that triggers the secondary verification, with an initial value of 1, which is a dimensionless counting integer. The preset initial basic correlation threshold is a dimensionless pure number, ranging from 0.5 to 0.7, preferably 0.6; The preset limit tolerance correlation threshold is a dimensionless pure number, ranging from 0.9 to 0.99, preferably 0.95; The preset iterative decay penalty coefficient is a dimensionless parameter, ranging from 0.3 to 0.8, with 0.5 being preferred. The necessity and significant advantage of this variant formula lies in the fact that as the number of physical adjustments k increases, the system... This non-linear increasing factor smoothly and forcefully raises the correlation threshold for triggering mirror verification, enabling the system to maintain extremely high sensitivity to large-area glass curtain walls during initial placement. In subsequent fine-tuning, it can automatically ignore microscopic residual reflections that are extremely sensitive to viewing angles, have very small physical surfaces, and do not affect the overall placement quality. This mathematically eliminates the potential for hardware deadlock that could prevent algorithm convergence. The system, based on the calculated... Traverse the edge weights in the secondary cross-camera field of view association graph and select node pairs with weights higher than the threshold as secondary sub-region pairs to be verified.

[0063] Subsequently, when the system determines the existence of secondary sub-region pairs to be verified, it will re-trigger the aforementioned complete mirror verification loop, which includes homography matrix solving, dynamic time warping verification, and physical optical deviation extraction. To quantify the global mirror elimination degree and use it as the final termination condition, the system constructs a system-level mirror residual energy assessment model based on the results of the secondary verification loop. The calculation formula is as follows: . In the formula, represents the total residual energy of the system image under the k-th iteration, and is a dimensionless evaluation index; The total number of currently identified secondary sub-regions to be verified, dimensionless; is the geometric flip feature identifier of the j-th pair of regions. If the determinant of its homography matrix is ​​negative, it takes the value 1, otherwise it is 0. It is a dimensionless constant. To utilize the trajectory mirror synchronization data of the j-th pair of output regions, the values ​​range from (0,1] and are dimensionless; The deviation of the overall physical properties of the j-th pair of regions calculated above is dimensionless; The preset optical confidence weighting amplification factor is a dimensionless coefficient ranging from 0.5 to 2.0, preferably 1.0. This formula deeply integrates the absolute rigidity constraint of geometric determination, the time coordination constraint of kinematics, and the characteristics of optical energy loss at a multiplicative level. Only when all three simultaneously and highly conform to the objective physical laws of optical reflection will a significant image energy value be generated. The system will calculate the... With respect to the preset system steady-state convergence minimum Perform a comparison, if The relabeled object-mirror pair is then transformed into a new repulsive and attractive force field, and the point optimization iteration is restarted again with the current coordinates as the local starting point; until a certain cycle is completed. If there are no physical-mirror pairs that meet the preset mirror judgment conditions in the secondary cross-camera field of view association graph, the system locks and outputs the final hardware coordinates, completing the automated deployment of the entire life cycle.

[0064] The above content is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described or use similar methods to replace them, as long as they do not deviate from the concept of the invention, they should all fall within the protection scope of the present invention.

[0065] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0066] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention.

Claims

1. A method for automatically determining camera placement locations based on video analysis, characterized in that, Includes the following steps: Acquire synchronous video streams of multiple candidate placement points, process pixel distribution information in synchronous video streams, generate reflection mask data corresponding to the reflective medium region, and delineate the non-masked effective region in synchronous video streams in reverse based on the reflection mask data. Extract the set of static feature points and the sequence of dynamic target trajectories within the effective area of ​​the non-masked area. Calculate the matching correlation degree between the static feature point set and the images corresponding to different candidate point locations. Construct a cross-camera field-of-view correlation map based on the matching correlation degree. Extract the node pairs in the cross-camera field-of-view correlation map with matching correlation degree higher than the preset correlation threshold as sub-region pairs to be verified. Trigger a mirror verification loop for the sub-region pair to be verified. The mirror verification loop includes: calculating the homography matrix between the sub-region pair to be verified, extracting the determinant sign of the homography matrix to determine the geometric symmetry features; performing dynamic time warping calculation on the dynamic target trajectory sequence within the sub-region pair to be verified, and obtaining trajectory mirror synchronization data; extracting the optical feature parameters of the sub-region pair to be verified, and calculating the physical characteristic deviation including brightness attenuation and contrast loss values. When the geometric symmetry features, trajectory mirror synchronization data, and physical characteristic deviations all meet the preset mirror judgment conditions, the sub-region pair to be verified is marked as a physical-mirror pair, and the value reassessment logic is initiated. The value reassessment logic includes: calculating the conditional information entropy of the mirror region relative to the physical region in the physical-mirror pair; if the conditional information entropy is lower than the preset redundancy threshold, the placement value weight corresponding to the mirror region is reset to zero; if the conditional information entropy is higher than the preset gain threshold, the mirror region is marked as a virtual gain field of view and the placement value weight corresponding to the mirror region is increased. The reflective medium region after the change of the placement value weight is mapped onto the three-dimensional spatial map to generate a placement spatial potential energy field containing repulsive force field and attractive force field. Among them, the mirror region after zeroing generates a repulsive force field, and the virtual gain field of view generates an attractive force field. Using the spatial potential energy field of the deployment point as the environmental constraint boundary, the deployment point optimization algorithm is used for iterative calculation to output the camera physical coordinate parameters and gimbal attitude parameters that avoid the repulsive force field and tend to the attractive force field.

2. The method for automatically determining camera placement based on video analysis according to claim 1, characterized in that, Generate reflection mask data corresponding to the reflective medium region, and delineate the non-masked effective region in the synchronized video stream based on the reflection mask data, including: Semantic feature extraction is performed on the synchronized video stream to obtain material texture feature data containing material attribute labels and spatial distribution feature data containing planar normal vector information; A preliminary reflection probability map of each frame is constructed using material texture feature data. Spatial distribution feature data is mapped to the preliminary reflection probability map for weight correction, and the reflection confidence score of each pixel in the synchronous video stream is calculated. Extract the spatial distribution displacement vector of the reflection confidence score in multiple consecutive frames, identify the set of pixels whose spatial distribution displacement vector is lower than a preset drift threshold within a preset time period, and use the set of pixels as a high-frequency reflection region. The intrinsic and extrinsic parameters of the cameras at the candidate locations are obtained. Inverse projection transformation is performed in combination with the pixel coordinates of the high-frequency reflection area to determine the geometric boundary of the high-frequency reflection area in the 3D spatial map. Reflection mask data is generated based on the geometric boundary. The non-masked effective area is delineated by removing the pixel data corresponding to the reflection mask data in the synchronous video stream.

3. The method for automatically determining camera placement based on video analysis according to claim 1, characterized in that, Calculate the matching correlation degree between the static feature point set and the corresponding images at different candidate point locations, and construct a cross-camera field-of-view correlation map based on the matching correlation degree, including: Timestamp alignment is performed on synchronized video streams with different candidate placement locations, and local descriptors of static feature point sets are extracted between frames with consistent timestamps; Calculate the Euclidean distance between local descriptors of different frames and establish initial feature point matching pairs; The epipolar geometric constraint algorithm is used to remove mismatched points from the initial feature point matching pairs to obtain the set of valid matching points; The density distribution values ​​of the effective matching point set within the preset grid area are statistically analyzed. Grid areas with density distribution values ​​greater than the preset density threshold are identified as field-of-view association nodes. Field-of-view association nodes are connected to generate a cross-camera field-of-view association map.

4. The method for automatically determining camera placement based on video analysis according to claim 1, characterized in that, Calculate the homography matrix between pairs of sub-regions to be verified, and extract the sign of the determinant of the homography matrix to determine the geometric symmetry features, including: Obtain the set of edge contour coordinate points within the two regions of the sub-region to be verified; The spatial mapping relationship between two edge contour coordinate point sets is calculated using the singular value decomposition method, and the third-order homography matrix is ​​obtained by solving. Calculate the eigenvalues ​​of the homography matrix and solve for the product of the eigenvalues ​​to obtain the determinant value; When the determinant value is negative and the absolute value of the determinant value is within the preset deformation tolerance range, it is determined that the sub-region to be verified satisfies the spatial flip relationship and generates geometric symmetry features.

5. The method for automatically determining camera placement based on video analysis according to claim 1, characterized in that, Dynamic time warping calculations are performed on the dynamic target trajectory sequence within the sub-region to be verified to obtain trajectory mirror synchronization data, including: The dynamic target trajectory sequence is decomposed into a spatiotemporal state vector containing velocity and direction components; Perform a mirror flip transformation on the spatiotemporal state vector of one region in the sub-region pair to be verified to generate a reference state vector; The cumulative distance matrix between the spatiotemporal state vector of another region and the reference state vector is calculated using the dynamic time warping algorithm. Extract the cost value of the shortest twisted path in the cumulative distance matrix, and normalize the cost value to obtain trajectory mirror synchronization data.

6. The method for automatically determining camera placement based on video analysis according to claim 1, characterized in that, Extract the optical characteristic parameters of the sub-region pairs to be verified, and calculate the physical characteristic deviations including brightness attenuation and contrast loss values, including: The image format of the sub-region pair to be verified is converted into color space data, and the luminance channel component and chrominance channel component in the color space data are separated. Calculate the global mean of the two regions in the luminance channel component of the sub-region to be verified respectively, and subtract the two global means to obtain the luminance attenuation. The variance data of the two regions on the grayscale histogram is calculated, and the contrast loss value is obtained by quotienting the two variance data. Then, the overall physical characteristic deviation is calculated based on the brightness attenuation and the contrast loss value.

7. The method for automatically determining camera placement based on video analysis according to claim 1, characterized in that, Calculate the conditional information entropy of the mirror region relative to the real region in a real-mirror pair, including: Extract the pixel gray-level co-occurrence matrix of the real area and the mirror area respectively, and calculate the edge information entropy of the real area and the joint information entropy of the mirror area; Subtracting the edge information entropy of the physical region from the joint information entropy yields the independent information increment in the mirror region, which is independent of the physical region. The conditional information entropy is calculated by combining independent information increments with the resolution parameters of cameras at candidate locations.

8. The method for automatically determining camera placement based on video analysis according to claim 1, characterized in that, Generate a spatial potential energy field containing both repulsive and attractive force fields, including: Extract the global 3D coordinates of each candidate point location in the 3D spatial map, and establish a 3D influence sphere with the global 3D coordinates as the origin; The geometric center of the mirrored region after zeroing is defined in the 3D spatial map as the repulsive force source point. The reciprocal of the distance from the repulsive force source point to each coordinate point in the 3D influence sphere is calculated to generate a repulsive force field that decays with a distance gradient. The intersection of the extensions of the normals of the virtual gain field of view in the 3D spatial map is defined as the attraction source point. The spatial curvature corresponding to the attraction source point is calculated, and an attraction field with a logarithmic distribution is generated based on the spatial curvature.

9. The method for automatically determining camera placement based on video analysis according to claim 1, characterized in that, Using the spatial potential energy field of the points as the environmental constraint boundary, iterative calculations are performed using a point optimization algorithm, including: The camera's physical coordinate parameters and the gimbal's attitude parameters are encoded into a population of particles, and the position and velocity states of the population of particles are randomly initialized in a 3D spatial map. The repulsive field is added as a penalty coefficient to the fitness function of the population particles, and the attractive field is added as a reward coefficient to the fitness function. The current fitness value of each particle in the population is calculated using the fitness function, and the global optimal position and local optimal position of the particles in the population are updated based on the current fitness value. Adjust the position and velocity states based on the global and local optimal positions until the current fitness value of the population particles converges, and output the camera physical coordinate parameters and gimbal attitude parameters corresponding to the population particles in the converged state.

10. The method for automatically determining camera placement based on video analysis according to claim 9, characterized in that, After outputting the camera's physical coordinate parameters and the gimbal's attitude parameters, the following further steps are included: Update the candidate placement positions according to the camera physical coordinate parameters and gimbal attitude parameters in the converged state, and obtain the secondary synchronous video stream after updating the candidate placement positions; Reconstruct the secondary cross-camera field-of-view correlation graph using the secondary synchronous video stream; Determine whether there are any pairs of secondary sub-regions to be verified in the secondary cross-camera field-of-view association graph that are higher than a preset association threshold; When there are secondary sub-region pairs to be verified, the mirror verification loop is retried until there are no physical-mirror pairs that meet the preset mirror judgment conditions in the secondary cross-camera field of view association graph.

11. A computer-readable storage medium, characterized in that: The system stores instructions that, when executed on a computer, cause the computer to perform the automatic camera placement method based on video analysis as described in any one of claims 1 to 10.