Cameras for Heterogeneous Information System Integration Based on the ROSO Model
By integrating cameras into a heterogeneous information system based on the ROSO model, a 3D point cloud model was constructed and reconstructed, solving the problems of blind spots and security risks in the monitoring system, and realizing dynamic adjustment and accurate target tracking.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HEFEI HYDROGEN ANIMAL UNION TECH CO LTD
- Filing Date
- 2026-03-04
- Publication Date
- 2026-05-26
AI Technical Summary
Existing monitoring systems cannot dynamically adjust their monitoring area after the camera port changes, resulting in blind spots and security risks, and making it impossible to achieve targeted tracking and monitoring.
The heterogeneous information system integrated camera based on the ROSO model constructs a point cloud set in the 3D model by recording the depth and optical images of each camera port, performs point cloud reconstruction and label merging, and selects the optimal camera port for target tracking and monitoring network reconstruction.
It enables dynamic adjustment of the monitoring network after the monitoring area is changed at the camera port, reducing blind spots and ensuring accurate tracking of targets and maximum monitoring range.
Smart Images

Figure CN122093531A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of surveillance system technology, specifically to a camera for heterogeneous information system integration based on the ROSO model. Background Technology
[0002] Current surveillance technologies typically employ multi-camera setups for joint monitoring, ensuring wide coverage and no blind spots. It's now possible to detect and track abnormal movements within optical lenses, with brands like Hikvision and InSight automatically identifying unusual activity in monitored areas, such as focusing on individuals entering the area. Furthermore, integrating depth cameras with optical cameras, such as the Face ID-enabled depth sensor on iPhones and Microsoft Azure Kinect, allows for both visual recognition and depth information acquisition. However, differences in systems, protocols, and data exist between brands. When a surveillance system uses multiple brands or types of devices, data conflicts may arise. The ROSO model, as a data processing service and protocol framework, acts as a "protocol adapter" to communicate with various front-end devices. It provides a unified access method to upper layers using standardized service interfaces, resolving data communication issues between different brands and types of devices and facilitating the collaborative construction of larger surveillance networks.
[0003] In the prior art, the technical document with publication number CN115834983A discloses a digital environment monitoring method and system that integrates multi-source information. The method includes: acquiring visible light video information, infrared field information, and acoustic field information through a camera and an acoustic sensor; calibrating the horizontal rotation angle parameters and pitch angle parameters of the camera facing different target objects to obtain a sequence of horizontal rotation angle and pitch angle parameters for different target objects; controlling the camera to adjust the monitoring field of view based on the parameter sequence; constructing a geometric feature 3D model based on the visible light video information using 3D reconstruction technology; and projecting the acquired infrared field information and acoustic field information onto the target environment and target objects corresponding to the geometric feature 3D model to generate a digital feature scene that integrates the infrared field, acoustic field, and geometric features.
[0004] While the publicly available technical documents demonstrate the method of constructing 3D models through a monitoring system, they cannot be used for targeted tracking and monitoring in real-world scenarios. When a camera port changes its monitoring area, this method cannot mobilize other monitoring ports to rebuild the monitoring network, which can easily create blind spots and leave unsafe conditions.
[0005] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this invention is to provide a camera for heterogeneous information system integration based on the ROSO model, so as to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: A camera based on the ROSO model for heterogeneous information system integration, the camera comprising several camera ports, the camera ports communicating with a central control platform via a service interface defined by the ROSO model, the camera ports being electrically connected to camera units, and the central control platform controlling the camera units via the service interface, the control method including: Step 1: Record the depth image of each camera port at each rotation angle, project the depth image into the 3D model to build a point cloud set, and add a label to each point. The label content includes the camera port number and the corresponding rotation angle. Step 2: Reconstruct the point cloud set. Point cloud reconstruction means merging the point clouds at overlapping positions into a single point, and integrating the labels of all points at the overlapping position into a sub-point cloud. Merge the sub-point clouds of all camera ports into a total point cloud in the 3D model, and reconstruct the total point cloud. Step 3: Discover the target through optical images and project the target position into the reconstructed total point cloud. Determine the target label based on the distance relationship between the target position and the point cloud set corresponding to each label in the total point cloud. Move the camera port corresponding to the target label according to the corner marked by the target label, and mark the other camera ports as monitoring ports. Step 4: Select a tag belonging to each monitoring port to form a tag combination, obtain all tag combinations, and determine the abnormal points based on the matching relationship between the point cloud set corresponding to the tag combination in the total point cloud and the point cloud set corresponding to the target tag in the total point cloud. Step 5: Analyze the point cloud and anomalies based on the labels of each monitoring port to generate a combined score for each label combination. Select the label combination with the highest combined score and adjust the angle of each monitoring port according to the number and angle in the label combination.
[0008] Furthermore, depth and optical images of each camera port are acquired separately, the relative position and orientation of each camera port in the real coordinate system are recorded, and all camera ports are numbered.
[0009] Furthermore, the camera port generates a depth image at each unit rotation angle. The pixel information of the depth image is extracted by the central control platform and input into the 3D model. The coordinate system of the 3D model is the same as the real coordinate system, and the X, Y and Z axes adopt the default direction of the 3D model. The pixel information in each depth image is converted into a point cloud set in the 3D model, and each point in the point cloud set is labeled. The label indicates the rotation angle of the camera port when the depth image of that point was generated and the number of that camera port. A coordinate point is selected in the 3D model to represent the position of the camera port in the 3D model. Using this point as the anchor point, the point cloud sets generated by the depth images of all cameras at each angle are aligned according to the camera port's intrinsic parameters and rotation angle.
[0010] Furthermore, the point cloud set formed by the camera port is used for point cloud reconstruction. The point cloud reconstruction logic is as follows: Voxels are divided into all regions containing point clouds. Voxel side lengths are set, and the number and coordinates of points in each voxel are obtained. The centroid of each voxel is also obtained using the following formula: in, , and They represent the first The centroid of an individual lies on the X, Y, and Z axes. , and They represent the first The first individual element The coordinates of a point on the X, Y, and Z axes. The variable representing the retrieval of a point in a voxel. , , Indicates the first The number of points in a single individual. The retrieval variable represents the voxel. , , Indicates the total number of voxels; The centroid of each voxel represents all points in that voxel. At the same time, the centroid of the voxel is labeled with a label that includes the labels of all points in that voxel. The point cloud formed by the centroids of all voxels is marked as the sub-point cloud of the camera port.
[0011] Furthermore, sub-point clouds of all camera ports are acquired, and the relative positions between the camera ports are mapped in the 3D model. The sub-point clouds of each camera port are then aligned in the 3D model according to the relative positional relationships between the camera ports. The specific alignment method is as follows: Each camera port has an independent default orientation, and a sub-point cloud is built based on this. The mapping points representing the camera ports are determined in the 3D model according to the relative positional relationship between the camera ports. At the same time, the sub-point clouds are connected according to the relative orientation between the camera ports in the real environment. Point cloud reconstruction is performed again in the 3D model to obtain the total point cloud of the area captured by all camera ports. Each point in the total point cloud contains the label of all points in that voxel.
[0012] Furthermore, the target within the camera's field of view is acquired through optical images, and the target is projected into the overall point cloud using a depth camera. The target label is selected based on the target's coordinates in the overall point cloud, and the selection logic is as follows: Obtain all labels from the total point cloud and assign them numbers. Each label corresponds to a point cloud set containing that label. Generate a position score based on the positional relationship between the target coordinates and the point cloud set corresponding to each label, using the following formula: in, Indicates the first The location rating of each tag, , and These represent the target's coordinates on the X, Y, and Z axes, respectively. , and They represent the first In the point cloud set corresponding to the label, the first The coordinates of a point on the X, Y, and Z axes. The retrieval variable represents a point in a point cloud set. , , Indicates the first The total number of points in the point cloud set corresponding to each label. This represents the tag retrieval variable. , This indicates the total number of tags.
[0013] Furthermore, the label with the lowest location score is selected as the target label. The camera port number and corner contained in the target label are sent to the corresponding camera port and controlled to reach the corner in the label. At the same time, the point cloud set corresponding to the target label is marked as the tracking set, and other camera ports are marked as monitoring ports.
[0014] Furthermore, a tag combination is set, in which each monitoring port selects a tag belonging to it, and all tag combinations are obtained; Obtain the point cloud corresponding to each label combination in the total point cloud, and obtain the total number of points in the point cloud, where the total number is the total number of points in the point cloud. The logic for identifying outliers in the overall point cloud is as follows: If a point in the corresponding point cloud carries two or more labels from the label combination, then the point is marked as the first anomalous point, the first anomalous point is numbered, and the number of labels in the label combination carried by the point is recorded. If a point in the corresponding point cloud appears in both the point cloud corresponding to the tracking set and the point cloud corresponding to the label combination, it is marked as the second anomaly. Get the point cloud that is not in the point cloud corresponding to the label combination in the total point cloud, get the points that carry two or more labels at the same time, mark them as third anomalous points, number the third anomalous points, and get the number of labels carried by each third anomalous point.
[0015] Furthermore, the number of all outliers is obtained, and a combined score for each label combination is constructed based on the following formula: in, Indicates the combined score. This indicates the total number of points in the point cloud corresponding to the label combination. Indicates the first The number of tags in the tag combination carried by the first anomaly. The retrieval variable representing the first outlier. , , This represents the total number of the first outlier. Indicates the number of second outliers. Indicates the first The number of tags carried by each third anomaly The retrieval variable representing the third outlier. , , This indicates the total number of third outliers. , and They represent the weights, , , , ; Select the tag combination with the highest combined score, and send the camera port number and corresponding rotation angle contained in the tag combination to the corresponding camera port, so that each camera port reaches the rotation angle in the selected tag combination.
[0016] Furthermore, the target position is reacquired at equal time intervals, and the target label and label combination are updated so that the camera port corresponding to the target label and label combination reaches the corresponding rotation angle.
[0017] Compared with the prior art, the beneficial effects of the present invention are: This invention constructs a total point cloud based on depth images from multiple camera ports, representing the three-dimensional information of the monitored area. During the construction process, the labels of the camera ports and corners that generate each point are recorded. The monitoring area of each camera port is mapped to the pose of the camera port. The target is acquired through the optical images of the camera ports and projected into the total point cloud. A position score is established based on the positional relationship between the point cloud set corresponding to each label and the target. The camera ports and corners most conducive to monitoring the target are selected, and label combinations for each camera port at different corners are established. The distribution of each label combination in the remaining point cloud is judged, and the label combination that can monitor more areas and minimize blind spots is selected. The angles of the remaining camera ports are adjusted according to the selected label combinations. This enables the simultaneous mobilization of other camera ports to reconstruct the monitoring network while a target is being tracked and inspected at a certain camera port. This ensures accurate monitoring of the target while maintaining the maximum monitoring range and reducing the monitoring blind spots. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the overall method flow of the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.
[0020] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0021] Please see Figure 1 The present invention provides a technical solution: A camera based on the ROSO model for heterogeneous information system integration, the camera comprising several camera ports, the camera ports communicating with a central control platform via a service interface defined by the ROSO model, the camera ports being electrically connected to camera units, and the central control platform controlling the camera units via the service interface, the control method including: Step 1: Acquire the depth image and optical image of each camera port respectively, record the depth image of each camera port at each rotation angle, and project the depth image into the 3D model to construct a point cloud set. At the same time, label each point with the camera port number and the rotation angle. Step 1 includes the following: Acquire depth and optical images for each camera port, record the relative position and orientation of each camera port in the real coordinate system, and number all camera ports.
[0022] The camera port generates a depth image at each unit rotation angle. The pixel information of the depth image is extracted by the central control platform and input into the 3D model. The coordinate system of the 3D model is the same as the real coordinate system, and the X, Y and Z axes adopt the default direction of the 3D model. The pixel information in each depth image is converted into a point cloud set in the 3D model, and each point in the point cloud set is labeled. The label indicates the rotation angle of the camera port when the depth image of that point was generated and the number of that camera port. A coordinate point is selected in the 3D model to represent the position of the camera port in the 3D model. Using this point as the anchor point, the point cloud sets generated by the depth images of all cameras at each angle are aligned according to the camera port's intrinsic parameters and rotation angle.
[0023] In a preferred embodiment, the camera port generates a depth image at each corner. The anchor point of this series of depth images is the position of the camera port. The specific value of each corner is also obtained according to the default orientation marked in the camera port's intrinsic parameters. Therefore, when aligning the point clouds generated by a series of depth images in the 3D model, the point clouds generated at different corners can be matched based on the position of the camera port and the default orientation in the intrinsic parameters, so as to reproduce the stereo data of the area monitored by the camera port in the 3D model.
[0024] This step constructs a raw 3D perception base with complete spatiotemporal traceability capabilities, laying the data foundation for the collaborative intelligence of the entire heterogeneous camera system. By acquiring depth images by rotating each camera port at different angles, not only is multi-view image information obtained, but more importantly, by recording the relative position and orientation of each camera port, and converting each depth image pixel into a 3D point cloud with camera port number and corner label, the system achieves dimensionality upgrading and structuring from 2D perception to 3D spatial structure. This process, under a unified real coordinate system, uses each camera port as an anchor point to precisely align observation data from different angles, thus forming a series of "traceable" point cloud sets. Each point clearly records "who saw where, under what conditions," which is equivalent to establishing a reverse index library in 3D space. This design enables the system to quickly trace back and schedule optimal observation resources based on spatial location, shifting from passive environmental recording to active perception scheduling. This provides an indispensable initial environmental model with complete metadata for multi-source data fusion in step 2 and intelligent target perspective tracing in step 3.
[0025] Step 2: Perform point cloud reconstruction on the point cloud set. Point cloud reconstruction means merging the point clouds at overlapping positions into a single point, and integrating the labels of all points at the overlapping positions into a sub-point cloud. Merge the sub-point clouds of all camera ports into a total point cloud in the 3D model, and then perform point cloud reconstruction on the total point cloud. Step 2 includes the following: Step 201: Reconstruct the point cloud from all the point cloud sets generated by the camera port. The point cloud reconstruction logic is as follows: Voxels are divided into all regions containing point clouds. Voxel side lengths are set, and the number and coordinates of points in each voxel are obtained. The centroid of each voxel is also obtained using the following formula: in, , and They represent the first The centroid of an individual lies on the X, Y, and Z axes. , and They represent the first The first individual element The coordinates of a point on the X, Y, and Z axes. The variable representing the retrieval of a point in a voxel. , , Indicates the first The number of points in a single individual. The retrieval variable represents the voxel. , , Indicates the total number of voxels; This formula calculates the centroid coordinates of each voxel, which is a core step in the regularization process of point cloud data. The dependent variable... , and Representing the The coordinates of the centroid of a voxel in three-dimensional space are, in physical terms, the geometric center of all points within that voxel, representing the overall position of that spatial unit. The formula is calculated by taking the arithmetic mean of the coordinates of all points within the voxel: , and As independent variables, they represent the first voxel within the voxel. The three-dimensional coordinates of each point. The average calculation means that the contribution of each point to the centroid coordinates is linear and equally weighted; a change in any coordinate point will cause a change in the centroid coordinates in the same direction, but is affected by the number of points. The averaging effect means that single-point fluctuations have a relatively small impact on the centroid position. The calculation of the centroid is essentially a compression and normalization of spatial data: it aggregates a discrete set of points into a representative coordinate system while preserving complete observation source information by inheriting the labels of all points. This elevates subsequent processing from the "point level" to the "spatial unit level," reducing the amount of data while maintaining spatial structure and information source, laying the foundation for efficient alignment and fusion of multi-view point clouds.
[0026] The centroid of each voxel represents all points in that voxel. At the same time, the centroid of the voxel is labeled with a label that includes the labels of all points in that voxel. The point cloud formed by the centroids of all voxels is marked as the sub-point cloud of the camera port.
[0027] In a preferred embodiment, if a voxel contains x points, each of which is acquired at different angles at the camera port, then each of these x points has a label related to the angle. When these x points are represented by the voxel's centroid, the centroid simultaneously has x labels, which are derived from these x points.
[0028] This step reconstructs the point cloud set of a single camera port through voxelized centroid calculation, achieving regularization, normalization, and intelligent compression of massive point cloud data while preserving information correlation to the maximum extent. This is not a simple downsampling, but a "spatial unitization" process: dividing the continuous space into regular voxels and using a centroid carrying the label information of all points within that unit to represent the entire unit. This has three practical implications: first, it significantly reduces the amount of data and improves system processing efficiency; second, by fusing closely spaced observation points within a small space, it naturally suppresses noise and improves the representativeness and stability of the data; and third, it transforms the originally discrete and uneven raw point cloud into a set composed of regular "information units." Each unit (centroid) becomes an independent data entity carrying multiple label attributes, which creates conditions for subsequent steps of unit-level comparison, fusion, and relationship analysis of cross-camera port data. This is a crucial preprocessing step that transforms raw data into computable data.
[0029] Step 202: Obtain the sub-point clouds of all camera ports, map the relative positions between the camera ports in the 3D model, and align the sub-point clouds of each camera port in the 3D model according to the relative positional relationships between the camera ports. The specific alignment method is as follows: Each camera port has an independent default orientation, and a sub-point cloud is built based on this. The mapping points representing the camera ports are determined in the 3D model according to the relative positional relationship between the camera ports. At the same time, the sub-point clouds are connected according to the relative orientation between the camera ports in the real environment. Point cloud reconstruction is performed again in the 3D model to obtain the total point cloud of the area captured by all camera ports. Each point in the total point cloud contains the label of all points in that voxel.
[0030] As a preferred embodiment, in a real environment, the relative positions of several camera ports can be obtained by measurement, including height differences. In the 3D model, points representing these camera ports are selected based on their relative positions. Based on the default orientation of these camera ports, corresponding default orientations can be established in the 3D model. The construction of each sub-point cloud is based on the default orientation of its respective camera port. In the 3D model, the sub-point clouds of different camera ports can be aligned again based on the default orientation.
[0031] In a preferred embodiment, during the point cloud reconstruction process, if a voxel contains y points, these y points are essentially the centroid of each sub-point cloud at that voxel position. Each point has multiple labels based on its origin (the set of labels generated by each corner in step 1). Assuming that each of these y points has x labels, then in the total point cloud, the centroid of that voxel position has x multiplied by y labels.
[0032] This step represents a qualitative leap from multiple independent local perceptions to a globally consistent perception network. Based on the real spatial relative relationships of the camera ports, coordinate system unification and information stitching are performed on each sub-point cloud. By aligning the mapping points of each camera port and their sub-point clouds in the 3D model, all data is ensured to be integrated into a seamless, globally unified 3D scene representation. More importantly, during this process, labels from different ports and perspectives are superimposed and merged on overlapping spatial units, so that each point in the final generated total point cloud carries a rich set of labels. This effectively constructs a three-dimensional "observation relationship network," clearly recording the history of which cameras and from which angles each location in the environment has been covered. This network is the foundation for the system to understand its overall perception capabilities and blind spots, directly supporting the rapid location of the optimal observation perspective based on target coordinates in step 3, and also providing a panoramic data base map for analyzing the coverage quality, conflicts, and vulnerabilities of the multi-camera collaborative layout in steps 4 and 5.
[0033] Step 3: Locate the target using optical images and project the target position into the overall point cloud. Determine the target label based on the distance relationship between the target position and the point cloud set corresponding to each label in the overall point cloud. Move the camera port corresponding to the target label according to the corner marked by the target label and mark the other camera ports as monitoring ports. Step 3 includes the following: The target within the camera's field of view is acquired using optical images. The target is then projected into the overall point cloud using a depth camera. The target label is selected based on its coordinates in the overall point cloud, following the following selection logic: Obtain all labels from the total point cloud and assign them numbers. Each label corresponds to a point cloud set containing that label. Generate a position score based on the positional relationship between the target coordinates and the point cloud set corresponding to each label, using the following formula: in, Indicates the first The location rating of each tag, , and These represent the target's coordinates on the X, Y, and Z axes, respectively. , and They represent the first In the point cloud set corresponding to the label, the first The coordinates of a point on the X, Y, and Z axes. The retrieval variable represents a point in a point cloud set. , , Indicates the first The total number of points in the point cloud set corresponding to each label. This represents the tag retrieval variable. , This indicates the total number of tags.
[0034] This formula calculates the deviation of the target location from the center of the historical observation field of view of each label, in order to quantitatively assess the applicability of different camera angles. Dependent variable Representing the The location score of each label is physically represented by the squared Euclidean distance between the target point and the average 3D coordinates of the corresponding label point cloud set. A smaller score indicates that the camera corresponding to that label is more aligned with the current target in a certain observation direction. The formula is derived by comparing the target's real-time coordinates. Average coordinates of the label point cloud To achieve the evaluation. The magnitude of the change is directly related to the coordinate difference: the closer the target position is to the centroid of a certain label's field of view, The smaller the value, the larger it is; conversely, the larger the value, the better. The system selects the smallest value. By using this value, the most suitable historical observation perspective can be quickly located, enabling precise scheduling.
[0035] As a preferred embodiment, identifying targets through changes within 2D images is a common technical feature in the field, such as the TP-Link Tapo C260 / C560WS and the Axis Q6411-LE products. When a person enters the field of view, the product can select the person and track them with focus. These products can achieve the acquisition of targets within the field of view of the camera port through optical images as described in this invention. The specific implementation path will not be elaborated here.
[0036] In a preferred embodiment, once a target is detected, the target is bounded in the optical image, and the depth image is projected onto the 3D model based on the depth information of the bounded position, thus obtaining the position of the target in the 3D model.
[0037] In a preferred embodiment, each point in the total point cloud may have multiple labels. According to the formation path of the total point cloud, it can be found that the set of point clouds with the same type of labels is a continuous region because the label is generated at a specific corner of a certain camera port. By searching each continuous region according to the label, when the target appears in the center region of a certain corner of a certain camera port, then the camera port must be the optimal monitoring angle. By having this camera port perform the detection, the optimal tracking and detection effect can be obtained.
[0038] Select the label with the lowest location score as the target label, send the camera port number and corner contained in the target label to the corresponding camera port and control it to reach the corner in the label. At the same time, the point cloud set corresponding to the target label is marked as the tracking set, and other camera ports are marked as monitoring ports.
[0039] This step represents a leap from "target discovery" to "precise scheduling," utilizing a globally consistent 3D model to perform reverse optimal localization and mobilization of sensing resources. After determining the target's 3D coordinates in the overall point cloud using optical images and depth information, the system doesn't allow cameras to search blindly. Instead, it transforms the problem into a data-driven historical optimal viewpoint query: by calculating the distance score between the target coordinates and the average position of the point cloud set corresponding to each label, the system quantitatively evaluates the quality of each historical observation posture aligned with the current target position. Selecting the label with the lowest score as the target label essentially automatically retrieves and reproduces the best historical viewpoint (camera port and angle) for that position. This achieves "memory tracking" of moving targets, with fast response and high accuracy. Simultaneously, this step clarifies the division of labor within the system: one camera acts as the "tracker," locking onto the target, while the others become "monitors," thus initiating a dynamic multi-role collaborative working mode and setting the task prerequisite for the subsequent optimized layout of the monitoring network.
[0040] Step 4: Select a tag belonging to each monitoring port to form a tag combination, obtain all tag combinations, and determine the abnormal points based on the matching relationship between the point cloud set corresponding to the tag combination in the total point cloud and the point cloud set corresponding to the target tag in the total point cloud. Step 4 includes the following: Define tag combinations, where each monitoring port selects a tag belonging to it within the tag combination, and obtain all tag combinations; Obtain the point cloud corresponding to each label combination in the total point cloud, and obtain the total number of points in the point cloud, where the total number is the total number of points in the point cloud. The logic for identifying outliers in the overall point cloud is as follows: If a point in the corresponding point cloud carries two or more labels from the label combination, then the point is marked as the first anomalous point, the first anomalous point is numbered, and the number of labels in the label combination carried by the point is recorded. If a point in the corresponding point cloud appears in both the point cloud corresponding to the tracking set and the point cloud corresponding to the label combination, it is marked as the second anomaly. Get the point cloud that is not in the point cloud corresponding to the label combination in the total point cloud, get the points that carry two or more labels at the same time, mark them as third anomalous points, number the third anomalous points, and get the number of labels carried by each third anomalous point.
[0041] This step involves a systematic virtual simulation and defect pre-detection of all possible future surveillance network layouts for all "monitors." By enumerating all possible tag combinations (each monitoring port selects one observation angle) and defining three types of anomalies for each combination scheme in the overall point cloud model, the complex collaborative performance evaluation is transformed into a calculable and quantifiable analysis. The first anomaly (label overlap within the combination) detects whether the scheme has internal field-of-view redundancy, i.e., multiple monitoring cameras repeatedly cover the same area, resulting in wasted resources. The second anomaly (overlap with the tracking set) detects whether the scheme has task conflicts, i.e., whether the monitoring field of view unexpectedly interferes with the core tracking task. The third anomaly (multiple tag points outside the combination) is used to accurately identify potential monitoring blind spots, i.e., areas that historical data indicates should have good coverage but are missed under this scheme, while also preventing each camera port from ignoring areas where different camera ports connect in pursuit of maximizing the field of view, ensuring the continuity of the monitoring area. This step generates a detailed "health check report" for each possible collaborative layout scheme, providing comprehensive and structured input on coverage, efficiency, and security risks for the final optimization decision in step 5.
[0042] Step 5: Generate a combination score based on the point cloud set and anomalies corresponding to each tag combination in the total point cloud. Select the tag combination with the highest combination score and adjust the angle of each monitoring port according to the number and angle in the tag combination.
[0043] Step 5 includes the following: The number of all outliers is obtained, and a combined score for each label combination is constructed based on the following formula: in, Indicates the combined score. This indicates the total number of points in the point cloud corresponding to the label combination. Indicates the first The number of tags in the tag combination carried by the first anomaly. The retrieval variable representing the first outlier. , , This represents the total number of the first outlier. Indicates the number of second outliers. Indicates the first The number of tags carried by each third anomaly The retrieval variable representing the third outlier. , , This indicates the total number of third outliers. , and They represent the weights, , , , ; This formula is used to quantitatively evaluate the overall effectiveness of each tag combination (i.e., the selection of viewing angles for a group of surveillance cameras). Dependent variable The score represents the combination; the higher the value, the better the monitoring layout is in terms of coverage, resource efficiency, and task coordination. This formula balances multiple key factors through weighted addition and subtraction: The total number of points covered by the combination directly contributes positive points, reflecting the breadth of the monitoring range; (The sum of the number of combined labels carried by the first anomaly) measures internal visual redundancy, and its weighted subtraction (weight) Punish resource waste; (Number of second anomalies) represents the conflict with the core tracking area, through... Weighted subtraction ensures that the tracking task is not disturbed; (Items related to the number of tags for the third anomaly) Evaluate the monitoring blind spots in different camera port connection areas, with weights. The most severe penalties will be imposed on security vulnerabilities. Increase The increase is linear; while the increase in redundancy, conflict, and blind zone related terms decreases linearly through their corresponding weights. Weight setting ( This reflects the priorities: avoiding blind spots is paramount, followed by preventing interference with tracking, and finally reducing redundancy. Therefore, the system maximizes... It automatically selects the optimal monitoring layout to achieve the best balance between coverage, efficiency, and security.
[0044] Select the tag combination with the highest combined score, and send the camera port number and corresponding rotation angle contained in the tag combination to the corresponding camera port, so that each camera port reaches the rotation angle in the selected tag combination.
[0045] This step implements an automated optimization decision-making process under multi-objective constraints. It uses a carefully designed combined scoring formula to comprehensively and quantitatively evaluate and select the best among all label combination schemes. This formula has a clear strategy orientation: maximizing monitoring coverage while mitigating three types of problems through weighted penalties. The weight settings reflect the decision priority: the most severe penalties are imposed on monitoring blind spots (the third anomaly), as they represent security vulnerabilities; strict protection is ensured for tracking tasks without interference (the second anomaly); and then efficiency optimization and reducing unnecessary field-of-view overlap are prioritized (the first anomaly). Selecting the highest-scoring label combination and driving the camera to rotate signifies that the system automatically issues a collaborative instruction that, at the current moment, achieves the optimal balance between coverage and resource efficiency while ensuring core tracking and blind-spot-free monitoring. This essentially replaces a human scheduler, enabling real-time, intelligent optimization of complex monitoring network layouts.
[0046] In a preferred embodiment, the target position is reacquired at equal time intervals, and the target label and label combination are updated so that the camera port corresponding to the target label and label combination reaches the corresponding rotation angle. This injects the entire system with dynamic adaptability and continuous optimization capabilities, transforming it from a static optimization system into an intelligent organism with negative feedback closed-loop control. Reacquiring the target position periodically and repeating steps 3 to 5 means that the system establishes two tightly coupled closed loops: a tracking closed loop, enabling the "tracker" to continuously adapt to the target's movement and maintain stable locking; and a monitoring network optimization closed loop, allowing the layout of all "monitors" to dynamically adjust according to changes in the target position and environment. This mechanism ensures that the system's performance does not degrade due to changes in initial conditions, but rather that it can continuously self-evaluate and self-adjust throughout the task, thereby maintaining optimal or near-optimal global monitoring performance in a dynamic environment for a long time, greatly improving the system's robustness and practicality.
[0047] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0048] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.
[0049] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0050] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A camera for heterogeneous information system integration based on the ROSO model, the camera comprising several camera ports, the camera ports being communicatively connected to a central control platform via a service interface defined by the ROSO model, the camera ports being electrically connected to camera units, and the central control platform controlling the camera units via the service interface, characterized in that, Control methods include: Step 1: Record the depth image of each camera port at each rotation angle, project the depth image into the 3D model to build a point cloud set, and add a label to each point. The label content includes the camera port number and the corresponding rotation angle. Step 2: Perform point cloud reconstruction on the point cloud set. Point cloud reconstruction means merging the point clouds at overlapping positions into a single point, and integrating the labels of all points at the overlapping positions into a sub-point cloud. Merge the sub-point clouds of all camera ports into a total point cloud in the 3D model, and then perform point cloud reconstruction on the total point cloud. Step 3: Locate the target through optical images and project the target position into the reconstructed total point cloud. Determine the target label based on the distance relationship between the target position and the point cloud set corresponding to each label in the total point cloud. Move the camera port corresponding to the target label according to the corner marked by the target label and mark the other camera ports as monitoring ports. Step 4: Select a tag belonging to each monitoring port to form a tag combination, obtain all tag combinations, and determine the abnormal points based on the matching relationship between the point cloud set corresponding to the tag combination in the total point cloud and the point cloud set corresponding to the target tag in the total point cloud. Step 5: Analyze the point cloud and anomalies based on the labels of each monitoring port to generate a combined score for each label combination. Select the label combination with the highest combined score and adjust the angle of each monitoring port according to the number and angle in the label combination.
2. The camera for heterogeneous information system integration based on the ROSO model according to claim 1, characterized in that: Acquire depth and optical images for each camera port, record the relative position and orientation of each camera port in the real coordinate system, and number all camera ports.
3. The camera for heterogeneous information system integration based on the ROSO model according to claim 2, characterized in that: The camera port generates a depth image at each unit rotation angle. The pixel information of the depth image is extracted by the central control platform and input into the 3D model. The coordinate system of the 3D model is the same as the real coordinate system, and the X, Y and Z axes adopt the default direction of the 3D model. The pixel information in each depth image is converted into a point cloud set in the 3D model, and each point in the point cloud set is labeled. The label indicates the rotation angle of the camera port when the depth image of that point was generated and the number of that camera port. A coordinate point is selected in the 3D model to represent the position of the camera port in the 3D model. Using this point as the anchor point, the point cloud sets generated by the depth images of all cameras at each angle are aligned according to the camera port's intrinsic parameters and rotation angle.
4. The camera for heterogeneous information system integration based on the ROSO model according to claim 3, characterized in that: The point cloud set generated from the camera port is reconstructed using the following logic: Voxels are divided into all regions containing point clouds. Voxel side lengths are set, and the number and coordinates of points in each voxel are obtained. The centroid of each voxel is also obtained using the following formula: in, , and They represent the first The centroid of an individual lies on the X, Y, and Z axes. , and They represent the first The first individual element The coordinates of a point on the X, Y, and Z axes. The variable representing the retrieval of a point in a voxel. , , Indicates the first The number of points in a single individual. The retrieval variable represents the voxel. , , Indicates the total number of voxels; The centroid of each voxel represents all points in that voxel. At the same time, the centroid of the voxel is labeled with a label that includes the labels of all points in that voxel. The point cloud formed by the centroids of all voxels is marked as the sub-point cloud of the camera port.
5. The camera for heterogeneous information system integration based on the ROSO model according to claim 4, characterized in that: Obtain the sub-point clouds of all camera ports, map the relative positions between the camera ports in the 3D model, and align the sub-point clouds of each camera port in the 3D model according to the relative positional relationships between the camera ports. The specific alignment method is as follows: Each camera port has an independent default orientation, and a sub-point cloud is built based on this. The mapping points representing the camera ports are determined in the 3D model according to the relative positional relationship between the camera ports. At the same time, the sub-point clouds are connected according to the relative orientation between the camera ports in the real environment. Point cloud reconstruction is performed again in the 3D model to obtain the total point cloud of the area captured by all camera ports. Each point in the total point cloud contains the label of all points in that voxel.
6. The camera for heterogeneous information system integration based on the ROSO model according to claim 5, characterized in that: The target within the camera's field of view is acquired using optical images. The target is then projected into the overall point cloud using a depth camera. The target label is selected based on its coordinates in the overall point cloud, following the following selection logic: Obtain all labels from the total point cloud and assign them numbers. Each label corresponds to a point cloud set containing that label. Generate a position score based on the positional relationship between the target coordinates and the point cloud set corresponding to each label, using the following formula: in, Indicates the first The location rating of each tag, , and These represent the target's coordinates on the X, Y, and Z axes, respectively. , and They represent the first In the point cloud set corresponding to the label, the first The coordinates of a point on the X, Y, and Z axes. The retrieval variable represents a point in a point cloud set. , , Indicates the first The total number of points in the point cloud set corresponding to each label. Indicates the tag retrieval variable. , This indicates the total number of tags.
7. The camera for heterogeneous information system integration based on the ROSO model according to claim 6, characterized in that: The label with the lowest position score is selected as the target label. The camera port number and rotation angle contained in the target label are sent to the corresponding camera port and controlled to reach the rotation angle in the label. At the same time, the point cloud set corresponding to the target label is marked as the tracking set, and other camera ports are marked as monitoring ports.
8. The camera for heterogeneous information system integration based on the ROSO model according to claim 7, characterized in that: Define tag combinations, where each monitoring port selects a tag belonging to it within the tag combination, and obtain all tag combinations; Obtain the point cloud corresponding to each label combination in the total point cloud, and obtain the total number of points in the point cloud, where the total number is the total number of points in the point cloud. The logic for identifying outliers in the overall point cloud is as follows: If a point in the corresponding point cloud carries two or more labels from the label combination, then the point is marked as the first anomalous point, the first anomalous point is numbered, and the number of labels in the label combination carried by the point is recorded. If a point in the corresponding point cloud appears in both the point cloud corresponding to the tracking set and the point cloud corresponding to the label combination, it is marked as the second anomaly. Get the point cloud that is not in the point cloud corresponding to the label combination in the total point cloud, get the points that carry two or more labels at the same time, mark them as third anomalous points, number the third anomalous points, and get the number of labels carried by each third anomalous point.
9. The camera for heterogeneous information system integration based on the ROSO model according to claim 8, characterized in that: The number of all outliers is obtained, and a combined score for each label combination is constructed based on the following formula: in, Indicates the combined score. This indicates the total number of points in the point cloud corresponding to the label combination. Indicates the first The number of tags in the tag combination carried by the first anomaly. The retrieval variable representing the first outlier. , , This represents the total number of the first outlier. Indicates the number of second outliers. Indicates the first The number of tags carried by each third anomaly The retrieval variable representing the third outlier. , , This indicates the total number of third outliers. , and They represent the weights, , , , ; Select the tag combination with the highest combined score, and send the camera port number and corresponding rotation angle contained in the tag combination to the corresponding camera port, so that each camera port reaches the rotation angle in the selected tag combination.
10. The camera for heterogeneous information system integration based on the ROSO model according to claim 9, characterized in that: The target position is reacquired at equal time intervals, and the target label and label combination are updated so that the camera port corresponding to the target label and label combination reaches the corresponding rotation angle.