Method for monitoring multiple cameras based on one indoor unit
Through the methods of timestamp alignment, spatial mapping and dynamic code rate allocation, intelligent collaborative monitoring of multiple cameras is realized, solving the operational complexity and resource waste of traditional building intercom systems, and improving video quality and user experience.
Patent Information
- Application Number
- CN202510469589.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-01
AI Technical Summary
Traditional building intercom systems cannot effectively monitor multiple cameras, resulting in complex operations, high costs, wasted bandwidth and storage resources, reduced video quality and poor user experience, and lack of intelligent linkage and dynamic optimization.
Timestamp alignment and spatial mapping algorithm are used to synthesize a single logical video stream, dynamically adjust the bit rate, use YOLO or Transformer models for object detection and cross-camera tracking, and combine FPGA hardware acceleration processing to realize intelligent collaborative monitoring of multiple cameras.
Reduce hardware costs, reduce bandwidth and storage redundancy, improve resource utilization, improve video resolution and latency performance, enhance the value of security linkage application, and enhance user experience.
Smart Images

Figure CN120238629A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of indoor intercom devices for buildings, and particularly relates to a method for monitoring multiple cameras based on one indoor unit. Background Art
[0002] With the development of society and the improvement of safety awareness, building intercom systems are increasingly widely used in residential and commercial buildings. However, traditional building intercom systems generally have certain limitations;
[0003] 1. Traditional building intercom systems usually adopt the "one device - one camera" mode, that is, each indoor unit can only be bound to one camera. Since each indoor unit can only control one camera, when monitoring multiple areas, users need to frequently switch the screen or configure multiple indoor units, which not only increases the complexity of operation but also raises the cost of the system;
[0004] 2. In the case of independent transmission of multiple video streams, each video stream needs to occupy a certain amount of bandwidth resources, which will lead to waste of bandwidth. At the same time, since the video data generated by each camera needs to be stored separately, there will also be redundancy in storage resources;
[0005] 3. Traditional building intercom systems cannot achieve intelligent linkage analysis between multiple cameras;
[0006] 4. When displaying multiple - screen videos, due to limitations in system resources and processing capabilities, the resolution often decreases, the video delay increases. At the same time, there is a lack of a dynamic optimization mechanism and it is unable to dynamically adjust the video quality according to network bandwidth and user requirements, further affecting the user experience;
[0007] 5. Some systems display multiple video streams through split - screen, but do not solve core problems such as intelligent synthesis of video streams, dynamic bit - rate allocation, and collaborative analysis across cameras;
[0008] To solve the above problems, a method for monitoring multiple cameras based on one indoor unit is proposed in this application. Summary of the Invention
[0009] The present invention provides a method for monitoring multiple cameras based on one indoor unit, which can effectively solve the problems raised in the above - mentioned background art.
[0010] To achieve the above object, the present invention provides the following technical solution: A method for monitoring multiple cameras based on one indoor unit, including the following steps:
[0011] S1: Collect and access video streams collected by multiple cameras that support existing network protocols such as ONVIF and RTSP;
[0012] S2: Process the collected video stream using timestamp alignment and spatial mapping algorithms to synthesize a single logical video stream S(t);
[0013]
[0014] where S(t) is the synthesized video stream, W i (t) is the dynamic weight of the i-th camera, Δt i is the timestamp compensation amount, V i is the original video frame data collected by the i-th camera at timestamp t, i = 1 is the starting index of the summation for the first camera, and n is the number of effectively working cameras;
[0015] S3: Dynamically adjust the bitrates of each camera according to the network bandwidth and the information entropy of the video stream;
[0016]
[0017] where R i is the bitrate allocated to camera i, H(V i ) is the information entropy (complexity) of the video stream V i , P i is the priority weight, B is the total bandwidth, is the weighted sum of the information entropy and priority of all cameras;
[0018] S4: Use the YOLO or Transformer model to detect objects in the synthesized video stream in real-time and implement object tracking through the cross-camera coordinate mapping algorithm, specifically including:
[0019] Calculate the mapping error of the coordinates (x, y) of the object in camera i to camera j:
[0020]
[0021] where Track(x, y) is the tracking function between cameras for the object, T i→j is the coordinate transformation matrix from camera i to j, (x, y) is the actually detected coordinates of the object in camera j, is to find the camera index j that minimizes the error;
[0022] S5: Provide the user with a screen to view the video streams from multiple cameras on the same screen through video stream decoding, frame segmentation, frame synthesis, and display output.
[0023] Preferably, in the method for monitoring multiple cameras based on one indoor unit according to the present invention, in step S2, timestamp alignment includes adjusting the timestamps of the video streams of each camera to achieve time synchronization of the video streams of multiple cameras, and the spatial mapping algorithm includes mapping corresponding points in space for the video frames captured by multiple cameras to achieve spatial synchronization of the video streams of multiple cameras.
[0024] Preferably, in the method for monitoring multiple cameras based on one indoor unit according to the present invention, in step S3, dynamic bitrate allocation further includes automatically reducing the bitrate of the video stream when the network bandwidth is limited to maintain the stability of network transmission.
[0025] Preferably, in the method for monitoring multiple cameras based on one indoor unit according to the present invention, in step S4, the targets detected in the synthesized video stream can be people or vehicles.
[0026] Preferably, in the method for monitoring multiple cameras based on one indoor unit according to the present invention, in step S5, using FPGA hardware acceleration reduces the operation latency of video stream processing and display output.
[0027] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0028] The method for monitoring multiple cameras based on one indoor unit significantly reduces the hardware cost through the function of supporting multiple cameras by a single indoor unit. Through the dynamic bitrate allocation technology, this method can adjust the bitrate of the video stream according to the network condition and the complexity of the video content, reducing the bandwidth occupancy and storage redundancy caused by the independent transmission of multiple video streams and improving the resource utilization rate. In terms of intelligent collaborative monitoring, the system improves the accuracy of cross-camera target tracking through advanced algorithm optimization, greatly enhancing its application value in the security linkage scenario. When displaying multiple pictures, it can maintain a high resolution and low latency, and at the same time introduces a dynamic optimization mechanism to ensure the smoothness and real-time nature of the monitoring pictures, improving the user's monitoring experience; it is compatible with existing cameras and network protocols and supports cloud collaboration, enabling the system to flexibly adapt to the needs of different scenarios, protecting the user's investment and supporting the upgrade and expansion of future technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0030] Figure 1 is the flowchart of the method for monitoring multiple cameras based on one indoor unit according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0032] An embodiment, such as Figure 1 shown, a method for monitoring multiple cameras based on one indoor unit includes the following steps:
[0033] S1: Collect video streams captured by multiple cameras that support existing network protocols such as ONVIF and RTSP and are connected and accessed;
[0034] S2: Using the timestamp alignment and spatial mapping algorithms, the video streams captured by multiple cameras can be processed and synthesized into a single logical video stream S(t). Timestamp alignment adjusts the timestamps of each camera's video stream to ensure the temporal synchronization of multiple camera video streams; while the spatial mapping algorithm performs spatial corresponding point mapping on the video frames captured by multiple cameras, achieving the spatial synchronization of the video streams. These two steps of processing ensure that the video frames from different cameras can be accurately corresponded and fused, and finally merged into a coherent and synchronized single logical video stream S(t), providing a unified temporal and spatial reference for monitoring and analysis;
[0035]
[0036] wherein, S(t) is the synthesized video stream, W i (t) is the dynamic weight of the i-th camera, Δt i is the timestamp compensation amount, V i is the original video frame data captured by the i-th camera at timestamp t, i = 1 is the starting index of the summation for the first camera, and n is the number of effectively working cameras;
[0037] S3: Dynamically adjust the bitrate of each camera according to the network bandwidth and the information entropy (complexity) of the video stream, which means that the system will monitor the changes in the network status and video content in real time. When the network bandwidth is limited or the complexity of the video stream increases, the system will automatically reduce the bitrate of the video stream to ensure the best balance between the stability of network transmission and video quality. This dynamic bitrate allocation strategy can effectively adapt to the changes in the network status and maintain the continuity and reliability of video stream transmission;
[0038]
[0039] wherein, R i is the bitrate allocated to camera i, H(V i ) is the video stream Vi The information entropy (complexity) of P i is the priority weight, B is the total bandwidth, is the weighted sum of the information entropy and priority of all cameras;
[0040] S4: Use the YOLO or Transformer model to detect targets (such as people or vehicles, etc.) in the synthesized video stream in real time, and implement target tracking through the cross-camera coordinate mapping algorithm, specifically including:
[0041] Calculate the mapping error of the coordinates (x, y) of the target in camera i to camera j:
[0042]
[0043] where Track(x, y) is the tracking function between cameras, T i→j is the coordinate transformation matrix from camera i to j, (x, y) is the coordinates actually detected by the target in camera j, is to find the camera index j that minimizes the error;
[0044] S5: Through video stream decoding, frame segmentation, frame synthesis and display output, provide the user with a screen to view video streams from multiple cameras on the same screen, and reduce the operation latency of video stream processing and display output through FPGA hardware acceleration.
[0045] This invention realizes the function of a single indoor unit supporting multiple cameras. This design significantly reduces the hardware cost. Through the dynamic bitrate allocation technology, this method can adjust the bitrate of the video stream according to the network conditions and the complexity of the video content, reduces the bandwidth occupation and storage redundancy caused by the independent transmission of multiple video streams, improves the resource utilization rate. In terms of intelligent collaborative monitoring, the system improves the accuracy of cross-camera target tracking through advanced algorithm optimization, greatly enhances its application value in the security linkage scenario, can maintain a high resolution and low latency when displaying multiple pictures, and at the same time introduces a dynamic optimization mechanism to ensure the smoothness and real-time nature of the monitoring pictures, improving the user's monitoring experience; is compatible with existing cameras and network protocols, and supports cloud collaboration, enabling the system to flexibly adapt to the needs of different scenarios, protecting the user's investment and supporting the upgrade and expansion of future technologies.
[0046] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for monitoring multiple cameras based on one indoor unit, characterized in that , including the following steps: S1: Collect and access video streams from multiple cameras that support ONVIF and RTSP existing network protocols; S2: Use timestamp alignment and spatial mapping algorithms to process the captured video streams to synthesize them into a single logical video stream S(t); Among them, S(t) is the synthetic video stream, W i (t) is the dynamic weight of the i-th camera, Δt i is the timestamp compensation amount, V i is the raw video frame data collected by the i-th camera at time stamp t, i=1 is the first camera of the summation start index, and n is the number of effectively working cameras; S3: Dynamically adjust the bit rate of each camera according to the network bandwidth and information entropy of the video stream; Among them, R i The bit rate assigned to camera i, H(V i ) is the video stream V i Information entropy (complexity), P i is the priority weight, B is the total bandwidth, is the weighted sum of information entropy and priority of all cameras; S4: Use YOLO or Transformer models to detect targets in synthetic video streams in real time and implement target tracking through cross-camera coordinate mapping algorithms, including: Calculate the mapping error of the target coordinates (x, y) in camera i to camera j: Among them, Track(x, y) is the tracking function of the target between cameras, T i→j is the coordinate transformation matrix from camera i to j, (x, y) is the coordinate of the target actually detected in camera j, To find the camera index j that minimizes the error; S5: Through video stream decoding, screen segmentation, screen synthesis and display output, it provides users with a screen that allows them to view video streams from multiple cameras on the same screen.
2. The method for monitoring multiple cameras based on one indoor unit according to claim 1, characterized in that: In S2, the timestamp alignment includes adjusting the timestamp of each camera video stream to achieve time synchronization of multiple camera video streams, and the spatial mapping algorithm includes mapping corresponding points in space on video frames captured by multiple cameras to achieve spatial synchronization of multiple camera video streams.
3. The method for monitoring multiple cameras based on one indoor unit according to claim 1, characterized in that: In S3, the dynamic bit rate allocation further includes automatically reducing the bit rate of the video stream when the network bandwidth is limited to maintain the stability of network transmission.
4. The method for monitoring multiple cameras based on one indoor unit according to claim 1, characterized in that: In S4, the target detected in the synthesized video stream may be a person or a vehicle.
5. The method for monitoring multiple cameras based on one indoor unit according to claim 1, characterized in that: In S5, FPGA hardware acceleration is used to reduce the operation delay of video stream processing and display output.