A wireless monitoring screen method for a construction site electrical area based on free space optical communication and AI vision linkage
Patent Information
- Application Number
- CN202610720958.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-08-18
AI Technical Summary
[0007]上述两种对准维持方案存在一个共同的本质缺陷:其感知反馈信号依赖于通信链路本身的存在
Smart Images

Figure CN122601074A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of free-space optical communication and intelligent video surveillance, and in particular to a wireless monitoring method for temporary power supply areas on construction sites based on free-space optical communication and AI vision linkage. Background Technology
[0002] In the construction of smart construction sites and power transmission and transformation projects, construction sites often contain a large number of high-voltage power equipment and live power lines under construction, forming so-called "temporary power areas." The strong power frequency electromagnetic fields and broadband electromagnetic noise generated by corona discharge in these areas pose a serious physical layer interference to conventional radio frequency wireless communication. However, in order to achieve real-time safety management, construction quality traceability, and remote expert guidance in these areas, it is precisely necessary to reliably transmit high-definition video monitoring footage from the site back to the monitoring center in real time. This contradiction between the "strong electromagnetic interference environment" and the "real-time video transmission requirement" constitutes the primary technical challenge for monitoring and communication in temporary power areas of construction sites.
[0003] Currently, there are three main types of solutions in the industry for the above scenarios, but each has its own limitations.
[0004] The first type of solution is radio frequency (RF)-based wireless video transmission. For example, common construction site monitoring systems use 5.8GHz Wi-Fi, 4G / 5G cellular networks, or industrial-grade data radios for video streaming. However, in areas near high-voltage power lines, strong electromagnetic interference can drastically degrade the signal-to-noise ratio (SNR) of the RF channel. To combat this interference, systems typically need to reduce data transmission rates, increase the number of data retransmissions, or employ more complex channel coding. These measures introduce uncontrollable latency and video interruptions, making it difficult to meet the requirements of real-time security management for video continuity and low latency. Furthermore, the broadcast nature of RF signals makes them vulnerable to eavesdropping; in construction site scenarios involving critical infrastructure, information security issues are equally critical.
[0005] The second option is wired fiber optic transmission. In areas with particularly severe electromagnetic interference, laying armored fiber optic cables can completely avoid this problem. However, fiber optic laying requires pre-planned routes and trench excavation, making it difficult to adapt to the characteristics of construction site operations characterized by frequent construction progress and dynamic movement of machinery. Fiber optic cables are prone to breakage and damage under the unique physical risks of construction sites, such as being crushed by heavy machinery or falling objects from heights, resulting in high maintenance costs. More importantly, fiber optic cables cannot extend to handheld mobile monitoring terminals and temporarily installed monitoring points, limiting the realization of the core function of "mobile collaborative monitoring" in smart construction sites.
[0006] The third type of solution is free-space optical communication. Free-space optical communication uses laser or infrared light as a carrier to transmit data point-to-point in the atmosphere. Because its operating band is far from the radio frequency range, it has natural resistance to electromagnetic interference and provides communication bandwidth far exceeding that of radio frequency solutions. In existing technologies, free-space optical communication has been applied in scenarios such as fixed backbone network bridging between buildings. However, traditional free-space optical communication equipment is usually installed on fixed building structures, and its line-of-sight alignment is completed and locked manually by mechanical adjustment, remaining stationary once installed. In construction site scenarios, construction machinery and mobile equipment carrying communication terminals inevitably experience continuous and slow attitude drift due to vibration, thermal deformation, and foundation settlement, causing the line-of-sight alignment of free-space optical communication to continuously degrade. To address the challenge of maintaining alignment in this dynamic environment, existing free-space optical communication solutions mainly rely on the following two technical approaches: First, closed-loop tracking based on received optical power feedback, which involves real-time monitoring of changes in the optical power at the receiving end, and driving the gimbal to perform fine-tuning search to restore peak power when the optical power drops below a threshold; Second, spot position feedback based on independent four-quadrant detectors or position-sensitive detectors, which involves driving the gimbal to compensate for pointing deviation by detecting the offset of the focused spot on the detector.
[0007] Both alignment maintenance schemes mentioned above share a common fundamental flaw: their sensing feedback signals depend on the existence of the communication link itself. When the free-space optical communication link is interrupted due to temporary obstruction or excessive relative displacement between the transmitter and receiver, the received optical power drops to zero, the light spot completely disappears from the detector's photosensitive surface, and the sensing feedback signal is lost entirely. At this time, the system cannot determine the current spatial location of the other end, nor can it determine which direction to adjust its own pointing to re-capture the other end; it can only initiate a blind scan search across the entire field of view. The time taken for blind scanning depends on the size of the uncertain area and the scan step size, typically requiring several seconds to tens of seconds. During this period, communication is completely interrupted, and the video transmission image freezes, which is unacceptable for security management scenarios requiring continuous real-time monitoring. In other words, the existing schemes suffer from a vicious cycle of "communication interruption leading to loss of sensing, and loss of sensing hindering communication recovery," the root cause of which lies in the deep physical coupling between the sensing channel and the communication channel, preventing them from operating independently.
[0008] In summary, existing technologies have three main shortcomings when addressing the need for wireless video surveillance in construction site areas with temporary power supply: radio frequency solutions are limited by electromagnetic interference, wired solutions lack deployment flexibility, and traditional free-space optical communication solutions lack the ability to perceive and predict the pose of the other end independently of the communication state. Summary of the Invention
[0009] The technical problem to be solved by this invention is: how to construct a peer orientation sensing mechanism decoupled from the communication state for a free-space optical communication link in the context of strong electromagnetic interference in the power supply area of a construction site, so that when the free-space optical communication link is interrupted due to continuous drift or temporary blockage of the transceiver end, the system can still independently obtain the spatial orientation information of the peer end, thereby replacing blind scanning search with predictive hold, and realizing the rapid recovery of the free-space optical communication link and the continuous and reliable backhaul of real-time video stream.
[0010] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0011] In a first aspect, the present invention provides a wireless monitoring method for temporary power areas at construction sites based on the linkage of free-space optical communication and AI vision. This method is applied to a system comprising at least one front-end terminal and one back-end terminal, with each terminal equipped with a communication-vision shared optical path integrated gimbal module. The method includes the following steps:
[0012] Step 1: AI Vision Active Alignment
[0013] The integrated gimbal module of the first terminal uses its monitoring camera to capture image frames containing the cooperation identifier on the second terminal. Using an AI vision algorithm, it detects multiple constituent light points of the cooperation identifier in real time from the image frames, calculates the six-degree-of-freedom spatial pose offset of the cooperation identifier relative to the optical axis of the first terminal, and generates gimbal servo compensation commands to drive the first terminal gimbal, ensuring that the geometric center of the cooperation identifier remains at the preset optical center position of the image frame. The communication-vision shared optical path integrated gimbal module uses a multi-band beam splitter to transmit visible light to the monitoring camera and reflect near-infrared light to the free-space optical communication transceiver unit, ensuring that the visual optical axis and the communication optical axis are coaxial behind the objective lens, thereby ensuring that visual alignment is equivalent to line-of-sight alignment in free-space optical communication.
[0014] Step 2: Free-space optical communication transmission
[0015] Under the line-of-sight alignment conditions established and maintained in step one, the first terminal and the second terminal communicate bidirectionally in the atmospheric channel via the free-space optical communication transceiver unit using the near-infrared band, carrying monitoring video streams, PTZ control commands, and AI model data.
[0016] Step 3: Predictive Preservation During Communication Interruption
[0017] When the cooperative identifier fails to perform visual detection in step one due to temporary occlusion, the current spatial position of the cooperative identifier is predicted by extrapolation of the motion model based on the pose history sequence of the cooperative identifier calculated from the last N frames before the occlusion. The gimbal is then driven to be preset to the predicted position and held. When the occlusion is removed and the cooperative identifier becomes visible again, visual detection is restored and the free space optical communication link is rebuilt.
[0018] The core innovation of the above method lies in replacing the feedback sensing source upon which free-space optical communication relies for alignment with the received optical power or spot position of the communication link itself with AI visual pose perception that exists independently of the communication state. Visual perception in the visible light band and communication transmission in the near-infrared band achieve spatial unification through a shared optical path design on the physical optical path, but are completely decoupled at the information level—visual perception does not depend on the existence of the communication link. Therefore, at any moment of communication interruption, as long as the cooperative identifier is not completely obscured, the system can continuously acquire the spatial orientation information of the other end.
[0019] As a further description of the above technical solution, the communication-vision shared optical path integrated gimbal module includes: a motorized zoom objective lens, a multi-band beam splitter prism, a visible light monitoring camera module, a free-space optical communication transceiver module, and a dual-axis servo gimbal. The multi-band beam splitter prism has a transmittance of no less than 92% for the 400nm to 700nm wavelength band and a reflectance of no less than 95% for the 780nm to 1600nm wavelength band. This beam splitting parameter setting ensures that the visible light monitoring camera obtains sufficient exposure, while the loss of the near-infrared signal in the free-space optical communication path is controlled to within 5%.
[0020] As a further description of the above technical solution, the cooperation identifier consists of 16 infrared LEDs arranged in a 4×4 square array, with a spacing of 40mm between adjacent LEDs. The infrared LEDs emit a wavelength of 850nm, which falls within the near-infrared response range of the visible light surveillance camera's CMOS, appearing as bright spots in the monitoring image. The infrared LEDs emit pulses at a frequency of 200Hz with a pulse width of 2ms. The 16 infrared LEDs are divided into four groups, each group encoding a unique device identifier for the cooperation identifier using a preset amplitude modulation mode. This design allows the AI vision algorithm to uniquely identify the target end by demodulating the device identifier when multiple terminals have cooperation identifiers in the same field of view, avoiding misalignment.
[0021] As a further description of the above technical solution, the AI vision algorithm in step one employs a two-pronged parallel strategy for detecting the cooperative identifier: one path uses a target detection network to identify candidate regions for light spots in the image frame, while the other path uses prior knowledge of the pulse frequency of the cooperative identifier to perform differential analysis on two consecutive image frames to enhance the pulse light source and suppress the background. The candidate regions output from the two paths are spatially matched and weighted by confidence; only those with a fusion confidence score higher than 0.85 are confirmed as valid light spots. This dual-path fusion detection strategy enables the algorithm to have high resistance to ambient light interference under normal operating conditions, and can still maintain basic detection functions by relying on the target detection network even under abnormal operating conditions of the LED pulse circuit.
[0022] As a further description of the above technical solution, the six-degree-of-freedom spatial pose offset mentioned in step one is calculated using the Perspective-n-Point algorithm. When the number of detected constitutive light points is no less than eight, the translation vector and rotation vector are calculated using the two-dimensional pixel coordinates of the constitutive light points and their corresponding three-dimensional world coordinates on the cooperative identification physical panel. The translation vector includes horizontal offset, vertical offset, and axial distance. Real-time calculation of the axial distance enables the system to dynamically evaluate the current communication distance, providing a basis for adaptive adjustment of the free-space optical communication transmission power.
[0023] As a further description of the above technical solution, in step one, after the cooperative identifier is successfully detected stably for 10 consecutive frames, the system switches to sparse optical flow tracking mode. In sparse optical flow tracking mode, based on the detected light point positions in the previous frame, the corresponding position in the current frame is predicted using the inverse optical flow method, and the predicted position is used as the input for pose calculation. Only when the number of tracked light points drops below 8 or the pose calculation residual after 5 consecutive frames of tracking increases by more than 50% compared to the start of tracking, the system reverts to the full detection pipeline. This tracking mode switching mechanism reduces the average computation latency of AI vision alignment from approximately 33ms to approximately 8ms, significantly improving the effective bandwidth of the servo control closed loop.
[0024] As a further description of the above technical solution, the motion model described in step three is a constant acceleration model, and its expression is x(t) = x0 + v0·t + ½·a·t 2Where x represents the spatial position of the cooperative identifier, t represents the time elapsed since the visual lock loss, x0 represents the spatial position of the cooperative identifier at the instant of visual lock loss, v0 represents the velocity of the cooperative identifier at the instant of visual lock loss, and a represents the acceleration of the cooperative identifier. x0, v0, and a are estimated by weighted least-squares fitting of the last 16 frames of valid pose data in a 32-frame circular history buffer, with the most recent frames assigned higher fitting weights. In the construction site scenario, the motion of the counterpart mainly originates from gimbal mechanical creep, low-frequency vibrations transmitted by construction machinery, and slow drift caused by thermal deformation. The constant acceleration model can control the predicted position deviation within half the width of the free-space optical communication receiving window within 2 seconds after the lock loss.
[0025] As a further description of the above technical solution, the free-space optical communication transceiver unit continuously monitors the received optical power indication and packet error rate at the physical layer. When the received optical power indication drops below a preset threshold or the packet error rate rises above a preset limit value within five consecutive time slots, the second terminal sends a link quality warning control frame to the first terminal via the free-space optical communication downlink. After receiving the link quality warning control frame, the first terminal increases the execution frequency of the visual inspection in step one from 30Hz to 60Hz and switches the monitoring video encoding format from H.265 to MJPEG. This adaptive adjustment mechanism enables the system to proactively respond to the instantaneous quality degradation of the link with a denser visual inspection frequency and a more robust low-bitrate encoding format before the alignment deviation deteriorates severely to the point of complete communication interruption.
[0026] Secondly, this invention also provides a wireless monitoring system for temporary power supply areas at construction sites based on free-space optical communication and AI vision linkage. The system includes at least one front-end portable communication monitoring terminal and one back-end monitoring center receiving terminal. Each terminal is equipped with a communication-vision shared optical path integrated gimbal module. The communication-vision shared optical path integrated gimbal module includes a motorized zoom objective lens, a multi-band beam splitter prism, a visible light monitoring camera module, a free-space optical communication transceiver module, and a dual-axis servo gimbal. The system is configured to execute the steps of the aforementioned wireless monitoring method for temporary power supply areas at construction sites based on free-space optical communication and AI vision linkage.
[0027] Thirdly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the aforementioned wireless monitoring method for temporary power areas at construction sites based on free-space optical communication and AI vision linkage.
[0028] Beneficial effects:
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] The aforementioned wireless monitoring method for temporary power areas on construction sites, based on free-space optical communication and AI vision linkage, fundamentally breaks the vicious cycle of "communication interruption leading to loss of perception, and loss of perception hindering communication recovery." This invention physically unifies the visual and communication optical axes through a shared optical path design for communication and vision, while simultaneously enabling visual perception to be independent of the communication state at the information level through band separation. When the free-space optical communication link is interrupted due to obstruction or drift, the AI visual perception module can still independently acquire the spatial orientation information of the counterpart's cooperative identifier through the visible light band, providing clear spatial guidance for the rapid recovery of the communication link and completely avoiding the predicament of traditional solutions entering a blind scanning search without prior knowledge after communication interruption.
[0031] The aforementioned wireless monitoring method for temporary power areas on construction sites, based on free-space optical communication and AI vision linkage, replaces mechanical scanning search with visual servo prediction and hold, significantly shortening the recovery time after communication interruption. Traditional free-space optical communication schemes require driving a gimbal to perform helical or grating scanning within an uncertain area after the light spot is lost, with recovery times typically ranging from several seconds to tens of seconds. This invention, based on the pose history sequence before occlusion, extrapolates the current position of the other end using a constant acceleration motion model and drives the gimbal to preset hold. In typical transient occlusion scenarios on construction sites, the recovery time of the communication link after the occlusion is removed depends only on the physical speed of the removal of the occluder, and can typically be rebuilt within 0.2 seconds, meeting the requirements of real-time video transmission for link continuity.
[0032] The aforementioned wireless monitoring method for temporary power areas on construction sites, based on free-space optical communication and AI vision linkage, eliminates calibration drift errors between the communication optical axis and the tracking sensor optical axis. In existing tracking schemes based on independent four-quadrant detectors, the communication and tracking optical paths are separate on the focal plane, and their coordinate mapping depends on mechanical assembly accuracy and drifts with temperature and time, requiring periodic offline calibration. This invention employs a multi-band beam splitter prism to achieve a coaxial design of visible and near-infrared light behind the objective lens. The correspondence between the visual and communication optical axes is determined by the accuracy of the color separation coating rather than mechanical assembly accuracy, exhibiting inherent long-term stability and eliminating the need for periodic calibration and maintenance.
[0033] The aforementioned wireless monitoring method for temporary power areas on construction sites, based on free-space optical communication and AI vision linkage, enhances the system's target identification capability in multi-terminal scenarios. The cooperative identifier uses 16 infrared LEDs encoded with a unique device identifier via grouped amplitude modulation. The AI vision algorithm demodulates the device identifier while detecting the spatial position of the light spot. This ensures that when multiple terminals have cooperative identifiers in the same field of view, the system can uniquely lock onto the target terminal without misaligning it with other non-target terminals, providing a reliable visual identification foundation for one-to-many concurrent monitoring under a star-shaped network architecture.
[0034] The aforementioned wireless monitoring method for temporary power areas on construction sites, based on the linkage of free-space optical communication and AI vision, achieves proactive preventative maintenance of link quality through deep integration of vision and communication. The received optical power indication and packet error rate monitoring of the free-space optical communication physical layer can provide early warnings before alignment deviations severely deteriorate to the point of communication interruption, triggering an increase in visual inspection frequency and a switch to video encoding degradation. This allows for more intensive alignment monitoring and a more robust encoding strategy to proactively address instantaneous link degradation, transforming a passive response to link interruption into proactive prevention. Attached Figure Description
[0035] Figure 1 This is an overall flowchart of the wireless monitoring method for temporary power areas on construction sites based on free space optical communication and AI vision linkage in this embodiment of the invention.
[0036] Figure 2 yes Figure 1 The algorithm flowchart for the first step of the AI vision active alignment process.
[0037] Figure 3 yes Figure 1 A flowchart illustrating the predictive hold process during step three of the middle stage when communication is interrupted.
[0038] Figure 4 This is a schematic diagram of the process of vision-communication linkage and adaptive adjustment of link quality in an embodiment of the present invention.
[0039] Figure 5 This is a schematic diagram of the overall architecture of a wireless monitoring system for temporary power areas at construction sites based on free-space optical communication and AI vision linkage, provided in an embodiment of the present invention. Detailed Implementation
[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0041] Example 1: Fixed-point monitoring screen in temporary power areas of power transmission and transformation projects
[0042] This embodiment provides a wireless monitoring method for temporary power areas on construction sites based on free-space optical communication and AI vision linkage. This method is applied to power transmission and transformation engineering construction sites to achieve reliable wireless transmission of monitoring videos in areas near high-voltage live equipment.
[0043] Reference Figure 5The system based on this embodiment includes at least one front-end portable communication monitoring terminal and one back-end monitoring center receiving terminal. The front-end portable communication monitoring terminal is deployed at monitoring points within the power supply area, and the back-end monitoring center receiving terminal is deployed at the on-site monitoring center far from areas with strong electromagnetic interference. Each of the front-end portable communication monitoring terminal and the back-end monitoring center receiving terminal is equipped with a communication-vision shared optical path integrated PTZ module.
[0044] The communication-vision shared optical path integrated gimbal module is the core physical device of this system. Its structure includes an electric zoom objective lens, a multi-band beam splitter prism, a visible light monitoring camera module, a free space optical communication transceiver module, and a dual-axis servo gimbal.
[0045] The motorized zoom objective has a focal length range of 10mm to 80mm and an F-number of 1.8. At the 80mm telephoto end, the field of view is approximately ±2.5°, which is on the same order of magnitude as the divergence angle of the laser beam used in free-space optical communication, ensuring that the counterpart cooperation marker remains clearly visible within the communication distance. At the 10mm wide-angle end, the field of view is approximately ±18°, which is used for rapid acquisition of the counterpart cooperation marker over a wide area during the initial installation phase.
[0046] A multi-band beam splitter is positioned between the motorized zoom objective and the image plane. The coating parameters of the beam splitter are: transmittance of no less than 92% for the visible light band (400nm to 700nm) and reflectance of no less than 95% for the near-infrared band (780nm to 1600nm). The incident light beam from the scene is converged by the motorized zoom objective and then separated by band at the multi-band beam splitter: the visible light band is transmitted to the CMOS sensor surface of the visible light monitoring camera module to generate monitoring video frames; the near-infrared band is reflected into the free-space optical communication transceiver module.
[0047] The visible light surveillance camera module uses a 1 / 1.8-inch CMOS sensor with an effective pixel count of 1920×1080 and a frame rate of 30fps. It is mounted on the back focal plane of the transmission direction of the multi-band beam splitter.
[0048] The free-space optical communication transceiver module is mounted on the reflection direction of a multi-band beam splitter prism. Internally, it includes: a distributed feedback laser as the emission source, with a center wavelength of 1550nm, an adjustable output power of 50mW, and a divergence angle of approximately 1.5mrad after collimation; and an InGaAs avalanche photodiode as the receiving detector, with a photosensitive surface diameter of 200μm and a transimpedance amplifier with a bandwidth of 1GHz at the front end. The emitted laser and received light share a single external optical port via a built-in optical circulator, achieving coaxial multiplexing for transmission and reception. The incident communication laser from the other end is reflected by the multi-band beam splitter prism and focused onto the InGaAs avalanche photodiode; the emitted laser from this end is reflected in the opposite direction by the same reflective surface of the multi-band beam splitter prism and collimated into a narrow beam by an electrically adjustable zoom objective lens before being emitted towards the other end.
[0049] The aforementioned optical arrangement ensures a crucial geometric relationship: when the image of a distant point light source falls precisely at the pixel coordinates on the CMOS sensor of the visible light monitoring camera module, the light emitted from that point light source will inevitably pass through the pupil center of the motorized zoom objective lens; and the laser emitted by the free-space optical communication transceiver module will also inevitably exit along the same reverse optical path, pointing towards the spatial location of the point light source. The visual optical axis and the communication optical axis are strictly coaxial behind the objective lens, and the angle between them is theoretically zero and less than 0.01° within the mechanical assembly tolerance range.
[0050] The dual-axis servo gimbal carries all the aforementioned optical components. Its azimuth travel is ±180°, its pitch travel is ±45°, its angular resolution is 0.001°, its maximum angular velocity is 30° / s, and its positioning accuracy is 0.01°.
[0051] Continue to refer to Figure 5 Each communication-vision shared optical path integrated PTZ module has a cooperation logo mounted on its panel. The physical center of the cooperation logo coincides with the mechanical center of the external optical port of the free-space optical communication transceiver module. The cooperation logo consists of 16 infrared LEDs arranged in a 4×4 square array, with an adjacent LED spacing of 40mm, and the overall array dimensions are 120mm×120mm. The infrared LEDs emit a wavelength of 850nm, which is within the near-infrared response range of the visible light monitoring camera module's CMOS, appearing as bright spots in the monitoring image. Each LED is fitted with a narrowband filter with a center wavelength of 850nm and a full width at half maximum (FWHM) of 10nm. The LEDs emit pulses at a frequency of 200Hz, with a pulse width of 2ms and a duty cycle of 40%. The 16 LEDs are divided into four groups of four, with each group of four LEDs modulated using a preset binary amplitude modulation mode to encode the unique device identifier of the cooperation logo.
[0052] Reference Figure 1 The wireless monitoring method in this embodiment includes the following steps.
[0053] Step 1: AI vision actively aligns.
[0054] The communication-vision shared optical path integrated gimbal module of the front-end portable communication monitoring terminal uses its visible light monitoring camera module to capture image frames containing cooperative identifiers on the back-end monitoring center receiving terminal. Through AI vision algorithms deployed on the edge AI processor built into the integrated gimbal module, multiple constituent light points of the cooperative identifier are detected in real time from the image frame. The six-degree-of-freedom spatial pose offset of the cooperative identifier relative to the optical axis of the front-end portable communication monitoring terminal is calculated, and based on this, gimbal servo compensation commands are generated to drive the dual-axis servo gimbal, ensuring that the geometric center of the cooperative identifier remains at the preset optical center position of the image frame.
[0055] Reference Figure 2 The AI vision active alignment algorithm consists of four interconnected stages.
[0056] The first stage is the detection of cooperative identifier light spots. The input is a grayscale image of the current monitoring video frame with a resolution of 640×480 pixels. The light spot detection adopts a two-way parallel strategy: one path runs the target detection network YOLOv5-nano with approximately 1.9M parameters to detect all candidate light spot regions in the image; the other path utilizes the prior knowledge of the 200Hz pulse frequency of the LED in the cooperative identifier to perform difference analysis on two consecutive frames. In the differenced image, the static background is suppressed while the pulsed light source is enhanced. The difference result is then adaptively thresholded and binarized to output a candidate light spot mask. The candidate light spot regions output from the two paths are then spatially matched using the intersection-over-union (IoU) ratio and fused using confidence weighting. Only when the fused confidence score is higher than 0.85 is the candidate region confirmed as a valid cooperative identifier constituting a light spot.
[0057] The second stage is 3D pose calculation. The pose calculation phase begins when at least 8 of the 16 constituent light points are successfully detected. The Perspective-n-Point algorithm is used, with the input being the 2D pixel coordinates of the detected light points and their corresponding 3D world coordinates on the cooperative marker physical panel. EPnP is used as the initialization method, followed by Levenberg-Marquardt iterative optimization to minimize the reprojection error. The output is a six-DOF pose vector defined in the camera coordinate system. The translation vector's three components represent the horizontal, vertical, and axial offsets of the cooperative marker center relative to the local camera's optical center, respectively; the rotation vector, after Rodrigues transform, yields three Euler angles.
[0058] The third stage involves generating gimbal servo compensation commands. The horizontal and vertical offsets obtained from pose calculation are transformed from the camera coordinate system to the gimbal angular coordinate system. The transformation formulas are: azimuth deviation equals the arctangent of the horizontal offset divided by the axial distance, and pitch deviation equals the arctangent of the vertical offset divided by the axial distance. The servo controller uses a PID control law to generate stepping commands for the dual-axis servo gimbal motors. The PID parameters are tuned as follows: proportional coefficient 1.2, derivative coefficient 0.3, and integral coefficient 0.02. The control cycle is synchronized with the visual processing frame rate, at 33ms.
[0059] The fourth stage is sparse prediction tracking. After 10 consecutive frames of stable detection of the cooperative identifier, the system switches to sparse optical flow tracking mode. In sparse optical flow tracking mode, the complete detection and PnP solution pipeline is no longer run for each frame. Instead, the positions of the detected constituent points in the previous frame are maintained, and the corresponding positions of the constituent points in the current frame are predicted using the Lucas-Kanade inverse optical flow method. The predicted positions are then used as input for the PnP solution. The system reverts to the complete detection pipeline only if either of the following conditions is met: the number of tracked points drops to 8 or less; or the PnP solution residual after 5 consecutive frames of tracking increases by more than 50% compared to the residual at the start of tracking.
[0060] Through the above four stages, AI vision active alignment continuously locks the geometric center of the cooperative identifier at the preset optical center position of the image frame. Since the visual optical axis and the communication optical axis in the communication-vision common optical path integrated gimbal module are strictly coaxial behind the objective lens, the image center is the direction of the communication optical axis, and vision alignment is equivalent to line-of-sight alignment in free space optical communication.
[0061] Step two, free space optical communication transmission.
[0062] Under the line-of-sight alignment conditions established and maintained in step one, the front-end portable communication monitoring terminal and the back-end monitoring center receiving terminal conduct bidirectional data communication in the atmospheric channel through their respective free-space optical communication transceiver modules using the 1550nm near-infrared band.
[0063] The bidirectional free-space optical communication link adopts a half-duplex mode, carrying bidirectional data streams on the same wavelength in a time-division duplex manner. The time slot length is 2ms, with downlink and uplink time slots each occupying 1ms, alternating. The physical layer uses on-off keyed direct detection modulation, with a communication rate of 500Mbps. The link layer employs an ARQ error control protocol based on selective repeat, with a maximum retransmission count of 3.
[0064] The front-end portable communication monitoring terminal encodes the high-definition monitoring video stream captured by the visible light monitoring camera module using H.265 and transmits it back to the back-end monitoring center receiving terminal in real time via a free-space optical communication uplink. The PTZ remote control commands and AI model update data generated by the back-end monitoring center receiving terminal are sent to the front-end portable communication monitoring terminal via a free-space optical communication downlink.
[0065] Step 3: Predictive preservation during communication interruptions.
[0066] When the cooperation logo is temporarily obscured, causing the visual detection in step one to fail, the system enters predictive hold mode. (Refer to...) Figure 3 The specific execution process of this mode is as follows.
[0067] Throughout the normal operation of the system, the PnP solution results of each frame in step one—that is, the six-DOF pose of the cooperative identifier in the camera coordinate system—are continuously recorded in a circular history buffer with a length of 32 frames, which covers a time span of about 1 second.
[0068] When a visual detection failure is triggered, the system extracts the last 16 frames of valid pose data from the circular history buffer and constructs a constant acceleration motion model for extrapolation prediction. The expression for the constant acceleration model is x(t) = x0 + v0·t + ½·a·t 2 , where x represents the spatial position of the cooperative marker, t represents the time elapsed since the moment of visual loss of lock, x0 represents the spatial position of the cooperative marker at the instant of visual loss of lock, v0 represents the velocity of the cooperative marker at the instant of visual loss of lock, and a represents the acceleration of the cooperative marker. x0, v0, and a are obtained by weighted least squares fitting estimation of the 16 frames of effective pose data, with the data of the most recent frames being given higher fitting weights.
[0069] Based on the above model, the system extrapolates the spatial position sequence of the cooperative identifier within 2 seconds after the lock is lost, at 33ms intervals. The servo controller of the dual-axis servo pan-tilt unit uses these predicted positions as feedforward commands and superimposes them onto the output of the PID controller, driving the dual-axis servo pan-tilt unit to continuously track the predicted trajectory. During the prediction hold period, the free-space optical communication link may be interrupted due to momentary occlusion. Newly generated monitoring video data frames are temporarily stored in the front-end local buffer, awaiting burst transmission after the link is restored.
[0070] When the temporary obstruction is removed and the cooperative identifier is re-exposed in the field of view of the visible light monitoring camera module, the first-stage detection pipeline immediately resumes operation, re-capturing the constituent light spot of the cooperative identifier. Since the dual-axis servo pan-tilt unit remains near the predicted position according to the prediction model, the light spot of the cooperative identifier is highly likely still within the field of view of the monitoring camera, enabling instantaneous recovery of visual detection. The time from visual detection recovery to free-space optical communication link reconstruction is mainly determined by the clearing speed of the link layer selective retransmission buffer, which is approximately 0.2 seconds under typical operating conditions in this embodiment.
[0071] In the operation of this embodiment, a real-time linkage mechanism also exists between the AI vision module and the free-space optical communication module. (Refer to...) Figure 4 The specific working method of this mechanism is as follows:
[0072] The free-space optical communication transceiver module continuously monitors two link quality indicators at the physical layer: received optical power indication and packet error rate. When the received optical power indication drops by more than 3dB within five consecutive time slots, or the packet error rate rises to more than 1%, the receiving terminal at the back-end monitoring center sends a link quality warning control frame to the front-end portable communication monitoring terminal via the free-space optical communication downlink.
[0073] After receiving the link quality early warning control frame, the edge AI processor of the front-end portable communication monitoring terminal performs two adaptive operations: First, it temporarily increases the execution frequency of AI visual detection in step one from 30Hz to 60Hz to increase the monitoring density of small alignment deviations; second, it switches the monitoring video encoding format from H.265 to MJPEG and reduces the resolution to 1280×720 to reduce the amount of data per frame to adapt to the current lossy channel conditions.
[0074] Once the received optical power indicator returns to the normal range and remains stable for 10 time slots, the AI visual detection frequency returns to 30Hz, and the video encoding format switches back to H.265.
[0075] Example 2: Multi-terminal star network and mobile collaborative monitoring
[0076] Based on Embodiment 1, this embodiment provides a wireless monitoring system that supports star-shaped networking between multiple front-end terminals and one back-end terminal.
[0077] At long-distance pipeline construction sites, multiple construction points operate simultaneously, with each point deploying a front-end portable communication monitoring terminal. The back-end monitoring center receiving terminal is equipped with multiple communication-vision shared optical path integrated PTZ modules, each independently aligned with one front-end portable communication monitoring terminal. A star topology is formed between the front-end and back-end terminals.
[0078] Multiple communication-vision shared optical path integrated gimbal modules of the backend monitoring center receiving terminal are mounted on a common rotating base, but each has an independent pitch adjustment axis. The control software of the backend monitoring center analyzes the position of the cooperation marker in the video returned by each front-end terminal in real time, and drives the independent dual-axis servo gimbal of each module in parallel to achieve one-to-many concurrent real-time monitoring.
[0079] For the handheld mobile monitoring terminal, the cooperation logo undergoes continuous large-scale displacement due to personnel movement. An additional six-axis MEMS inertial measurement unit is integrated on the mobile terminal to acquire three-axis acceleration and three-axis angular velocity data at a frequency of 200Hz. Under normal conditions where the AI visual detection in step one is continuously effective, the inertial measurement unit data and the visual PnP pose are loosely coupled and fused through an extended Kalman filter: the visual PnP result serves as the position state update filter for position observation, and the integral of the inertial measurement unit's angular velocity serves as the attitude increment update filter for attitude state.
[0080] During the brief visual lockout in step three, the extended Kalman filter uses only inertial measurement unit (IMU) data to drive inertial navigation for pose extrapolation. The position error calculated by the IMU alone diverges squared over time, accumulating to approximately 0.5° within 2 seconds. This deviation remains within the field of view of the visible light monitoring camera module. Compared to a completely unpredictable mechanical locking scheme, this significantly improves both the convergence speed of visual recovery and the interruption time of communication recovery.
[0081] Example 3: Remote Update of AI Model
[0082] This embodiment, based on Embodiment 1, provides an implementation method for remotely pushing updates to the AI model via a free-space optical communication link.
[0083] When the backend monitoring center receiving terminal analyzes the received monitoring video frames and finds that the detection accuracy of a certain type of security violation has dropped below a preset threshold, and it is necessary to update the AI detection model deployed at the front end, the backend monitoring center receiving terminal will push the updated target detection model weight file to the front-end portable communication monitoring terminal through the free space optical communication downlink.
[0084] The model weight file is approximately 20MB in size, and at a communication rate of 500Mbps on a free-space optical communication link, the complete transmission takes only about 0.3 seconds. If the free-space optical communication link is interrupted during the push process, the front-end portable communication monitoring terminal will feed back the currently received file offset to the back-end monitoring center receiving terminal via a control frame cached locally after the link is restored. The back-end monitoring center receiving terminal will then continue pushing from that offset position, achieving breakpoint resume transmission.
[0085] Finally, it should be noted that the above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A wireless monitoring method for temporary power supply areas on construction sites based on free-space optical communication and AI vision linkage, applied to a system comprising at least one front-end terminal and one back-end terminal, wherein each terminal is equipped with a communication-vision shared optical path integrated gimbal module, characterized in that, Includes the following steps: Step 1: AI Vision Active Alignment The integrated gimbal module of the first terminal uses its monitoring camera to capture image frames containing the cooperation identifier on the second terminal; it uses AI vision algorithms to detect multiple constituent light points of the cooperation identifier in real time from the image frames, calculates the six-degree-of-freedom spatial pose offset of the cooperation identifier relative to the optical axis of the first terminal, and generates gimbal servo compensation commands to drive the gimbal of the first terminal, so that the geometric center of the cooperation identifier is kept at the preset optical center position of the image frame; wherein, the communication-vision common optical path integrated gimbal module transmits the visible light band to the monitoring camera and reflects the near-infrared band to the free space optical communication transceiver unit through a multi-band beam splitter, so that the visual optical axis and the communication optical axis are coaxial behind the objective lens; Step 2: Free-space optical communication transmission Under the line-of-sight alignment established and maintained in step one, the first terminal and the second terminal communicate bidirectionally in the atmospheric channel via the free-space optical communication transceiver unit using the near-infrared band, carrying monitoring video streams, PTZ control commands and AI model data. Step 3: Predictive Preservation During Communication Interruption When the cooperative identifier fails to perform visual detection in step one due to temporary occlusion, the current spatial position of the cooperative identifier is predicted by extrapolation of the motion model based on the pose history sequence of the cooperative identifier calculated from the last N frames before the occlusion. The gimbal is then driven to be preset to the predicted position and held. When the occlusion is removed and the cooperative identifier becomes visible again, visual detection is restored and the free space optical communication link is rebuilt.
2. The wireless monitoring method according to claim 1, characterized in that, The communication-vision common optical path integrated gimbal module includes: a motorized zoom objective lens, a multi-band beam splitter prism, a visible light monitoring camera module, a free-space optical communication transceiver module, and a dual-axis servo gimbal; the multi-band beam splitter prism has a transmittance of not less than 92% in the 400nm to 700nm band and a reflectance of not less than 95% in the 780nm to 1600nm band.
3. The wireless monitoring method according to claim 1, characterized in that, The cooperation logo consists of 16 infrared LEDs arranged in a 4×4 square array, with a spacing of 40mm between adjacent LEDs; the infrared LEDs emit light at a wavelength of 850nm, pulse at a frequency of 200Hz, and a pulse width of 2ms; the 16 infrared LEDs are divided into four groups, and each group is encoded with a unique device identifier for the cooperation logo using a preset amplitude modulation mode.
4. The wireless monitoring method according to claim 1, characterized in that, In step one, the AI vision algorithm uses a two-way parallel strategy to detect the cooperative identifier: one path uses a target detection network to identify candidate regions of light spots in the image frame, and the other path uses prior knowledge of the pulse frequency of the cooperative identifier to perform differential analysis on two consecutive image frames to enhance the pulse light source and suppress the background; after spatial matching and confidence-weighted fusion, the candidate regions output by the two paths are only confirmed as valid light spots when the fusion confidence is higher than 0.
85.
5. The wireless monitoring method according to claim 1, characterized in that, The six-degree-of-freedom spatial pose offset mentioned in step one is calculated using the Perspective-n-Point algorithm. When the number of detected constituent light points is not less than 8, the translation vector and rotation vector are calculated using the two-dimensional pixel coordinates of the constituent light points and their corresponding three-dimensional world coordinates on the physical panel of the cooperative identifier. The translation vector includes horizontal offset, vertical offset and axial distance.
6. The wireless monitoring method according to claim 1, characterized in that, In step one, after the cooperative identifier is successfully detected stably for 10 consecutive frames, the system switches to sparse optical flow tracking mode. In sparse optical flow tracking mode, based on the detected light point positions in the previous frame, the corresponding positions in the current frame are predicted using the inverse optical flow method, and the predicted positions are used as inputs for pose calculation. The system only reverts to the complete detection pipeline when the number of tracked light points drops to less than 8 or the pose calculation residual after 5 consecutive frames of tracking increases by more than 50% compared to the beginning of tracking.
7. The wireless monitoring method according to claim 1, characterized in that, The motion model described in step three is a constant acceleration model, and its expression is x(t) = x0 + v0·t + ½·a·t 2 Where x represents the spatial position of the cooperative identifier, t represents the time elapsed since the moment of visual loss of lock, x0 represents the spatial position of the cooperative identifier at the moment of visual loss of lock, v0 represents the motion velocity of the cooperative identifier at the moment of visual loss of lock, and a represents the motion acceleration of the cooperative identifier; x0, v0, and a are obtained by weighted least squares fitting estimation of the last 16 frames of effective pose data in a circular history buffer of 32 frames, with the data of the most recent frames being given higher fitting weights.
8. The wireless monitoring method according to claim 1, characterized in that, The free-space optical communication transceiver unit continuously monitors the received optical power indication and packet error rate at the physical layer. When the received optical power indication drops below a preset threshold or the packet error rate rises above a preset limit within five consecutive time slots, the second terminal sends a link quality warning control frame to the first terminal via the free-space optical communication downlink. After receiving the link quality early warning control frame, the first terminal increases the execution frequency of visual detection in step one from 30Hz to 60Hz and switches the monitoring video encoding format from H.265 to MJPEG.
9. A wireless monitoring system for temporary power supply areas at construction sites based on free-space optical communication and AI vision linkage, comprising at least one front-end portable communication monitoring terminal and one back-end monitoring center receiving terminal, each terminal being equipped with a communication-vision common optical path integrated gimbal module, wherein the communication-vision common optical path integrated gimbal module comprises an electric zoom objective lens, a multi-band beam splitter prism, a visible light monitoring camera module, a free-space optical communication transceiver module, and a dual-axis servo gimbal, characterized in that, The system is configured to perform the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method as described in any one of claims 1 to 8.