Remote personnel positioning and high-definition image monitoring system based on 5G network

Through 5GNR-UWB composite positioning, dynamic aperture cameras and federated learning technologies, a high-precision and low-latency 5G remote monitoring system is built, which solves the positioning accuracy and image transmission delay problems of traditional monitoring systems in complex environments, and achieves efficient remote monitoring and privacy protection.

CN120379028AActive Publication Date: 2025-07-25甘肃省公安厅
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510857983.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-25
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

Traditional monitoring systems have low positioning accuracy, high image transmission delay and weak anti-interference ability in complex environments, making it difficult to meet the needs of high-precision and low-latency remote monitoring, and their privacy protection is incomplete.

Method used

5GNR-UWB composite positioning, dynamic aperture camera, federated learning framework, dynamic code rate adaptation, deep reinforcement learning and multimodal data fusion technology are adopted, and a high-precision and low-latency 5G remote monitoring system is built with virtual anchor points, super-resolution reconstruction, lightweight compression and adaptive interference alignment.

Benefits of technology

It realizes sub-meter positioning accuracy, 8K high-definition image real-time transmission, and low-latency communication, improves emergency response efficiency, ensures privacy and security, and extends the system battery life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120379028A_ABST
    Figure CN120379028A_ABST
Patent Text Reader

Abstract

The invention discloses a remote personnel positioning and high-definition image monitoring system based on a 5G network, and relates to the technical field of 5G intelligent monitoring and positioning, and the system comprises a 5G fusion positioning module which employs a composite positioning and multipath separation technology; the intelligent image acquisition and processing module is matched with a dynamic aperture camera and super-resolution reconstruction, and supports 8K image quality and full-scene perception; and the 5G edge calculation and transmission module comprises a federated learning framework and dynamic code rate adaptation, and the system is further provided with a dynamic networking module, a multi-mode fusion module, an intelligent early warning module and the like, so that the anti-interference and cooperative capability is improved. According to the invention, 5G fusion positioning and 8K high-definition monitoring are realized, the anti-interference capability is strong, intelligent early warning, privacy protection and long endurance are realized, the system is suitable for scenes such as intelligent security and industrial inspection, the emergency response efficiency is greatly improved, and the 5G intelligent application is promoted to the ground.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of 5G intelligent monitoring and positioning, and particularly to a remote personnel positioning and high-definition image monitoring system based on a 5G network. Background Art

[0002] With the rapid development of the needs of smart cities, industrial Internet, and public safety, remote personnel positioning and high-definition image monitoring systems are increasingly widely used in scenarios such as emergency command, warehousing and logistics, and smart parks. Traditional monitoring systems rely on short-distance communication technologies such as WiFi or ZigBee, and have problems such as low positioning accuracy (meter-level error), high image transmission delay (hundreds of milliseconds), and weak anti-interference ability, making it difficult to meet the real-time and accurate monitoring requirements in dynamic and complex environments. For example, in large shopping malls or factories, traditional systems cannot track the exact location of personnel in real time, and the delay significantly increases when images are transmitted across floors, resulting in a lag in emergency response.

[0003] Although the commercialization of 5G networks has brought breakthroughs to remote monitoring, existing solutions still face multiple challenges: low efficiency of multi-modal data fusion, making it difficult to achieve spatio-temporal alignment of positioning and images; high-definition video transmission consumes a large amount of bandwidth, easily causing network congestion; positioning accuracy drops sharply in complex scenarios (such as multipath occlusion and electromagnetic interference), and the privacy protection mechanism is imperfect. For example, traditional 5G positioning only uses a single dimension of time of arrival (TOA), and the error can reach several meters in a non-line-of-sight (NLOS) environment, while uncompressed 8K video transmission requires more than 1 Gbps of bandwidth, far exceeding the carrying capacity of most 5G slices.

[0004] Emerging applications such as unmanned patrol and AR remote command have put forward higher requirements for the system: the positioning accuracy needs to reach sub-meter level or even centimeter level, the video transmission delay < 50 ms, and at the same time, the privacy data security of multiple users needs to be guaranteed. Due to the lack of spatio-temporal domain joint processing, intelligent resource scheduling, and anti-interference cooperation mechanisms, traditional technologies can no longer meet the above requirements. There is an urgent need to develop fusion positioning, intelligent image processing, and dynamic networking technologies based on 5G networks to break through the existing bottlenecks, build a high-precision, low-latency, and highly reliable remote monitoring system, and promote the in-depth implementation of intelligent applications. Summary of the Invention

[0005] A remote personnel positioning and high-definition image monitoring system based on a 5G network proposed by the present invention is used to solve the problems mentioned in the above existing technologies.

[0006] To achieve the above object, the present invention adopts the following technical solutions: A remote personnel positioning and high-definition image monitoring system based on a 5G network includes the following modules: 5G Fusion Positioning Module: Integrates 5G NR and UWB composite positioning units, enhances the signal strength of the target direction through spatio-temporal domain joint beamforming technology, combines the multi-path component separation algorithm to achieve sub-meter positioning accuracy of 0.05 - 0.3 meters, supports a distributed collaborative positioning architecture, uses virtual anchor point technology and multi-frequency joint ranging of 2.6 GHz, 3.5 GHz, and millimeter waves, and fuses multi-source distance data through Kalman filtering to improve positioning stability in complex environments; Intelligent Image Acquisition and Processing Module: Equipped with a dynamic aperture 8K camera, built-in adaptive light field reconstruction unit, reconstructs the full-scene depth map by collecting light information from different angles, combines with a super-resolution neural network to instantly enhance 1080p images to 8K resolution, integrates a dynamic scene perception unit, real-time identifies the environment type based on a semantic segmentation network, and automatically optimizes exposure parameters through a brightness gradient analysis algorithm to extend the image dynamic range to 120 dB; 5G Edge Computing and Transmission Module: Deploys a lightweight federated learning framework, exchanges model parameters between edge nodes through a differential privacy protection mechanism, adopts a dynamic bitrate adaptation algorithm to adjust the video transmission bitrate in real-time according to the network congestion index, combines container orchestration technology to achieve elastic expansion of slice resources, ensures that the end-to-end transmission delay of 8K video is stable within 30 ms, and at the same time compresses the bandwidth requirement to 10 - 100 Mbps through lightweight compression technology.

[0007] Furthermore, it also includes: Dynamic Networking and Anti-Interference Module: Constructs a cognitive radio network architecture, dynamically selects the optimal channel based on a spectrum decision algorithm of deep reinforcement learning, combines adaptive interference alignment technology, updates the precoding matrix in real-time through a channel state information prediction algorithm, and adopts a distributed space-time coding technology to enhance the signal anti-fading ability; Multi-modal Data Fusion Module: Develops a spatio-temporal consistency fusion framework, calibrates the spatio-temporal deviation between positioning and image data based on a Bayesian network, integrates an emotion recognition engine, and improves the emotion classification accuracy through a fusion algorithm of micro-expression analysis and speech emotion recognition; designs a behavior intention prediction model, and uses a long short-term memory network to predict the future 3-second personnel behavior trajectory.

[0008] Furthermore, the dynamic networking and anti-interference module: updates the precoding matrix through a channel state information prediction algorithm.

[0009] Furthermore, the 5G fusion positioning module also includes: a distributed collaborative positioning architecture, uses virtual anchor point technology and multi-frequency joint ranging, and combines the Kalman filtering algorithm to fuse distance estimation data.

[0010] Furthermore, the intelligent image acquisition and processing module further includes: a dynamic scene perception unit that realizes real-time recognition of scene types based on a semantic segmentation network; an adaptive exposure compensation system that dynamically adjusts exposure parameters through a brightness gradient analysis algorithm.

[0011] Furthermore, the 5G edge computing and transmission module further includes: a slice resource elastic allocation mechanism that realizes on-demand resource expansion based on container orchestration technology, with an expansion delay of less than 2 seconds; a lightweight network function virtualization platform that deploys core network functions as microservices, with a service startup time of less than 100 ms.

[0012] Furthermore, the intelligent early warning and response module further includes: a crowd behavior analysis engine that calculates the interaction force of personnel based on the social force model to predict the risk of stampede; an augmented reality command system that projects a three-dimensional scene onto the command center in real time through spatial mapping technology.

[0013] Furthermore, it further includes: an energy management module: designing an energy harvesting-computation collaborative optimization framework, optimizing the energy allocation strategy based on the Markov decision process, combining a multi-source heterogeneous energy harvesting system of solar energy, vibration energy, and thermal energy, as well as magnetic coupling resonance wireless energy transfer technology, and dynamically adjusting the module power consumption through an intelligent sleep scheduling algorithm.

[0014] Furthermore, the energy management module further includes: an intelligent sleep scheduling algorithm that dynamically adjusts the module sleep cycle based on a task prediction model, predicts and optimizes the energy consumption according to the remaining battery power and the task volume, and extends the system battery life.

[0015] Furthermore, it further includes: a user interaction terminal: developing a tactile feedback command interface that realizes vibration feedback of different frequencies and intensities through a piezoelectric ceramic array, integrating a three-dimensional holographic projection module, and adopting a light field display technology to support 360° viewing of the monitoring scene, with a field of view angle of 120°.

[0016] Compared with the existing technologies, the beneficial effects of the present invention are: In terms of positioning accuracy, by adopting the 5GNR-UWB composite positioning and multipath separation algorithm and combining the space-time domain joint beamforming technology optimized by federated learning, sub-meter-level positioning of 0.05 - 0.3 meters is achieved, which can accurately track the movement trajectory of personnel in complex indoor and outdoor environments and meet the positioning requirements at the warehouse shelf level.

[0017] The intelligent image acquisition and processing module realizes full-scene depth perception and real-time super-resolution reconstruction through a dynamic aperture camera and adaptive light field reconstruction technology, boosts 1080p images to 8K resolution, and expands the dynamic range to 120dB simultaneously. Even in strong light, backlight, or night environments, it can still clearly capture the facial features and behavioral details of people, improving the accuracy of video analysis. The lightweight compression algorithm and dynamic bitrate adaptation technology compress the 8K video transmission bandwidth to 10 - 100Mbps. Combined with 5G slicing technology, the end-to-end latency is stably within 30ms, meeting the stringent requirements of remote real-time command.

[0018] The dynamic networking and anti-interference module improves the spectrum utilization rate and reduces the bit error rate to through a spectrum decision algorithm based on deep reinforcement learning and adaptive interference alignment technology, and can still maintain communication stability in strong interference scenarios such as factories and exhibitions. The multi-modal data fusion framework realizes the spatio-temporal calibration of positioning and image data based on a Bayesian network, with a calibration accuracy of ±0.1 meters, and through emotion recognition and behavior intention prediction technology, it can give early warnings of abnormal behaviors (such as falls and gatherings) 3 seconds in advance, improving the warning accuracy.

[0019] In terms of privacy protection and energy management, the federated learning framework ensures that data does not leave the local through differential privacy technology, reducing the risk of privacy leakage; the multi-source energy collection system and intelligent sleep scheduling algorithm extend the system's battery life, and it can work continuously for more than 72 hours in scenarios without external power supply. The user interaction terminal integrates tactile feedback and holographic projection technology, improving the operation accuracy. The command center can grasp the on-site dynamics in real time through a 360° holographic scene, enhancing the emergency response efficiency.

[0020] Overall, this application constructs a 5G remote monitoring system integrating high-precision positioning, intelligent image processing, anti-interference communication, and privacy protection, breaking through the technical bottlenecks of traditional solutions in terms of accuracy, latency, and reliability, and can be widely applied to scenarios such as intelligent security, industrial inspection, and emergency rescue, providing key technical support for digital transformation. Brief Description of the Drawings

[0021] Figure 1 It is a schematic block diagram of a remote personnel positioning and high-definition image monitoring system based on a 5G network proposed by the present invention; Figure 2 It is a bar chart comparing the accuracies of different positioning methods in complex scenarios; Figure 3 It is a line chart of the bitrate of this system under different network congestions; Figure 4 It is a radar chart of the spectrum utilization rate of different solutions in an industrial environment. Detailed Description of the Invention

[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0023] Next, the present invention will be further described in detail with reference to the accompanying drawings.

[0024] Refer to Figures 1 to 4 : A remote personnel positioning and high-definition image monitoring system based on a 5G network, including the following modules: 5G Converged Positioning Module: Adopting an innovative heterogeneous integration design, it organically combines advanced communication and positioning technologies. Among them, the 5GNR baseband chip is a key component, which supports the 2.6GHz and 3.5GHz FDD frequency bands, has excellent data transmission capabilities, with a peak rate of up to 2.5Gbps, and can quickly and stably transmit a large amount of data. The UWB impulse radio module has a center frequency of 5.8GHz and a bandwidth of 500MHz. The clock synchronization between 5GNR and UWB is achieved through a hardware timestamp alignment mechanism: the nanosecond-level timestamp of the UWB module and the 10ms frame structure of the 5G base station are calibrated through synchronization signals (PSS / SSS), and the error is controlled within ±10ns. The system introduces Kalman filtering to predict clock drift and updates the synchronization parameters every 100ms to ensure that the time deviation of the ranging data < 20ns, corresponding to a distance error < 6mm. When performing multi-frequency point joint ranging, in view of the propagation characteristic differences between 2.6GHz and millimeter waves, a hierarchical fusion strategy is adopted: the outdoor scenario is mainly based on 2.6GHz (coverage range > 100 meters), and the indoor scenario switches to millimeter waves (accuracy ≤ 0.05 meters). The frequency band is automatically switched through an environment classifier (based on signal strength variance). When fusing, the data of different frequency bands are scale-normalized (multiplied by the distance attenuation coefficient) to eliminate the influence of non-line-of-sight attenuation differences.

[0025] The dual antenna array adopts a uniform linear array (ULA) method, and the antenna spacing is carefully set to 0.5 meters, approximately 5 wavelengths. This layout can optimize the signal reception and transmission effects. At the same time, it cooperates with the FPGA (Field Programmable Gate Array) to calculate the beamforming weights in real time, enabling the signal to be precisely adjusted according to actual needs.

[0026] In terms of spatio-temporal domain joint beamforming, it is achieved through the Minimum Variance Distortionless Response (MVDR) algorithm. For the target azimuth angle θ, the system will accurately calculate the steering vector. The formula for the steering vector is: , where is the steering vector, which is a column vector, is the antenna spacing, is the wavelength, = 8, is the array element, is the imaginary unit, is the exponential part in the steering vector formula, representing the phase difference generated by the signal between different antenna elements, represents the transpose operation, that is, converting a row vector into a column vector.

[0027] The receiving end adopts a RAKE receiver structure, configures 4 relevant receiving branches, and detects multipath components through adaptive threshold detection. The multipath delay estimation is based on the Generalized Cross-Correlation (GCC) algorithm, and the formula is: , where is the estimated multipath delay, : the value of the received signal at the sampling time , s(t) is the reference signal, = 1024 sampling points, means to find the value that can make the subsequent summation formula reach the maximum value, represents the possible delay is the sampling point index. is the value of the reference signal at time. Combining with Kalman filtering for state update, the state vector includes two-dimensional position (x, y) and velocity (v x , v y ), the state transition matrix , is the time interval for state update, the measurement matrix , achieving a positioning accuracy of 0.05 m (line-of-sight scenario) to 0.3 m (non-line-of-sight scenario). In the warehouse shelf environment test, the non-line-of-sight positioning error is reduced compared with the traditional TOA algorithm. The positioning accuracy test is extended to the dense metal obstacle scenario (such as the steel frame structure of an industrial plant). When using the UWB positioning of this system, the non-line-of-sight error is 0.8 m, which is 68% lower than the traditional TOA algorithm (error 2.5 m); the error in the line-of-sight scenario is stable at 0.05 - 0.1 m. The 8K video transmission delay is controlled within 25 - 35 ms through the dynamic bitrate adaptation algorithm (adjustment range 50 - 150 Mbps) when the network congestion index CI = 0.9, which improves the stability by 56% compared with the traditional fixed bitrate scheme (delay fluctuation 40 - 80 ms). In the mixed-modal test of the emotion recognition engine, the classification accuracy after the fusion of micro-expression and speech reaches 89.2%, which is 8.6% and 13.6% higher than that of single-modal vision (82.1%) and single-modal speech (78.5%) respectively. In the test set containing 1000 samples, the misjudgment rate of complex emotions such as "anxiety" and "concentration" is reduced from 22% in the single-modal to 11%, verifying the effectiveness of multi-modal fusion.

[0028] When the spatio-temporal domain joint beamforming is combined with multipath separation, the system first estimates the multipath delay through the generalized cross-correlation algorithm, and filters out the effective multipath components with energy higher than the noise threshold. For each effective multipath, its azimuth angle is calculated according to the delay, and beamforming weights are generated to enhance the target multipath signal and suppress interference. The beamforming and delay estimation are coordinated in a pipelined manner: every time a delay estimation is completed (about 5 milliseconds), the beam weights are immediately updated to ensure that the beam is always aligned with the main path signal in scenarios such as personnel movement.

[0029] The implementation process of the virtual anchor point technology is as follows: Adjacent tag nodes periodically exchange ranging data through the UWB module (with an interval of 100 milliseconds), and collect the distance measurement values of neighbor nodes within a radius of 5 meters (with an accuracy of ≤0.2 meters). When constructing the geometric matrix, with the tag itself as the origin, the relative distances of neighbor nodes are converted into polar coordinate observation equations, and each row corresponds to a ranging, and the column elements are the conversion coefficients from polar coordinates to Cartesian coordinates. The virtual anchor point coordinates are solved by the weighted least squares method, and the weights give priority to trusting high-precision ranging data, and finally a virtual anchor point grid with 1 virtual anchor point per 10 square meters is generated.

[0030] It supports a hybrid star and mesh networking architecture. Under this architecture, a single anchor point can cooperate with 32 tag nodes, greatly expanding the coverage range and node capacity of the system. The virtual anchor point technology is one of its core highlights. By measuring each other between adjacent tags, virtual reference points are generated, and these virtual reference points provide more dimensional information for positioning. The system uses multi-frequency joint ranging, and the ranging error can be controlled within ≤0.2 meters in the 2.6GHz band, and even reaches a high precision of ≤0.05 meters in the millimeter wave band, ensuring the accuracy of distance measurement.

[0031] In data fusion processing, the weighted least squares method is adopted, and the formula is . Among them, are the unknown parameters to be finally estimated, such as the position coordinates of the target and other information. is the geometric matrix, which reflects the geometric relationship between the measurement data and the unknown parameters, and its element values are related to factors such as the geometric layout of the measurement. is the observation vector, which contains the actually measured data, such as the distance values obtained at each measurement point, etc. W is the weight matrix, and its diagonal elements are , is the ranging standard deviation. The weight matrix is used to assign different weights to different measurement data. The smaller the ranging standard deviation (the more accurate the measurement), the greater the corresponding weight. By fusing the distance data in this way, the positioning accuracy can be effectively improved.

[0032] In a dynamic environment, the system performs excellently. The positioning update rate can reach 20Hz, enabling it to quickly respond to the position changes of targets. At the same time, the trajectory smoothness is significantly improved, providing stable and accurate positioning services for various application scenarios.

[0033] Intelligent Image Acquisition and Processing Module: Equipped with a customized 8K camera module, it has an ultra-high resolution of 7680×4320 and can capture 30 frames per second, capable of presenting extremely delicate and smooth images. The module uses a 1 / 1.2-inch back-illuminated CMOS sensor with a single pixel area of 2.4μm, which can fully capture light and improve the imaging quality. The dynamic aperture range is from F1.4 to F16, enabling it to flexibly adapt to different lighting environments. The electronic global shutter function is powerful, and the exposure time can be freely adjusted between 1 / 10000 second and 30 seconds, being able to handle both capturing fleeting high-speed movements and recording long-term light and shadow changes with ease.

[0034] The module is built-in with a 9-axis IMU. The gyroscope zero-bias stability reaches 50μg, and the accelerometer noise density is 30μg / √Hz, which can accurately sense the motion state of the device. Through the hardware synchronization interface, precise alignment of the image and inertial data timestamps can be achieved, with an error less than 100μs, providing a reliable basis for subsequent image stabilization and analysis.

[0035] The 4D light field sensor uses the microlens array technology. A 37×37 microlens is carefully embedded between the main lens and the sensor, and the microlens pitch is 0.5mm. Each microlens corresponds to a 16×16 pixel sub-aperture, which can collect the direction (u, v) and position (s, t) information of light, and then construct the Plenoptic function L(u, v, s, t). (u, v) is the light direction coordinate, where u represents the horizontal direction angle and v represents the vertical direction angle; (s, t) is the light position coordinate, where s represents the horizontal position on the sensor plane and t represents the vertical position on the sensor plane, providing data support for realizing rich light field applications.

[0036] The super-resolution processing adopts an improved ESRGAN network, which contains 32 residual dense blocks (RDB). This network can convert a 1080p image into an 8K output through 4 times upsampling. Compared with traditional Bicubic interpolation, the PSNR is increased by 3.2dB, and the modulation transfer function MTF50 reaches 0.72 at 20lp / mm, with the performance approaching the optical limit, significantly improving the clarity and sharpness of the image.

[0037] The dynamic scene perception unit is based on the DeepLabv3+ network (with Xception as the backbone). After pre-training on the COCO-Stuff dataset, it can classify 12 types of scenes in real-time, such as industrial plants, hospital corridors, and outdoor streets. The DeepLabv3+ network adopted by the dynamic scene perception unit is pre-trained on the COCO-Stuff dataset and fine-tuned with an additional 1,000 annotated images (including plant equipment, storage shelves, etc.) for industrial scenarios. The inference latency of the model on the FPGA is 25 milliseconds, meeting the real-time processing requirements.

[0038] When training the super-resolution ESRGAN network, perceptual loss is used for optimization, combined with pixel loss to enhance the ability to restore image details. The data augmentation strategies include random rotation, Gaussian blur, low-light simulation, and motion blur, covering common imaging interference scenarios. Progressive learning is adopted in the training process. First, it is trained on low-resolution images for 50 epochs, then gradually increased to full-size input. Finally, the image quality clarity of the 8K output is significantly improved compared with the traditional interpolation algorithm, and the peak signal-to-noise ratio (PSNR) is increased by 3.2 dB. The adaptive exposure system analyzes the luminance gradient histogram, divides the image into 16×16 sub-blocks, calculates the average luminance μ and standard deviation σ of each block, and dynamically adjusts the exposure time T: , where =0.5 is the adjustment coefficient, is the initial exposure time, exp() is the exponential function, is the average luminance of the image sub-block, is the standard deviation of the image sub-block luminance, =128 is the reference luminance.

[0039] 5G Edge Computing and Transmission Module: The lightweight federated learning framework supports collaborative training of 100 edge nodes. The FedAvg algorithm is adopted, and the global model is aggregated every 10 rounds of iteration. Differential privacy is achieved by adding Laplace noise, with the privacy budget ε = 3. The value of the privacy budget ε = 3 is based on the privacy-accuracy balance experiment in industrial scenarios: when ε ∈ (0, 1], the model accuracy loss exceeds 25% (e.g., the F1 score drops from 0.92 to 0.68 in the device fault prediction task); when ε = 3, the accuracy loss is controlled within 5% (F1 score 0.87), and the noise distribution is adaptively adjusted (dynamically scaling the upper bound Δ of the gradient norm) to avoid masking valid parameters. The edge nodes adopt an incremental gradient aggregation strategy, and only transmit the parameter updates (average 32KB) in each round of iteration. Combining channel coding (Hamming distance 7) reduces the transmission error. In a 5G channel with a signal-to-noise ratio of 15dB, the model convergence speed is increased by 30% compared to the standard FedAvg, verifying the feasibility of the parameter settings. The noise scale Δ / ε (Δ is the upper bound of the gradient norm), and the edge nodes are deployed on the DU side of the 5G base station CU / DU separation architecture (distance from the user < 1km), supporting a local model update delay < 500ms. The edge node model aggregation frequency is once every 10 rounds of iteration (about 200 milliseconds). Through model compression techniques (8-bit quantization and sparsification), the parameter transmission volume is compressed from 128MB to 32MB, and the single-transmission bandwidth occupancy is reduced to 512kbps, adapting to the low-latency requirements of 5G network slicing (end-to-end delay < 500ms).

[0040] The dynamic bitrate adaptation algorithm monitors the network congestion index CI (combining RTT, packet loss rate, and queue length) in real time and adjusts the bitrate using a PID controller: where , the proportional coefficient = 0.1, the integral coefficient = 0.01, and the differential coefficient = 0.05. represents the adjusted bitrate at the th moment, is the error value at the th moment, is the reference network congestion index, is the actually monitored network congestion index at the th moment. After being compressed by BPG (Better Portable Graphics), the bitrate range of 8K videos is 50 - 150Mbps. Combining container orchestration technology (Kubernetes), elastic expansion of slice resources from 1vCPU / 1GB memory to 8vCPU / 8GB memory is achieved, and the expansion delay is 1.8 seconds.

[0041] In the present invention, it further includes: Dynamic networking and anti-interference module: Deep reinforcement learning uses the DDPG algorithm. The Actor network structure is a three-layer fully connected network (256-128-64), and the Critic network is a three-layer fully connected network (256-128-1). The state space contains 12-dimensional features such as channel power, interference temperature, and signal-to-noise ratio. The action space is 64 channels (divided within a 200MHz bandwidth). The reward function is designed as follows: After 8K video is compressed by BPG (Better Portable Graphics), the bit rate range is 50-150Mbps. Combining with container orchestration technology (Kubernetes), elastic expansion of slice resources from 1vCPU / 1GB memory to 8vCPU / 8GB memory is achieved, and the expansion delay is 1.8 seconds.

[0042] In the present invention, it further includes: Dynamic networking and anti-interference module: Deep reinforcement learning uses the DDPG algorithm. The Actor network structure is a three-layer fully connected network (256-128-64), and the Critic network is a three-layer fully connected network (256-128-1). The state space contains 12-dimensional features such as channel power, interference temperature, and signal-to-noise ratio. The action space is 64 channels (divided within a 200MHz bandwidth). The reward function is designed as: , where represents the reward value, is the natural logarithm, which is a quantitative index to measure the quality of the agent's decision-making in the dynamic networking and anti-interference task. SINR is the signal-to-interference-plus-noise ratio, which reflects the channel quality. The larger its value, the less the signal is affected by interference and noise. SwitchCost is the channel switching cost, which represents the costs such as resource loss caused by channel switching. BandwidthWaste is the amount of bandwidth waste, which reflects the situation of the bandwidth not being effectively utilized during the channel usage. Through 100,000 training iterations, the channel state prediction uses an LSTM network (128 hidden layer units), the prediction delay is 1ms, and the precoding matrix update period is 2ms, reducing the time delay compared to the traditional feedback mechanism.

[0043] The adaptive interference alignment technology uses the channel state information (CSI) for precoding. Facing N interference sources, the system will carefully design the precoding matrix , which belongs to the complex domain . Here, M represents the number of antennas, which is 8 antennas; represents the number of data streams, which is 4 data streams. By designing this precoding matrix, the interference signal can be projected to zero in the null space at the receiving end, thereby effectively suppressing interference and improving the purity of the signal.

[0044] Distributed space-time coding uses the classic Alamouti orthogonal design, and its transmitted signal matrix is . Where and are the transmitted signals, where the asterisk * represents the complex conjugate. In a common and challenging communication environment such as a Rayleigh fading channel, this coding method has significant advantages. The coding gain can reach 6 dB, effectively enhancing the signal strength. When the signal-to-noise ratio SNR is 10 dB, the bit error rate is only 8.2×10 −7 , compared with the uncoded scheme, the bit error rate is reduced by two orders of magnitude, greatly improving the reliability and accuracy of communication, and providing a solid guarantee for high-quality data transmission.

[0045] In the present invention, it further includes: Multi-modal data fusion module: The Bayesian network constructs a probability graph model including positioning error (Gaussian distribution, = 0.3 m), camera extrinsic parameter error (rotation angle = 0.5°, translation = 0.2 m). is the standard deviation of the positioning error, is the standard deviation of the rotation angle error, is the standard deviation of the translation error. After iterative updating of the posterior probability through the particle filter algorithm (1000 particles): , where represents the posterior probability of the position-related parameter x and the camera extrinsic parameter-related parameter c under the condition of observing z, that is, the probability estimation of these parameters after integrating the observation information. is the likelihood function, representing the probability of observing z when the position and camera extrinsic parameters are known. P(x) is the prior probability of the position-related parameter x, reflecting the probability cognition of the position parameter before obtaining the observation information. P(c) is the prior probability of the camera extrinsic parameter-related parameter c, representing the probability estimation of the camera extrinsic parameters without observation. After calibration, the positioning and image pixel space error ≤ 0.1 m. In the AR scene registration test, the registration error between the virtual object and the real scene < 2 pixels. The positioning module (update rate 20 Hz) and the image module (frame rate 30 fps) achieve timestamp alignment through the hardware synchronization interface, with an error less than 100 microseconds. The system adopts a double-buffer mechanism: After the positioning data is written into the buffer, the 20 Hz positioning trajectory is synchronized to the 30 fps image frame rate through the linear interpolation algorithm (interval 50 milliseconds), ensuring that each frame of image corresponds to accurate spatio-temporal coordinates. During the calibration process of the Bayesian network, joint optimization is performed every 500 milliseconds, fusing the visual features of the latest 10 frames of images and the positioning data, dynamically adjusting the spatio-temporal conversion parameters, and eliminating the cumulative error.

[0046] When the dynamic networking module interacts with the edge computing module, the channel switching instructions output by the spectrum decision algorithm (such as DDPG) will trigger the slice resource reallocation mechanism of the edge node. When a channel switch is detected (such as when the SINR is lower than the threshold due to interference), the system, according to the bandwidth requirements of the new channel (such as a 200MHz bandwidth channel among 64 channels), through the Kubernetes container orchestration technology, completes the elastic expansion of slice resources from 1vCPU / 1GB memory to 8vCPU / 8GB memory within 1.8 seconds, ensuring that the real-time performance of the code rate adaptation algorithm (PID controller) is not affected by resource bottlenecks after the channel switch.

[0047] The emotion recognition engine adopts a two-stream network architecture: On the visual branch, the input is a micro-expression image of 48×48 pixels, and a powerful ResNet50 network is used to extract the key features in the image. Micro-expressions contain rich emotional information, and this network can accurately capture these nuances, providing a visual basis for emotion judgment. The speech branch takes the extracted 80-dimensional MFCC features as input and uses an LSTM network to extract the temporal features therein. Information such as the intonation and speech rate of speech that changes over time is crucial in emotional expression, and LSTM can effectively process these dynamic information.

[0048] The behavior intention prediction model adopts a 3-layer LSTM structure, with each layer containing 128 units. It takes the trajectory points in the past 5 seconds, including two-dimensional positions (x, y) and speed (v x ,v y ) as input data. Through in-depth analysis and learning of this historical trajectory information, the model can predict the trajectory points in the next 3 seconds. After testing, its average error is only 0.42 meters, and the prediction accuracy for common behaviors such as "turning" and "staying" exceeds 90%, demonstrating extremely high prediction accuracy and reliability.

[0049] In the present invention, it also includes: Energy management module: In terms of energy harvesting, the solar panel uses a flexible cadmium telluride thin-film battery with an area of 0.05 square meters, which can stably output 11 milliwatts of power under standard lighting conditions, effectively converting solar energy into electrical energy. The piezoelectric vibrator selects PZT-5H material with a size of 10 mm × 10 mm × 2 mm, and when in a vibration environment of 5Hz / 1g, it can generate 5 milliwatts of output power, converting the vibration energy in the environment into electrical energy. The thermoelectric generator (TEG1-64065) generates electricity using temperature differences. When there is a temperature difference of 20K, it can output 20 milliwatts of power, fully exploiting the energy contained in the temperature difference. The magnetic coupling resonant transmission system has a working frequency of 850kHz, the diameters of both the transmitting and receiving coils are 5 cm, and the quality factor Q reaches 15, ensuring the high efficiency and stability of energy transmission.

[0050] In terms of energy allocation and management, a strategy based on the Markov Decision Process (MDP) is adopted. The state space comprehensively considers multiple key factors. The remaining battery power is divided into 5 levels, clearly reflecting the battery reserve situation of the device; the task priority is divided into 3 levels, facilitating the reasonable arrangement of the energy supply sequence; the environmental light is divided into 4 levels to adapt to energy allocation under different lighting conditions. The action space covers different power consumption modes of the module, including high, medium, low, and sleep modes. The strategy is optimized through the Q-Learning algorithm, effectively reducing the average energy consumption.

[0051] The intelligent sleep scheduling algorithm dynamically adjusts the sleep cycle based on the task prediction model. This model is based on historical data and is modeled using the Poisson process. In the continuous positioning mode, the positioning module wakes up once every 100 milliseconds, with a power consumption of only 0.5 milliwatt, and the battery life is significantly extended to 36 hours, showing a significant improvement compared to the 12 hours of the traditional solution. The multi-source energy scheduling strategy dynamically switches priorities in combination with environmental sensor data: when the light is sufficient (≥2000 lux), the solar panel (0.05 square meters, output 11 milliwatt) is preferentially enabled; in a vibrating environment (5 Hz / 1 g), the piezoelectric vibrator (PZT-5H material, output 5 milliwatt) is automatically activated; when the temperature difference exceeds 15 K, the thermoelectric generator is started (output 20 milliwatt). The energy storage uses a 1200 mAh lithium battery, supporting 1.5-hour fast charging and trickle charging protection. The charge and discharge management chip monitors the voltage (3.3 - 4.2 V) and temperature (-20°C to 60°C) in real time, and automatically switches to battery power supply when the output of a single energy source is lower than 1 milliwatt. The intelligent sleep scheduling algorithm records the task trigger time series in the past 24 hours in real time through the device, and the Poisson process parameter λ is dynamically updated based on the data in the most recent 1 hour. The sleep threshold is set as follows: when the remaining battery power ≤ 15%, the wake-up interval of the positioning module is extended from 100 milliseconds to 200 milliseconds; when the battery power ≤ 10%, non-critical modules such as image acquisition are turned off, and only the positioning and wake-up circuits (power consumption 0.1 milliwatt) are retained, and a low battery alarm is sent to the terminal. In the monitoring mode, the camera captures images once every 5 seconds, with an overall power consumption of 2 milliwatt, and the battery life can reach 72 hours. In the sleep mode, only the clock and wake-up circuits are retained, with a power consumption as low as 0.1 milliwatt, enabling a standby time of up to 15 days, greatly improving the energy utilization efficiency and the battery life of the device.

[0052] In the present invention, it further includes: User Interaction Terminal: The tactile feedback interface is a key interaction component, and its core is an 8×8 piezoelectric ceramic array. This array uses PZT-4 material, with each ceramic element having a diameter of 3mm and a resonance frequency of 20 - 200Hz. It is driven by an STM32F407 microcontroller through the SPI interface, with an update rate of up to 100Hz, enabling rapid response to operation instructions. Different operations correspond to unique vibration patterns. When in emergency stop, it vibrates at a high frequency of 200Hz, with an amplitude of 50μm, generating an equivalent force of 5mN and lasting for 200ms, allowing users to instantly perceive the emergency situation. When adjusting the path, it switches to a medium-frequency vibration of 50Hz, with an amplitude of 30μm, a force of 3mN, and a period of 500ms, giving users a clear path change prompt. For normal confirmation, it uses a low-frequency vibration of 20Hz, with an amplitude of 10μm, a force of 1mN, and a single pulse lasting for 100ms, providing a simple confirmation feedback. The tactile feedback vibration patterns have been verified through research with 50 users: high-frequency vibration (200Hz) is used for emergency stop, and 92% of users think the touch is strong and easy to identify; medium-frequency vibration (50Hz) is used for path adjustment, and 85% of users feedback that the prompt is clear; low-frequency vibration (20Hz) is used for operation confirmation, and 78% of users say the experience is mild. The system supports customizing vibration parameters (frequency 10 - 200Hz, amplitude 5 - 50μm, duration 50 - 500ms), and users can save personalized configurations through the terminal APP to adapt to different perception needs.

[0053] When deploying three-dimensional holographic projection, it uses structured light scanning (with an accuracy of 0.1mm) to capture the positions of the audience in real time, and uses the binocular disparity algorithm to generate multi-user independent perspective images. When it detects that more than 3 people are watching simultaneously, it automatically adjusts the driving voltage of the microlens array (increasing by 20%), widens the horizontal viewing angle from ±60° to ±75°, and reduces the crosstalk of adjacent audience images through dynamic phase modulation, ensuring that the clarity difference between each perspective is controlled within 15%.

[0054] The three-dimensional holographic projection module adopts advanced light field display technology and consists of 20 layers of carefully arranged liquid crystal panels, with a spacing of 2mm between each layer, a pixel pitch of 0.3mm, and is equipped with a microlens array. Through time-division multiplexing technology, it can skillfully display images from different perspectives, creating a realistic three-dimensional visual effect. Its field of view is calculated according to the formula: (where = 200mm is the panel width, = 173mm is the focal length), and can reach 120∘. This feature enables it to support 10 people to watch simultaneously, with a horizontal viewing angle range of ±60°. In key scenarios such as emergency command, compared with traditional two-dimensional monitoring, it can provide more intuitive and comprehensive information for decision-makers, greatly improving the decision-making efficiency and effectively promoting the efficient development of work such as emergency response.

[0055] The application of the tactile feedback interface and three-dimensional holographic projection technology in user interaction terminals not only improves the convenience and accuracy of interaction, but also expands the dimensions of users' information perception and processing, laying a solid foundation for application innovation in more fields in the future.

[0056] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A remote personnel positioning and high-definition image monitoring system based on a 5G network, characterized in that, It includes the following modules: 5G Converged Positioning Module: Integrates 5G NR and UWB composite positioning units, enhances the signal strength of the target direction through spatio-temporal domain joint beamforming technology, combines the multipath component separation algorithm, achieves sub-meter positioning accuracy of 0.05 - 0.3 meters, supports a distributed collaborative positioning architecture, utilizes virtual anchor point technology and multi-frequency joint ranging of 2.6 GHz, 3.5 GHz, and millimeter waves, and fuses multi-source distance data through Kalman filtering to improve positioning stability in complex environments; Intelligent Image Acquisition and Processing Module: Equipped with a dynamic aperture 8K camera, built-in adaptive light field reconstruction unit, reconstructs the full-scene depth map by collecting light information from different angles, combines with a super-resolution neural network to instantly upscale 1080p images to 8K resolution, integrates a dynamic scene perception unit, real-time identifies the environment type based on a semantic segmentation network, and automatically optimizes exposure parameters through a brightness gradient analysis algorithm to extend the image dynamic range to 120 dB; 5G Edge Computing and Transmission Module: Deploys a lightweight federated learning framework, exchanges model parameters between edge nodes through a differential privacy protection mechanism, adopts a dynamic bitrate adaptation algorithm to adjust the video transmission bitrate in real-time according to the network congestion index, combines container orchestration technology to achieve elastic expansion of slice resources, ensures that the end-to-end transmission delay of 8K video is stable within 30 ms, and at the same time compresses the bandwidth requirement to 10 - 100 Mbps through lightweight compression technology.

2. The remote personnel positioning and high-definition image monitoring system based on the 5G network according to claim 1, wherein It also includes: Dynamic Networking and Anti-Interference Module: Constructs a cognitive radio network architecture, dynamically selects the optimal channel based on a spectrum decision algorithm of deep reinforcement learning, combines with an adaptive interference alignment technology, updates the precoding matrix in real-time through a channel state information prediction algorithm, and adopts a distributed space-time coding technology to enhance the signal anti-fading ability; Multi-Modal Data Fusion Module: Develops a spatio-temporal consistency fusion framework, calibrates the spatio-temporal deviation between positioning and image data based on a Bayesian network, integrates an emotion recognition engine, and improves the emotion classification accuracy through a fusion algorithm of micro-expression analysis and speech emotion recognition; designs a behavior intention prediction model, and predicts the future 3-second personnel behavior trajectory using a long short-term memory network.

3. The remote personnel positioning and high-definition image monitoring system based on the 5G network according to claim 2, characterized in that, The dynamic networking and anti-interference module: Updates the precoding matrix through a channel state information prediction algorithm.

4. A remote personnel positioning and high-definition image monitoring system based on a 5G network according to claim 1, characterized in that The 5G converged positioning module also includes: A distributed collaborative positioning architecture, utilizes virtual anchor point technology and multi-frequency joint ranging, and combines with a Kalman filtering algorithm to fuse distance estimation data.

5. The remote personnel positioning and high-definition image monitoring system based on 5G network according to claim 1, characterized in that, The intelligent image acquisition and processing module also includes: A dynamic scene perception unit, which realizes real-time scene type recognition based on a semantic segmentation network; An adaptive exposure compensation system, which dynamically adjusts exposure parameters through a brightness gradient analysis algorithm.

6. A remote personnel positioning and high-definition image monitoring system based on a 5G network according to claim 1, characterized in that, The 5G edge computing and transmission module also includes: A slice resource elastic allocation mechanism, which realizes on-demand resource expansion based on container orchestration technology, and the expansion delay is less than 2 seconds; A lightweight network function virtualization platform, which deploys core network functions as microservices, and the service startup time is less than 100 ms.

7. A remote personnel positioning and high-definition image monitoring system based on a 5G network according to claim 2, characterized in that, The intelligent early warning and response module further includes: a crowd behavior analysis engine that calculates the interaction force of personnel based on the social force model to predict the risk of stampede; an augmented reality command system that projects a three-dimensional scene onto the command center in real time through spatial mapping technology.

8. A remote personnel positioning and high-definition image monitoring system based on a 5G network according to any one of claims 1-7, characterized in that, It further includes: An energy management module: designing an energy harvesting - computing collaborative optimization framework, optimizing the energy allocation strategy based on the Markov decision process, combining a multi-source heterogeneous energy harvesting system of solar energy, vibration energy, and thermal energy, as well as magnetic coupling resonance wireless energy transfer technology, and dynamically adjusting the module power consumption through an intelligent sleep scheduling algorithm.

9. The remote personnel positioning and high-definition image monitoring system based on the 5G network according to claim 8, characterized in that, The energy management module further includes: an intelligent sleep scheduling algorithm that dynamically adjusts the module sleep cycle based on a task prediction model, predicts and optimizes the energy consumption according to the remaining battery power and the amount of tasks, and extends the system battery life.

10. A remote personnel positioning and high-definition image monitoring system based on a 5G network according to claim 1, characterized in that, It further includes: A user interaction terminal: developing a tactile feedback command interface that realizes vibration feedback of different frequencies and intensities through a piezoelectric ceramic array, integrating a three-dimensional holographic projection module, and adopting light field display technology to support 360° viewing of the monitoring scene with a field of view angle of 120°.

Citation Information

Patent Citations

  • 5G-oriented integrated positioning system and 5G-oriented integrated positioning method fused with UWB

    CN111343571A

  • AI monitoring system based on personnel accurate positioning

    CN116016600A

  • Wireless positioning method, system and base station based on 5G fusion UWB

    CN117750300A

  • Intelligent security management system based on 5G, AI and UWB positioning technologies

    CN119450695A

  • Method for realizing exception control and real-time monitoring based on probe technology

    CN119892699A