A remote personnel positioning and high-definition image monitoring system based on a 5G network
By employing technologies such as 5G NR-UWB composite positioning, dynamic aperture 8K cameras, and lightweight federated learning, the system solves the problems of positioning accuracy and image transmission latency in complex environments of traditional monitoring systems. It constructs a high-precision, low-latency, and interference-resistant remote monitoring system, achieving sub-meter-level positioning, 8K resolution image transmission, and privacy protection, and is suitable for scenarios such as smart security and industrial inspection.
Patent Information
- Application Number
- CN202510857983.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Traditional monitoring systems suffer from low positioning accuracy, high image transmission latency, and weak anti-interference capabilities in scenarios such as emergency command and warehousing logistics, making it difficult to meet the real-time and accurate monitoring needs in dynamic and complex environments. Furthermore, existing 5G solutions suffer from low efficiency in multimodal data fusion, high bandwidth consumption in high-definition video transmission, and inadequate privacy protection.
Employing technologies such as 5G NR-UWB composite positioning, dynamic aperture 8K camera, lightweight federated learning framework, deep reinforcement learning spectrum decision-making, multimodal data fusion, and energy management module, combined with spatiotemporal joint beamforming, super-resolution image processing, dynamic bit rate adaptation, virtual anchor point technology, adaptive light field reconstruction, differential privacy protection, and multi-source energy harvesting, a high-precision, low-latency, and interference-resistant remote monitoring system is constructed.
It achieves sub-meter level positioning accuracy, stabilizes 8K resolution image transmission latency within 30ms, improves communication stability, enhances privacy protection, extends system battery life, and improves emergency response efficiency, meeting the high requirements of scenarios such as smart security and industrial inspection.
Smart Images

Figure CN120379028B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of 5G intelligent monitoring and positioning technology, and in particular to a remote personnel positioning and high-definition image monitoring system based on a 5G network. BACKGROUND
[0002] With the rapid development of smart cities, industrial internet and public safety needs, remote personnel positioning and high-definition image monitoring systems are increasingly widely used in emergency command, warehouse logistics, smart parks and other scenarios. Traditional monitoring systems rely on short-range communication technologies such as WiFi or ZigBee, and have problems such as low positioning accuracy (meter-level error), high image transmission delay (hundreds of milliseconds), weak anti-interference ability, and are difficult to meet the real-time and accurate monitoring needs in dynamic and complex environments. For example, in large shopping malls or factories, traditional systems cannot track the precise location of personnel in real time, and the image transmission delay increases significantly across floors, resulting in delayed emergency response.
[0003] Although the commercialization of 5G networks has brought breakthroughs in remote monitoring, existing solutions still face multiple challenges: multi-modal data fusion efficiency is low, making it difficult to achieve spatio-temporal alignment of positioning and images; high-definition video transmission consumes a large amount of bandwidth, easily causing network congestion; positioning accuracy drops sharply in complex scenarios (such as multi-path obstruction, electromagnetic interference), and privacy protection mechanisms are not perfect. For example, traditional 5G positioning only uses time of arrival (TOA) in a single dimension, with errors of several meters in non-line-of-sight (NLOS) environments, while uncompressed 8K video transmission requires more than 1 Gbps bandwidth, far exceeding the carrying capacity of most 5G slices.
[0004] Emerging applications such as unmanned inspection and AR remote command require higher system requirements: positioning accuracy needs to reach sub-millimeter or even centimeter level, video transmission delay < 50 ms, while ensuring the security of multi-user privacy data. Traditional technologies lack spatio-temporal joint processing, intelligent resource scheduling and anti-interference coordination mechanisms, and cannot meet the above requirements. It is urgent to develop fusion positioning, intelligent image processing and dynamic networking technologies based on 5G networks to break through existing bottlenecks, build a high-precision, low-latency, high-reliability remote monitoring system, and promote the deep landing of intelligent applications. SUMMARY
[0005] The present application proposes a remote personnel positioning and high-definition image monitoring system based on a 5G network to solve the problems mentioned in the above prior art.
[0006] To achieve the above purpose, the present application adopts the following technical solutions:
[0007] A remote personnel positioning and high-definition image monitoring system based on a 5G network includes the following modules:
[0008] 5G fusion positioning module: integrates 5G NR and UWB composite positioning unit, enhances target direction signal strength through space-time domain joint beamforming technology, combines multipath component separation algorithm, realizes sub-meter level positioning accuracy of 0.05-0.3 meters, supports distributed collaborative positioning architecture, uses virtual anchor point technology and 2.6GHz, 3.5GHz, millimeter wave multi-frequency point joint ranging, fuses multi-source distance data through Kalman filter, improves positioning stability in complex environment;
[0009] Intelligent image acquisition and processing module: equipped with dynamic aperture 8K camera, built-in adaptive light field reconstruction unit, reconstructs full scene depth map by collecting light information at different angles, combines super-resolution neural network to improve 1080p image to 8K resolution in real time, integrates dynamic scene perception unit, identifies environment type in real time based on semantic segmentation network, automatically optimizes exposure parameters through brightness gradient analysis algorithm, expands image dynamic range to 120dB;
[0010] 5G edge computing and transmission module: deploy lightweight federated learning framework, exchange model parameters between edge nodes through differential privacy protection mechanism, use dynamic code rate adaptation algorithm to adjust video transmission code rate in real time according to network congestion index, realize elastic expansion of slice resources combined with container orchestration technology, ensure that 8K video end-to-end transmission delay is stable within 30ms, and at the same time, through lightweight compression technology, the bandwidth demand is compressed to 10-100Mbps.
[0011] Further, it also includes:
[0012] Dynamic networking and anti-interference module: build cognitive radio network architecture, dynamically select the optimal channel based on deep reinforcement learning spectrum decision algorithm, combine adaptive interference alignment technology, update precoding matrix in real time through channel state information prediction algorithm, use distributed space-time coding technology to enhance signal anti-fading ability;
[0013] Multi-modal data fusion module: develop space-time consistency fusion framework, calibrate the space-time deviation of positioning and image data based on Bayesian network, integrate emotion recognition engine, improve emotion classification accuracy through the fusion algorithm of micro-expression analysis and voice emotion recognition; design behavior intention prediction model, use long short-term memory network to predict future 3-second personnel behavior trajectory.
[0014] Further, the dynamic networking and anti-interference module updates the precoding matrix through the channel state information prediction algorithm.
[0015] Further, the 5G fusion positioning module further includes: distributed collaborative positioning architecture, uses virtual anchor point technology and multi-frequency point joint ranging, combines Kalman filter algorithm to fuse distance estimation data.
[0016] Further, the intelligent image acquisition and processing module further comprises: a dynamic scene perception unit, which realizes real-time identification of scene types based on a semantic segmentation network; and an adaptive exposure compensation system, which dynamically adjusts exposure parameters through a brightness gradient analysis algorithm.
[0017] Further, the 5G edge computing and transmission module further comprises: a slice resource elastic allocation mechanism, which realizes resource on-demand expansion based on a container orchestration technology, and the expansion delay is less than 2 seconds; and a lightweight network function virtualization platform, which deploys core network functions as microservices, and the service startup time is less than 100 ms.
[0018] Further, the intelligent early warning and response module further comprises: a group behavior analysis engine, which calculates personnel interaction force based on a social force model and predicts the risk of stampede; and an augmented reality command system, which projects a three-dimensional scene to a command center in real time through a space mapping technology.
[0019] Further, the energy management module is designed to: design an energy collection-computation collaborative optimization framework, optimize an energy distribution strategy based on a Markov decision process, combine a solar energy, vibration energy and thermal energy multi-source heterogeneous energy collection system and a magnetic coupling resonant wireless energy transmission technology, and dynamically adjust the power consumption of the module through an intelligent sleep scheduling algorithm.
[0020] Further, the energy management module further comprises: an intelligent sleep scheduling algorithm, which dynamically adjusts the sleep period of the module based on a task prediction model, optimizes energy consumption according to the remaining power and the amount of tasks, and prolongs the system endurance time.
[0021] Further, the user interaction terminal is developed to: develop a haptic feedback command interface, realize vibration feedback of different frequencies and intensities through a piezoelectric ceramic array, integrate a three-dimensional holographic projection module, and support 360° viewing angle observation of a monitoring scene by adopting light field display technology, with a field of view angle of 120°.
[0022] Compared with the prior art, the present application has the following advantages:
[0023] In terms of positioning accuracy, the 5G NR-UWB composite positioning and multipath separation algorithm is adopted, combined with the spatio-temporal domain joint beamforming technology optimized by federated learning, to realize sub-meter level positioning of 0.05-0.3 meters, accurately track the moving track of personnel in complex indoor and outdoor environments, and meet the warehouse shelf level positioning requirements.
[0024] The intelligent image acquisition and processing module, through a dynamic aperture camera and adaptive light field reconstruction technology, achieves full-scene depth perception and real-time super-resolution reconstruction, upscaling 1080p images to 8K resolution while expanding the dynamic range to 120dB. Even in strong light, backlight, or nighttime environments, it can still clearly capture facial features and behavioral details, improving video analysis accuracy. Lightweight compression algorithms and dynamic bitrate adaptation technology compress 8K video transmission bandwidth to 10-100Mbps. Combined with 5G slicing technology, end-to-end latency is consistently below 30ms, meeting the stringent requirements of remote real-time command.
[0025] The dynamic networking and anti-interference module improves spectrum utilization and reduces the bit error rate through a deep reinforcement learning-based spectrum decision-making algorithm and adaptive interference alignment technology, maintaining communication stability even in highly interference-prone scenarios such as factories and exhibitions. The multimodal data fusion framework uses Bayesian networks to achieve spatiotemporal calibration of positioning and image data with a calibration accuracy of ±0.1 meters. Furthermore, through emotion recognition and behavioral intent prediction technologies, it provides 3-second advance warnings of abnormal behaviors (such as falls or gatherings), improving warning accuracy.
[0026] In terms of privacy protection and energy management, the federated learning framework ensures that data does not leave the local storage location through differential privacy technology, reducing the risk of privacy leaks. A multi-source energy harvesting system and intelligent sleep scheduling algorithm extend the system's battery life, allowing it to operate continuously for over 72 hours without an external power source. The user interaction terminal integrates haptic feedback and holographic projection technology, improving operational accuracy. The command center can monitor the real-time situation through a 360° holographic view, enhancing emergency response efficiency.
[0027] Overall, this application constructs a 5G remote monitoring system that integrates high-precision positioning, intelligent image processing, anti-interference communication, and privacy protection. It breaks through the technical bottlenecks of traditional solutions in terms of accuracy, latency, and reliability, and can be widely used in scenarios such as smart security, industrial inspection, and emergency rescue, providing key technical support for digital transformation. Attached Figure Description
[0028] Fig. 1 This is a schematic block diagram of a remote personnel positioning and high-definition image monitoring system based on a 5G network proposed in this invention.
[0029] Fig. 2 A bar chart comparing the accuracy of different positioning methods in complex scenarios;
[0030] Fig. 3 This is a line graph showing the bitrate of this system under different network congestion conditions;
[0031] Fig. 4 Radar charts showing the spectrum utilization of different schemes in an industrial environment. Detailed Implementation
[0032] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0033] The present application will be further described in detail below with reference to the drawings.
[0034] Referring to Figs. 1 to 4 A remote personnel positioning and high-definition image monitoring system based on a 5G network includes the following modules:
[0035] The 5G fusion positioning module adopts an innovative heterogeneous integration design to organically combine advanced communication and positioning technologies. The 5G NR baseband chip is a key component, which supports 2.6 GHz and 3.5 GHz FDD frequency bands and has excellent data transmission capability with a peak rate of 2.5 Gbps, capable of quickly and stably transmitting a large amount of data. The UWB pulse radio module has a center frequency of 5.8 GHz and a bandwidth of 500 MHz. The clock synchronization of 5G NR and UWB is achieved through a hardware timestamp alignment mechanism. The nanosecond-level timestamp of the UWB module is calibrated with the 10 ms frame structure of the 5G base station through a synchronization signal (PSS / SSS), with an error controlled within ±10 ns. The system introduces Kalman filtering to predict clock drift, updating synchronization parameters every 100 ms to ensure that the time deviation of ranging data is less than 20 ns, corresponding to a distance error of less than 6 mm. In multi-frequency point joint ranging, a layered fusion strategy is adopted to address the differences in propagation characteristics between 2.6 GHz and millimeter waves. Outdoor scenes mainly use 2.6 GHz (coverage range > 100 meters), while indoor scenes switch to millimeter waves (accuracy ≤ 0.05 meters). The frequency bands are automatically switched by an environment classifier based on signal strength variance, and the different frequency band data are normalized in scale (multiplied by the distance attenuation coefficient) to eliminate the influence of non-line-of-sight attenuation differences.
[0036] The dual-antenna array adopts a uniform linear array (ULA) arrangement with an antenna spacing of 0.5 meters, which is approximately 5 wavelengths. This layout optimizes the reception and transmission of signals. At the same time, FPGA (Field Programmable Gate Array) is used to calculate beamforming weights in real time, allowing the signal to be accurately adjusted according to actual needs.
[0037] In the spatial-temporal domain joint beamforming, the minimum variance distortion response (MVDR) algorithm is used. For the target azimuth angle θ, the system accurately calculates the steering vector. The steering vector formula is: where is a column vector, is the antenna spacing, is the wavelength, = 8, is the array element, is the imaginary unit, is the phase difference between different antenna elements, represents the transpose operation, i.e., converting a row vector to a column vector.
[0038] The receiver uses a RAKE receiver structure with 4 fingers to detect multipath components by adaptive thresholding. The multipath delay estimation is based on the generalized cross-correlation (GCC) algorithm, whose formula is: where is the estimated multipath delay, : the value of the received signal at sampling time , s(t) is the reference signal, = 1024 sampling points, represents finding the value of that maximizes the following summation, represents the possible time delay is the sampling point index. is the value of the reference signal at time . The state update is combined with Kalman filtering, and the state vector contains two-dimensional position (x, y) and velocity (v x , v y ), the state transition matrix , is the time interval for state update, and the measurement matrix , achieve 0.05 meters (line-of-sight scenario) to 0.3 meters (non-line-of-sight scenario) positioning accuracy. In the warehouse shelf environment test, the non-line-of-sight positioning error is reduced compared with the traditional TOA algorithm. The positioning accuracy test is extended to the dense metal obstacle scene (such as industrial plant steel structure), and when using the UWB positioning system, the non-line-of-sight error is 0.8 meters, which is reduced by 68% compared with the traditional TOA algorithm (error 2.5 meters); in the line-of-sight scenario, the error is stable at 0.05-0.1 meters. The 8K video transmission delay is 0.9 when the network congestion index CI=0 is adjusted by the dynamic code rate adaptation algorithm (adjustment range 50-150Mbps), and the delay fluctuation is controlled between 25-35ms, which is improved by 56% compared with the traditional fixed code rate scheme (delay fluctuation 40-80ms). The emotion recognition engine in the mixed modal test, the classification accuracy of micro-expression and voice fusion is 89.2%, which is improved by 8.6% and 13.6% respectively compared with single-mode vision (82.1%) and single-mode voice (78.5%). In the test set containing 1000 samples, the misjudgment rate of complex emotions such as "anxiety" and "focus" is reduced from 22% to 11%, which verifies the effectiveness of multi-modal fusion.
[0039] When the space-time domain joint beamforming and multipath separation are combined, the system first estimates the multipath time delay by a generalized cross-correlation algorithm, and selects the effective multipath components with energy higher than the noise threshold. For each effective multipath, the azimuth angle is calculated according to the time delay, and the beamforming weight is generated to enhance the target multipath signal and suppress interference. Beamforming and time delay estimation use a pipeline: after each time delay estimation (about 5 milliseconds), the beam weight is updated immediately to ensure that the beam is always aligned with the main path signal in scenarios such as personnel movement.
[0040] The implementation process of the virtual anchor point technology is as follows: adjacent tag nodes periodically exchange ranging data (interval 100 milliseconds) through the UWB module, and collect distance measurement values (accuracy ≤0.2 meters) of neighbor nodes within a radius of 5 meters. When constructing the geometric matrix, the tag itself is taken as the origin, and the relative distance of the neighbor nodes is converted into polar coordinate observation equation, and each row corresponds to one ranging, and the column elements are the conversion coefficients from polar coordinates to Cartesian coordinates. The virtual anchor point coordinates are solved by weighted least squares method, and the high-precision ranging data is given priority in weight, and finally a virtual anchor point grid of 1 per 10 square meters is generated.
[0041] Supports a star and mesh hybrid networking architecture, under which a single anchor point can cooperate with 32 tag nodes, greatly expanding the system's coverage range and node capacity. Virtual anchor point technology is one of its core highlights. By measuring between adjacent tags, virtual reference points are generated, which provide more dimensional information for positioning. The system uses multi-frequency point joint ranging, with ranging error controlled at ≤0.2 meters at 2.6 GHz frequency band, and ≤0.05 meters at millimeter wave frequency band, ensuring the accuracy of distance measurement.
[0042] In data fusion processing, the weighted least squares method is used, and the formula is . Among them, is the unknown parameter to be estimated finally, such as the position coordinates of the target and other information. is the geometric matrix, which reflects the geometric relationship between the measurement data and the unknown parameters, and its element value is related to factors such as the geometric layout of the measurement. is the observation vector, which contains the actual measurement data, such as the distance values obtained at each measurement point. W is the weight matrix, whose diagonal elements are , is the ranging standard deviation, and the weight matrix is used to give different weights to different measurement data. The smaller the ranging standard deviation (the more accurate the measurement), the greater the corresponding weight. By fusing distance data in this way, the positioning accuracy can be effectively improved.
[0043] In a dynamic environment, the system performs excellently, with a positioning update rate of up to 20 Hz, which can quickly respond to changes in the target's position, and the trajectory smoothness is significantly improved, providing stable and accurate positioning services for various application scenarios.
[0044] Intelligent image acquisition and processing module: equipped with a customized 8K camera module, with a super-high resolution of 7680x4320, capable of capturing 30 frames of pictures per second, and can present extremely delicate and smooth images. The module uses a 1 / 1.2-inch back-illuminated CMOS sensor, with a single pixel area of 2.4μm, which can fully capture light and improve image quality. The dynamic aperture range is from F1.4 to F16, which can flexibly adapt to different lighting environments. The electronic global shutter function is powerful, with exposure time freely adjustable between 1 / 10000 seconds and 30 seconds, capable of capturing instantaneous high-speed motion or recording long light changes with ease.
[0045] The module is built-in 9-axis IMU, with gyro zero bias stability reaching 50μg, and accelerometer noise density of 30μg / √Hz, which can accurately perceive the motion state of the device. Through the hardware synchronization interface, the image and inertial data timestamp can be precisely aligned, with an error of less than 100μs, providing a reliable foundation for subsequent image stabilization and analysis.
[0046] 4D light field sensor uses microlens array technology, and 37x37 microlenses are embedded between the main lens and the sensor. The spacing between the microlenses is 0.5mm. Each microlens corresponds to a 16x16 pixel sub-aperture, which can collect the direction (u, v) and position (s, t) information of the light, and then construct the Plenoptic function L(u, v, s, t). (u, v) is the light direction coordinate, u represents the horizontal angle, and v represents the vertical angle. (s, t) is the light position coordinate, s represents the horizontal position of the sensor plane, and t represents the vertical position of the sensor plane. It provides data support for realizing rich light field applications.
[0047] The super-resolution processing uses an improved ESRGAN network, which contains 32 residual dense blocks (RDB). This network can convert 1080p images to 8K output through 4 times upsampling. Compared with traditional Bicubic interpolation, the PSNR is improved by 3.2dB, and the modulation transfer function MTF50 reaches 0.72 at 20lp / mm, which is close to the optical limit, significantly improving the clarity and sharpness of the image.
[0048] The dynamic scene perception unit is based on the DeepLabv3+ network (backbone Xception). After pre-training on the COCO-Stuff dataset, it can classify 12 types of scenes such as industrial plant, hospital corridor, and outdoor street in real time. The DeepLabv3+ network used by the dynamic scene perception unit is pre-trained based on the COCO-Stuff dataset, and 1000 labeled images (including plant equipment, storage shelves, etc.) are fine-tuned for industrial scenes. The inference delay of the model on the FPGA is 25ms, which meets the real-time processing requirements.
[0049] The super-resolution ESRGAN network uses perceptual loss optimization during training, combined with pixel loss to improve image detail restoration. Data enhancement strategies include random rotation, Gaussian blur, low light simulation, and motion blur, covering common imaging interference scenarios. The training process uses progressive learning, first training on low-resolution images for 50 rounds, then gradually increasing to full-size input, and finally achieving 8K output with significantly improved image clarity compared to traditional interpolation algorithms, with a peak signal-to-noise ratio (PSNR) of 3.2dB. The adaptive exposure system analyzes the brightness gradient histogram to divide the image into 16x16 sub-blocks, calculates the mean brightness μ and standard deviation σ of each block, and dynamically adjusts the exposure time T: where =0.5 is the adjustment coefficient, is the initial exposure time, exp() is the exponential function, is the mean brightness of the image sub-block, is the standard deviation of the brightness of the image sub-block, =128 is the reference luminance.
[0050] 5G edge computing and transmission module: a lightweight federated learning framework supports 100 edge nodes for collaborative training, using the FedAvg algorithm, and the global model is aggregated every 10 iterations. Differential privacy is achieved by adding Laplace noise, with a privacy budget ε = 3. The value of the privacy budget ε = 3 is based on the privacy-accuracy balance experiment in the industrial scenario: when ε ∈ (0, 1], the model accuracy loss exceeds 25% (such as the F1 score from 0.92 to 0.68 in the device failure prediction task); when ε = 3, the accuracy loss is controlled within 5% (F1 score 0.87), and the noise distribution is adjusted adaptively (dynamic scaling of gradient norm upper bound Δ) to avoid masking effective parameters. The edge node uses an incremental gradient aggregation strategy, transmitting only the parameter update amount (average 32 KB) per iteration, combined with channel coding (Hamming distance 7) to reduce transmission errors. In a 5G channel with a signal-to-noise ratio of 15 dB, the model convergence speed is improved by 30% compared to the standard FedAvg, verifying the feasibility of the parameter settings. Noise scale Δ / ε (Δ is the gradient norm upper bound), edge nodes are deployed on the DU side of the 5G base station CU / DU separation architecture (distance from user <1 km), supporting localized model update delay <500 ms. The edge node model aggregation frequency is every 10 iterations (about 200 milliseconds), and through model compression technology (8-bit quantization and sparsification), the parameter transmission amount is compressed from 128 MB to 32 MB, and the single transmission bandwidth occupancy is reduced to 512 kbps, adapting to the low latency demand of 5G network slicing (end-to-end delay <500 ms).
[0051] Dynamic bit rate adaptation algorithm monitors network congestion index CI (comprehensive RTT, packet loss rate, queue length) in real time, and adjusts the bit rate using a PID controller: where , the proportional coefficient =0.1, the integral coefficient =0.01, and the differential coefficient =0.05. denotes the adjusted bit rate at the th time, is the error value at the th time, is the reference network congestion index, is the actual network congestion index monitored at the th time. The 8K video is compressed by BPG (Better Portable Graphics) with a bit rate range of 50-150 Mbps, combined with container orchestration technology (Kubernetes), to achieve elastic expansion of slicing resources from 1 vCPU / 1 GB memory to 8 vCPU / 8 GB memory, with an expansion delay of 1.8 seconds.
[0052] In the present application, further comprising:
[0053] Dynamic networking and anti-interference module: deep reinforcement learning adopts DDPG algorithm, the Actor network structure is 3 layers full connection (256-128-64), the Critic network is 3 layers full connection (256-128-1), the state space contains 12-dimensional features such as channel power, interference temperature, signal-to-noise ratio, the action space is 64 channels (divided within 200MHz bandwidth). The reward function is designed as: 8K video is compressed by BPG (Better Portable Graphics) with code rate range 50-150Mbps, combined with container orchestration technology (Kubernetes), realize the elastic expansion of slice resources from 1vCPU / 1GB memory to 8vCPU / 8GB memory, the expansion delay is 1.8 seconds.
[0054] In the present application, further comprising:
[0055] Dynamic networking and anti-interference module: deep reinforcement learning adopts DDPG algorithm, the Actor network structure is 3 layers full connection (256-128-64), the Critic network is 3 layers full connection (256-128-1), the state space contains 12-dimensional features such as channel power, interference temperature, signal-to-noise ratio, the action space is 64 channels (divided within 200MHz bandwidth). The reward function is designed as: , wherein, represents the reward value, is the natural logarithm, which is a quantitative index for measuring the decision-making of the intelligent agent in the dynamic networking and anti-interference task. SINR is the signal-to-interference-plus-noise ratio, which reflects the channel quality, and the larger the value is, the smaller the signal is affected by interference and noise; SwitchCost is the channel switching cost, which represents the resource loss and other costs caused by switching channels; BandwidthWaste is the bandwidth waste, which reflects the situation of the bandwidth not effectively used in the process of using the channel. Through 100,000 training iterations, the channel state prediction adopts LSTM network (128 hidden layer units), the prediction delay is 1ms, and the precoding matrix update period is 2ms, which reduces the delay compared with the traditional feedback mechanism.
[0056] The adaptive interference alignment technology uses channel state information (CSI) for precoding. Facing N interference sources, the system will carefully design the precoding matrix , which belongs to the complex domain . Here, M represents the number of antennas, which is 8 antennas; , which is 4 data streams. By designing this precoding matrix, the interference signal can be projected to zero in the null space of the receiving end, thereby effectively suppressing interference and improving the purity of the signal.
[0057] Distributed space-time coding adopts the classic Alamouti orthogonal design, and the transmission signal matrix is wherein and are transmission signals, and an asterisk * represents a complex conjugate. In the common and challenging communication environment of Rayleigh fading channel, the coding mode has obvious advantages, the coding gain can reach 6dB, and the signal strength can be effectively enhanced. When the signal-to-noise ratio SNR is 10dB, the bit error rate is only 8.2*10 −7 Compared with the uncoded scheme, the bit error rate is reduced by two orders of magnitude, greatly improving the reliability and accuracy of communication, and providing a solid guarantee for high-quality data transmission.
[0058] In the present application, it also includes:
[0059] A multi-modal data fusion module: a Bayesian network is constructed to contain a probability graph model of positioning error (Gaussian distribution, = 0.3 meters), camera extrinsic error (rotation angle = 0.5°, translation = 0.2 meters), is the standard deviation of the positioning error, is the standard deviation of the rotation angle error, is the standard deviation of the translation error, and the posterior probability is iteratively updated by a particle filter algorithm (1000 particles): wherein, represents the posterior probability of the position-related parameter x and the camera extrinsic-related parameter c under the condition of observing z, that is, the probability estimation of these parameters after integrating the observation information. is a likelihood function, representing the probability of observing z when the position and camera extrinsic parameters are known. P(x) is the prior probability of the position-related parameter x, reflecting the probability cognition of the position parameter before the observation information is obtained. P(c) is the prior probability of the camera extrinsic-related parameter c, representing the probability estimation of the camera extrinsic without observation. After calibration, the positioning error and the image pixel space error are ≤0.1 meters, and in the AR scene registration test, the registration error of the virtual object and the real scene is <2 pixels. The positioning module (update rate 20Hz) and the image module (frame rate 30fps) realize timestamp alignment through a hardware synchronization interface, with an error less than 100 microseconds. The system adopts a double-buffer mechanism: after the positioning data is written into the buffer, the 20Hz positioning trajectory is synchronized to the 30fps image frame rate through a linear interpolation algorithm (interval 50 milliseconds), ensuring that each image corresponds to an accurate space-time coordinate. During the Bayesian network calibration process, joint optimization is performed every 500 milliseconds, and the latest 10 frames of image visual features and positioning data are fused to dynamically adjust the space-time conversion parameters and eliminate cumulative errors.
[0060] When the dynamic networking module is linked with the edge computing module, the channel switching instruction output by the spectrum decision algorithm (such as DDPG) will trigger the slice resource reallocation mechanism of the edge node. When channel switching (such as the SINR being lower than a threshold due to interference) is detected, the system completes the elastic expansion of the slice resource from 1vCPU / 1GB memory to 8vCPU / 8GB memory within 1.8 seconds according to the bandwidth requirement (such as a 200MHz bandwidth channel in 64 channels) of the new channel through the Kubernetes container orchestration technology, so that the real-time performance of the code rate adaptation algorithm (PID controller) after channel switching is not affected by the resource bottleneck.
[0061] The emotion recognition engine adopts a double-flow network architecture: in the visual branch, the input is a micro-expression image of 48x48 pixels, and the powerful ResNet50 network is used to extract the key features in the image. Micro-expression contains rich emotional information, and this network can accurately capture these subtle details to provide visual evidence for emotion judgment. The speech branch takes the extracted 80-dimensional MFCC features as input and uses the LSTM network to extract the temporal features. The tone, speed and other time-varying information of the speech are crucial in emotional expression, and LSTM can effectively process these dynamic information.
[0062] The behavior intention prediction model adopts a 3-layer LSTM structure, each layer containing 128 units. It takes the past 5 seconds of trajectory points, including two-dimensional position (x, y) and velocity (v x ,v y ) as input data. Through deep analysis and learning of these historical trajectory information, the model can predict the trajectory points in the next 3 seconds. Through testing, its average error is only 0.42 meters, and the prediction accuracy for common behaviors such as "turning" and "staying" is more than 90%, showing high prediction accuracy and reliability.
[0063] In the present application, it also includes:
[0064] The energy management module: In terms of energy harvesting, the solar panel uses a flexible cadmium telluride thin film battery with an area of 0.05 square meters, which can stably output 11 milliwatts of power under standard lighting conditions, effectively converting solar energy into electrical energy. The piezoelectric vibrator uses PZT-5H material with a size of 10mm x 10mm x 2mm, which can generate 5 milliwatts of output power when placed in a 5Hz / 1g vibration environment, converting vibration energy in the environment into electrical energy. The thermoelectric generator sheet (TEG1-64065) uses temperature difference to generate electricity, which can output 20 milliwatts of power when there is a temperature difference of 20K, fully exploiting the energy contained in the temperature difference. The magnetic coupling resonance transmission system has a working frequency of 850kHz, and the diameter of the transmitting and receiving coils is 5cm, with a quality factor Q of 15, ensuring the efficiency and stability of energy transmission.
[0065] In energy allocation and management, a strategy based on Markov Decision Process (MDP) is adopted. The state space comprehensively considers multiple key factors, the remaining power is divided into 5 levels, which clearly reflects the power reserve situation of the device; the task priority is divided into 3 levels, which is convenient for reasonable arrangement of energy supply sequence; the environmental light is divided into 4 levels to adapt to energy allocation under different light conditions. The action space covers different power consumption modes of the module, including high, medium, low and sleep modes. The strategy is optimized by Q-Learning algorithm, which effectively reduces the average energy consumption.
[0066] The intelligent sleep scheduling algorithm dynamically adjusts the sleep cycle according to the task prediction model. The model is based on historical data and uses Poisson process for modeling. In continuous positioning mode, the positioning module wakes up every 100 milliseconds, with power consumption of only 0.5 milliwatts, and the endurance time is greatly improved to 36 hours, which is a significant improvement compared with the traditional scheme of 12 hours. The multi-source energy scheduling strategy dynamically switches priority combined with environmental sensor data: when the light is sufficient (≥2000 lux), the solar panel (0.05 square meters, output 11 milliwatts) is preferentially enabled, in the vibration environment (5Hz / 1g), the piezoelectric vibrator (PZT-5H material, output 5 milliwatts) is automatically activated, and when the temperature difference exceeds 15K, the thermoelectric power sheet (output 20 milliwatts) is started. Energy storage uses a 1200mAh lithium battery, supports 1.5 hour fast charging and trickle charging protection, and the charge and discharge management chip monitors the voltage (3.3-4.2V) and temperature (-20°C-60°C) in real time. When the output of a single energy source is less than 1 milliwatt, it automatically switches to battery power. The intelligent sleep scheduling algorithm records the past 24 hours task trigger time sequence in real time, and the Poisson process parameter λ is dynamically updated based on the last 1 hour data. The sleep threshold is set as follows: when the remaining power is ≤15%, the positioning module wake-up interval is extended from 100 milliseconds to 200 milliseconds; when the power is ≤10%, the non-critical modules such as image acquisition are turned off, only the positioning and wake-up circuit (power consumption 0.1 milliwatt) is retained, and a low power alarm is sent to the terminal. In the monitoring mode, the camera captures an image every 5 seconds, and the overall power consumption is 2 milliwatts, and the endurance can reach 72 hours. In sleep mode, only the clock and wake-up circuit are retained, with power consumption as low as 0.1 milliwatt, which can achieve standby for up to 15 days, greatly improving energy utilization efficiency and device endurance.
[0067] In the present application, it also includes:
[0068] User interaction terminal: The tactile feedback interface is a key interactive component, and its core is an 8x8 piezoelectric ceramic array. The array uses PZT-4 material, and each ceramic element has a diameter of 3mm and a resonance frequency of 20-200Hz. It is driven by an STM32F407 microcontroller through an SPI interface, with an update rate of up to 100Hz, which can quickly respond to operation instructions. Different operations correspond to unique vibration patterns. In the case of emergency stop, a high-frequency vibration of 200Hz is presented, with an amplitude of 50μm, an equivalent force of 5mN, and a duration of 200ms, allowing users to instantly perceive the emergency. When adjusting the path, switch to a medium-frequency vibration of 50Hz, with an amplitude of 30μm, a force of 3mN, and a period of 500ms, giving users clear path change prompts. Normal confirmation uses low-frequency vibration of 20Hz, with an amplitude of 10μm, a force of 1mN, and a single pulse lasting 100ms, providing simple confirmation feedback. The tactile feedback vibration pattern has been verified by 50 users: high-frequency vibration (200Hz) for emergency stop, 92% of users think the touch is strong and easy to identify; medium-frequency vibration (50Hz) for path adjustment, 85% of users feedback clear prompts; low-frequency vibration (20Hz) for operation confirmation, 78% of users express mild experience. The system supports custom vibration parameters (frequency 10-200Hz, amplitude 5-50μm, duration 50-500ms), and users can save personalized configurations through the terminal APP to adapt to different perception needs.
[0069] When deploying three-dimensional holographic projection, the audience's position is captured in real time through structured light scanning (accuracy 0.1mm), and multi-user independent perspective images are generated using binocular parallax algorithm. When detecting that more than 3 people are watching at the same time, automatically adjust the driving voltage of the microlens array (increase by 20%), widen the horizontal viewing angle from ±60° to ±75°, and reduce the image crosstalk of adjacent viewers through dynamic phase modulation to ensure that the clarity difference of each perspective is controlled within 15%.
[0070] The three-dimensional holographic projection module uses advanced light field display technology, consisting of 20 layers of carefully arranged liquid crystal panels, each with a 2mm spacing and a 0.3mm pixel pitch, and is equipped with a microlens array. Through time division multiplexing technology, it can skillfully display images of different perspectives, creating a realistic three-dimensional visual effect. Its field of view angle can reach 120° according to the formula: where =200mm is the panel width, =173mm is the focal length. This feature allows it to support 10 people watching at the same time, with a horizontal viewing angle range of ±60°. In critical scenarios such as emergency command, it can provide decision-makers with more intuitive and comprehensive information than traditional two-dimensional monitoring, greatly improving decision-making efficiency and effectively promoting the efficient implementation of emergency response and other work.
[0071] The application of the haptic feedback interface and the three-dimensional holographic projection technology in the user interaction terminal not only improves the convenience and accuracy of interaction, but also expands the dimension of information perception and processing of the user, and lays a solid foundation for application innovation in more fields in the future.
[0072] The above merely describes the preferred embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can make equivalent replacements or changes within the technical range disclosed by the present application according to the technical scheme and the inventive concept of the present application, which should be covered within the protection scope of the present application.
Claims
1. A remote personnel positioning and high-definition image monitoring system based on a 5G network, characterized in that, The following modules are included: 5G fusion positioning module: integrates 5G NR and UWB composite positioning unit, enhances target direction signal strength through space-time domain joint beamforming technology, combines multipath component separation algorithm to achieve sub-meter level positioning accuracy of 0.05-0.3 meters, supports distributed collaborative positioning architecture, uses virtual anchor point technology and 2.6GHz, 3.5GHz, millimeter wave multi-frequency joint ranging, fuses multi-source distance data through Kalman filter to improve positioning stability in complex environments; Intelligent image acquisition and processing module: equipped with dynamic aperture 8K camera, built-in adaptive light field reconstruction unit, reconstructs full scene depth map by collecting light information at different angles, combines super-resolution neural network to upgrade 1080p image to 8K resolution in real time, integrates dynamic scene perception unit, realizes real-time identification of environment type based on semantic segmentation network, automatically optimizes exposure parameters through brightness gradient analysis algorithm, expands image dynamic range to 120dB; 5G edge computing and transmission module: deploys lightweight federated learning framework, exchanges model parameters between edge nodes through differential privacy protection mechanism, uses dynamic code rate adaptation algorithm to adjust video transmission code rate in real time according to network congestion index, realizes elastic expansion of slicing resources through container orchestration technology, guarantees 8K video end-to-end transmission delay within 30ms, and simultaneously compresses bandwidth demand to 10-100Mbps through lightweight compression technology; Dynamic networking and anti-interference module: builds cognitive radio network architecture, dynamically selects the optimal channel based on deep reinforcement learning spectrum decision algorithm, combines adaptive interference alignment technology, updates precoding matrix in real time through channel state information prediction algorithm, uses distributed space-time coding technology to enhance signal anti-fading ability; Multi-modal data fusion module: develops spatio-temporal consistency fusion framework, calibrates the spatio-temporal deviation of positioning and image data based on Bayesian network, integrates emotion recognition engine, improves emotion classification accuracy through the fusion algorithm of micro-expression analysis and voice emotion recognition; designs behavior intention prediction model, predicts future 3-second personnel behavior trajectory using long short-term memory network; Intelligent early warning and response module: group behavior analysis engine calculates personnel interaction force based on social force model to predict the risk of stampede; augmented reality command system projects three-dimensional scene to command center in real time through space mapping technology; User interaction terminal: develops tactile feedback command interface, realizes vibration feedback of different frequencies and intensities through piezoelectric ceramic array, integrates three-dimensional holographic projection module, supports 360° viewing angle observation of monitoring scene using light field display technology, with a field of view angle of 120°.
2. The remote personnel locating and high definition image monitoring system based on 5G network of claim 1, wherein, The 5G fusion positioning module further includes: a distributed collaborative positioning architecture that uses virtual anchor point technology and multi-frequency joint ranging, and combines a Kalman filter algorithm to fuse distance estimation data.
3. The remote personnel locating and high definition image monitoring system based on 5G network of claim 1, wherein, The intelligent image acquisition and processing module further includes: a dynamic scene perception unit that realizes real-time identification of scene type based on a semantic segmentation network; and an adaptive exposure compensation system that dynamically adjusts exposure parameters through a brightness gradient analysis algorithm.
4. The remote personnel locating and high definition image monitoring system based on 5G network of claim 1, wherein, The 5G edge computing and transmission module further comprises a slice resource elasticity allocation mechanism, which realizes resource on-demand expansion based on container orchestration technology, and the expansion delay is less than 2 seconds; and a lightweight network function virtualization platform, which deploys core network functions as microservices, and the service startup time is less than 100 ms.
5. The remote personnel locating and high definition image monitoring system based on 5G network according to any one of claims 1-4, characterized in that, Further comprising: An energy management module: design an energy collection-computation collaborative optimization framework, optimize the energy distribution strategy based on Markov decision process, combine a solar energy, vibration energy, thermal energy multi-source heterogeneous energy collection system, and a magnetic coupling resonant wireless energy transmission technology, and dynamically adjust the module power consumption through an intelligent sleep scheduling algorithm. 6.The remote personnel positioning and high-definition image monitoring system based on a 5G network according to claim 5, characterized in that, The energy management module further comprises an intelligent sleep scheduling algorithm, which dynamically adjusts the module sleep period based on a task prediction model, optimizes the energy consumption according to the residual power and task quantity prediction, and prolongs the system endurance time.
Citation Information
Patent Citations
5G-oriented integrated positioning system and 5G-oriented integrated positioning method fused with UWB
CN111343571A
Method for realizing exception control and real-time monitoring based on probe technology
CN119892699A
Camera monitoring system based on wireless communication and communication protocol optimization method
CN119893264A