Camera remote control method and system based on personal WiFi

By constructing a distributed network topology and using adaptive coding technology, the video stuttering and latency issues of remote control systems for cameras in portable WiFi environments were resolved, achieving stable and efficient video transmission and control response, and improving the real-time performance and reliability of the monitoring system.

CN120091222BActive Publication Date: 2026-03-24SHENZHEN NEW SAIBO TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-16
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Traditional remote camera control systems cannot effectively cope with dynamic changes in a portable WiFi network environment, resulting in video stuttering, control delays, and loss of critical information. Furthermore, the lack of a priority management mechanism affects the monitoring effect and real-time performance.

Method used

By constructing a distributed network topology, collecting camera node information and performing scene analysis, calculating the optimal encoding parameter configuration, dividing video frame priority regions, adopting differentiated compression processing and multi-level transmission queue management, and combining deep learning and adaptive encoding technology, the scheduling of data transmission and control commands is optimized.

Benefits of technology

Maintaining stable video transmission and control response under network fluctuation conditions significantly reduces video stuttering and control latency, ensures the quality of critical content, and improves the real-time performance and reliability of the monitoring system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091222B_ABST
    Figure CN120091222B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of camera remote control, and discloses a camera remote control method and system based on a personal WiFi. The method comprises the following steps: collecting position information, a coverage range and signal strength of a camera node, and constructing a distributed network topology structure; according to the distributed network topology structure, performing scene analysis on video content collected by the camera, and calculating an optimal encoding parameter configuration table; dividing video frames into different priority areas according to the optimal encoding parameter configuration table, performing differential compression processing, and forming video data packets with priority marks; distributing the video data packets with priority marks to a multi-level transmission queue, scheduling data transmission according to a current bandwidth state of the personal WiFi, and generating an orderly transmission stream. The application effectively balances the calculation load and energy consumption of each camera node, can still maintain stable video transmission and control response under network fluctuation conditions, and significantly reduces the video freezing rate and control delay.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of camera remote control, and in particular to a camera remote control method and system based on a personal WiFi. BACKGROUND

[0002] With the rapid popularization of security monitoring and remote video applications, distributed multi-camera monitoring systems have been widely researched and applied. Such a system usually connects multiple physically separated cameras to a central processor through a network to achieve joint video processing and all-around monitoring. However, traditional wired network connections have problems such as complex wiring, high installation cost, poor deployment flexibility, etc., while personal WiFi, as a portable wireless network solution, provides new possibilities for camera remote control. However, as a mobile network environment, the network state of personal WiFi is easily affected by factors such as environment, distance, obstacles, etc., resulting in large bandwidth fluctuations, unstable connection, high packet loss rate, etc. These uncertainties pose serious challenges to high-quality video transmission and precise control.

[0003] Existing camera remote control systems mostly use fixed parameter transmission strategies, which cannot effectively cope with the dynamic changes of the personal WiFi network environment. When the network condition deteriorates, the system either maintains high code rate transmission, resulting in severe video stuttering and control delay, or blindly reduces video quality, sacrificing monitoring effectiveness. In addition, traditional solutions often allocate limited bandwidth resources equally to all video content, ignoring the importance differences of different video area content, causing problems such as loss of critical information or excessive bandwidth resource occupation in non-critical areas. At the same time, the conventional camera control instruction transmission lacks a priority management mechanism, and in the case of network congestion, control instructions may be delayed or lost, seriously affecting the real-time and reliability of remote control. SUMMARY

[0004] The present application provides a camera remote control method and system based on personal WiFi, which effectively balances the computing load and energy consumption of each camera node, and still maintains stable video transmission and control response under network fluctuation conditions, significantly reducing video stuttering rate and control delay.

[0005] In a first aspect, the present application provides a camera remote control method based on personal WiFi, which comprises:

[0006] Collecting the position information, coverage range and signal strength of the camera nodes, and constructing a distributed network topology structure;

[0007] According to the distributed network topology structure, performing scene analysis on the video content collected by the cameras, and calculating an optimal encoding parameter configuration table;

[0008] According to the optimal encoding parameter configuration table, the video frame is divided into different priority areas, differential compression processing is performed, and a video data packet with a priority mark is formed;

[0009] The video data packet with the priority mark is distributed to a multi-level transmission queue, data transmission is scheduled according to a current bandwidth state of the personal WiFi, and an ordered transmission stream is generated.

[0010] In a second aspect, the application provides a camera remote control system based on personal WiFi, which comprises:

[0011] The acquisition module is configured to acquire position information, coverage range and signal strength of the camera node, and construct a distributed network topology structure.

[0012] The scene analysis module is configured to perform scene analysis on video content collected by the camera according to the distributed network topology structure, and calculate an optimal encoding parameter configuration table.

[0013] The compression processing module is configured to divide the video frame into different priority areas according to the optimal encoding parameter configuration table, perform differential compression processing, and form a video data packet with a priority mark.

[0014] The data transmission module is configured to distribute the video data packet with the priority mark to a multi-level transmission queue, schedule data transmission according to a current bandwidth state of the personal WiFi, and generate an ordered transmission stream.

[0015] In the technical solution provided by the application, by constructing a distributed network topology structure and monitoring the on-body WiFi connection state in real time, the system can quickly perceive network changes and make corresponding adjustments, and still maintain stable video transmission and control response under network fluctuation conditions, significantly reducing the video stutter rate and control delay. The adaptive encoding parameter configuration based on the network bandwidth quality model is combined with the differentiated compression strategy based on regional importance perception, so that the system can significantly reduce the overall data volume while ensuring the quality of key content. Through the multi-level transmission queue management and priority sorting mechanism, the control instructions can still be reliably transmitted under network limitation, and the gradual control strategy and real-time feedback mechanism are adopted to ensure the accuracy and real-time performance of remote control. The target detection technology based on deep learning is used to identify important regions in the video frame, and the differentiated encoding parameters and multi-level preprocessing operations are applied, combined with the super-resolution reconstruction algorithm and the anti-artifact filter processing network to deal with the low-quality regions caused by network limitation, which significantly improves the subjective and objective quality of the monitoring video. By designing the network exception handling strategy, adaptive segmented retransmission mechanism and forward error correction encoding technology, the system can effectively deal with network abnormal conditions such as signal strength drop and packet loss rate surge, and ensure the continuous operation of the monitoring system in poor network environment. By intelligently allocating computing and processing tasks, local preprocessing is performed on the camera terminal, data aggregation is performed on the edge processing unit, and video synthesis is performed on the central processor, forming a layered computing architecture, which effectively balances the computing load and energy consumption of each camera node, and improves the overall operation efficiency of the system. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0017] Figure 1 The flowchart of the camera remote control method based on on-body WiFi provided by the embodiment of the present application is shown.

[0018] Figure 2 The structural schematic block diagram of the camera remote control system based on on-body WiFi provided by the embodiment of the present application is shown. DETAILED DESCRIPTION

[0019] With reference to the accompanying drawings, the technical solutions in the embodiments of the present application will be clearly and completely described below, obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work belong to the protection scope of the present application.

[0020] The flow chart shown in the drawings is only an example, and does not necessarily include all the contents and operations / steps, and is not necessarily executed in the described order. For example, some operations / steps can also be decomposed, combined or partially merged, so that the actual execution order can be changed based on the actual situation.

[0021] It should also be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and the appended claims of the present application, unless otherwise clear from the context, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0022] It should be further understood that the term "and / or" used in the specification and the appended claims of the present application means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.

[0023] Some embodiments of the present application will be described in detail below with reference to the accompanying drawings. The following embodiments and features in the embodiments can be combined with each other without conflict.

[0024] Please refer to Figure 1 , Figure 1 The flow chart of the camera remote control method based on the personal WiFi provided by the embodiments of the present application is shown in FIG. 1, the camera remote control method based on the personal WiFi provided by the embodiments of the present application includes steps S100 to S400. Figure 1

[0025] Step S100, collect the position information, coverage range and signal strength of the camera node, and construct a distributed network topology structure;

[0026] It can be understood that the execution subject of the present application can be a camera remote control system based on personal WiFi, and can also be a terminal or a server, and the specific place is not limited. The embodiments of the present application take the server as the execution subject for example.

[0027] ​Specifically, each camera node is uniquely identified and coded by the Body WiFi module to ensure that all cameras in the network can be accurately distinguished, obtaining node identification data. Due to the dynamic changes in the Body WiFi network environment, based on this node identification data, the connection between the camera and the Body WiFi module is periodically detected to obtain the connection state parameters of the camera node. This parameter contains key indicators such as connection stability, signal strength, data throughput, and packet loss rate, which can effectively reflect the current communication status of the camera. According to the connection state parameters, the similarity of the camera nodes is calculated, and a dynamic grouping method is used to classify the camera nodes to form a camera node group. Similarity calculation is based on multiple dimensions, such as signal strength, location information, angle overlap, network connection stability, etc. By analyzing these factors comprehensively, camera groups with similar characteristics are divided. In the monitoring system, different cameras cover adjacent areas, or need to be cooperatively photographed due to scene requirements. Dynamic grouping helps to reduce data redundancy and improve the overall monitoring efficiency and data transmission efficiency of the system. After the grouping of the camera nodes is completed, the feature information of the fixed camera and the mobile camera is recorded respectively. For fixed cameras, record their geographic location coordinates, installation angle, and viewing angle range, etc. to accurately locate the monitoring coverage and viewing area of each camera when building the network topology. For mobile cameras, record their real-time location coordinates and additionally collect their moving trajectory data, including moving speed, direction change, path characteristics, etc. to dynamically adjust the data transmission strategy in the Body WiFi environment, ensuring that the mobile camera's picture can be stably transmitted and adapt to the changing network environment. In this way, a complete camera characteristic dataset containing the features of fixed cameras and mobile cameras is generated, covering the spatial distribution, network connection status, and motion state of the cameras. The camera characteristic dataset is transmitted to the edge processing unit deployed at the Body WiFi access point for local data aggregation processing. Remote control of cameras in the Body WiFi environment uses edge computing technology to reduce data transmission delay and improve real-time processing capability. After receiving the camera characteristic dataset, the edge processing unit performs data filtering, format conversion, and anomaly detection preprocessing operations, and performs preliminary data aggregation calculations based on the spatial distribution and connection state parameters of the cameras, generating preliminary processing results. Use the preliminary processing results and connection state parameters to build a network resource status map, considering the network connection quality, available bandwidth, and delay characteristics between camera nodes to accurately reflect the availability and transmission capacity of the current Body WiFi network environment. According to the connection state parameters of the cameras, analyze the signal strength and data transmission rate between nodes to calculate the network connection quality between cameras. By continuously monitoring the sending rate and packet loss rate of camera data packets, the available bandwidth of each node is evaluated, and a dynamic bandwidth allocation model is established according to the network load changes in different time periods.Meanwhile, by measuring the transmission delay of the camera video stream, the data transmission delay between nodes is calculated, and the network topology structure is combined to optimize the data transmission path to reduce the influence of delay on remote control. Through the above steps, the establishment of the distributed network topology structure is completed.

[0028] Step S200, according to the distributed network topology structure, scene analysis is performed on the video content collected by the camera, and the optimal encoding parameter configuration table is calculated;

[0029] Specifically, a multi-thread parallel collection mechanism is designed for each camera node. This mechanism can monitor the instantaneous uplink bandwidth, instantaneous downlink bandwidth, packet round-trip delay, and packet loss rate simultaneously, thereby obtaining raw network status data. The introduction of the multi-thread architecture ensures that the collection of different network parameters can be performed synchronously, avoiding the problem of inconsistent data caused by timing misalignment, while improving the collection efficiency and enabling the system to respond in real time to dynamic network changes in the personal WiFi environment. The raw network status data is marked with camera ID and time-stamped to form a time-series network status dataset. Each data record contains the corresponding camera node identifier, collection time information, and various network performance parameters. Due to the high volatility and instability of the personal WiFi environment, the time-series network status dataset is subjected to outlier filtering, smoothing processing, and trend extraction operations to obtain more stable preprocessed network status data. During the outlier filtering process, statistical methods are used to remove abnormal data points caused by sudden interference, and smoothing processing is used to reduce the impact of instantaneous bandwidth fluctuations, ensuring the stability of the data trend. Based on time series analysis, the trend of network status data is extracted to establish a more representative network feature model. The preprocessed network status data is input into the random forest algorithm to analyze the video quality level under different bandwidth conditions and establish a functional relationship between bandwidth and encoding parameters. The random forest algorithm plays a key role in data feature learning and mapping modeling in this process. By constructing multiple decision trees and combining the voting mechanism, it can effectively analyze the impact of bandwidth fluctuations on video encoding quality while avoiding the overfitting problem caused by a single model. Based on this calculation result, a preliminary bandwidth quality mapping is generated, which describes how camera video data should adjust encoding parameters under different network conditions to ensure optimal video quality and bandwidth utilization in the personal WiFi environment. According to the complexity index of the camera acquisition scene, the preliminary bandwidth quality mapping is classified and refined. Based on the motion characteristics of the video content, the monitoring scene is divided into static scenes, low-motion scenes, and high-motion scenes, and the corresponding bandwidth demand difference models for these scenes are constructed. Static scenes refer to monitoring areas with no obvious object movement, such as fixed facility monitoring, which uses a lower bit rate to save bandwidth. Low-motion scenes cover areas with slow-moving personnel or slight environmental changes, such as ordinary indoor monitoring, so the encoding bit rate is appropriately increased to ensure detail clarity. High-motion scenes include scenes with intense motion, such as traffic monitoring or sports scenes, which require higher instantaneous response capability for video encoding, so a higher bit rate is used and the inter-frame prediction strategy is optimized to reduce motion artifacts and improve picture smoothness. Through classification and refinement, a scene-adaptive mapping matrix is formed. By combining historical data with real-time data, the scene-adaptive mapping matrix is given a time decay weight to construct a more stable network bandwidth quality model.The introduction of time decay weights enables the system to take full advantage of long-term accumulated network performance trends while ensuring that recent data has a higher influence on the dynamic adjustment of encoding parameters, thereby achieving better adaptability in the presence of frequent bandwidth fluctuations in the personal WiFi environment. Using this network bandwidth quality model, the video content captured by the camera is analyzed for scene, and an optimal encoding parameter configuration table is calculated comprehensively to ensure efficient and stable video transmission under different network conditions, thereby optimizing the overall performance of the camera remote control.

[0030] The video content captured by the camera is subjected to pixel difference calculation, edge density statistics, and motion vector distribution analysis to accurately quantify the changing characteristics of the video frame and obtain the scene change coefficient. Pixel difference calculation is used to measure the brightness and color change between adjacent frames, thereby reflecting the dynamic change characteristics of the video content; edge density statistics calculates the distribution density of edge pixels in the picture based on edge detection algorithms to evaluate the complexity of the details in the picture; and motion vector distribution analysis identifies the main motion trend in the video scene by analyzing the motion vectors generated during inter-frame prediction in the video encoding process. By integrating these three calculation processes, the dynamic complexity of the camera monitoring area is effectively evaluated, and the scene change coefficient is determined. According to the scene change coefficient, the video scene is classified into low complexity scenes, medium complexity scenes, and high complexity scenes, and reference encoding parameter tables are established for different categories of scenes to form the scene classification result. Low complexity scenes correspond to static or less changing monitoring pictures, such as fixed backgrounds or scenes with less movement of personnel, and the encoding parameters for such scenes use lower bit rates and larger inter-frame intervals to reduce bandwidth consumption. Medium complexity scenes include monitoring areas with some motion but relatively stable change rate, and the encoding strategy needs to balance bit rate, frame rate, and encoding complexity to balance picture quality and transmission efficiency. High complexity scenes usually involve dramatic motion or rapidly changing pictures, and the encoding strategy uses higher bit rates and short inter-frame intervals, and optimizes the motion compensation mechanism to ensure the clarity and smoothness of the picture. Based on this classification result, initial encoding parameter tables are assigned to different categories of scenes to provide reasonable encoding references. The scene classification result is combined with the network bandwidth quality model to construct an encoding parameter search space. The network bandwidth quality model provides feedback information on real-time bandwidth status, and the scene classification result defines the encoding parameter range suitable for different scenes. After combining the two, the search area of the encoding parameters is limited to avoid ineffective optimization in unsuitable encoding parameter ranges. The encoding parameter search space covers multiple encoding dimensions such as bit rate, frame rate, GOP (Group of Pictures), and quantization parameter, and ensures that the search range meets the requirements of network resource availability and scene characteristics. After constructing the encoding parameter search space, based on the space, multi-objective optimization calculation is performed on the video quality index, transmission delay, and energy consumption level three objective functions to determine the optimal encoding parameter combination. The video quality index is used to measure the subjective and objective quality of the video under different encoding parameters, including peak signal-to-noise ratio, structural similarity index, and other indicators to ensure the clarity and stability of the video picture after encoding compression; the transmission delay reflects the transmission delay of video data in the Body WiFi network environment, and the optimization of this objective function aims to reduce the delay jitter of data transmission to improve the real-time performance of remote monitoring; the energy consumption level is used to measure the calculation cost and power consumption of encoding operation, especially in mobile devices or low-power scenarios, and the optimization of this objective function can effectively reduce device load and energy consumption.An optimal solution of the three objective functions is calculated under different combinations of encoding parameters by using a multi-objective optimization algorithm, and an optimal combination of encoding parameters that can achieve the best balance among video quality, transmission delay and energy consumption is finally determined. The optimal combination of encoding parameters is divided into levels to form a hierarchical set of encoding parameters, improving the adaptability of encoding and enabling it to dynamically adjust to different network environments. For example, in the case of sufficient network bandwidth, the high-quality encoding parameter layer is selected to ensure the best video clarity; while in the case of limited network bandwidth, the encoding parameter layer is downgraded to a lower quality to reduce data transmission pressure and improve smoothness. In order to ensure smooth switching between encoding parameters of different levels, a parameter smoothing control mechanism is introduced to the hierarchical set of encoding parameters. This mechanism avoids picture quality jitter or freezing caused by parameter mutation when the network fluctuates by using methods such as time-weighted average and gradual adjustment of encoding parameters. After hierarchical division and parameter smoothing control, an optimal encoding parameter configuration table is generated, which can achieve optimal video encoding control under different network conditions and video scenes, thereby ensuring the efficiency and stability of the camera remote control system based on the personal WiFi in complex environments.

[0031] Step S300, divide the video frame into different priority areas according to the optimal encoding parameter configuration table, perform differential compression processing, and form video data packets with priority labels;

[0032] Specifically, the video frames captured by the camera are input into a target detection neural network based on an improved MobileNet-SSD model to accurately identify key objects in the picture. The MobileNet-SSD model is lightweight and efficient in target detection, making it suitable for embedded devices and edge computing scenarios. The improved model further optimizes feature extraction and computation efficiency, allowing it to maintain high detection accuracy in resource-constrained on-body WiFi environments. The target detection neural network is used to identify people, vehicles, and abnormal activity areas in the picture, and generates a set of target location coordinates during the detection process to accurately mark the spatial distribution of key targets in the video frames. After target detection, pixel-level weight calculation is performed on the video frames based on the target location coordinate set to form an importance weight map. This calculation process combines the detection confidence of the target area, the spatial position distribution of the target, and the motion trajectory information of the target to determine the importance weight value of each pixel. For detected people, vehicles, and abnormal activity areas, higher weight values are assigned to ensure that these areas maintain high image quality during subsequent encoding. For background or unimportant areas, lower weight values are assigned to reduce the transmission pressure of redundant data under limited bandwidth conditions. The generation of the importance weight map enables the system to perform regional encoding optimization during video compression, more reasonably allocating network bandwidth resources. The importance weight map is used to adjust the quantization parameters in the optimal encoding parameter configuration table regionally, resulting in a regionally adaptive encoding parameter table. Regional adjustment of quantization parameters involves optimization at multiple encoding levels, such as using a lower quantization step size for high-weight areas to reduce data loss and improve picture clarity, and using a higher quantization step size for low-weight areas to further reduce data volume. Through the regionally adaptive encoding strategy, the visual quality of key target areas is ensured without significantly increasing data load, improving the effectiveness of remote monitoring. Based on the regionally adaptive encoding parameter table, spatial domain preprocessing is performed on the video data to obtain spatially optimized video data. Spatial domain preprocessing mainly involves noise reduction, edge enhancement, and background smoothing operations to optimize the detail level of the picture before encoding, making it more suitable for encoding strategies in on-body WiFi environments. Noise reduction helps reduce encoding redundancy caused by high-frequency noise, edge enhancement ensures that the outlines of key targets are clearer, and background smoothing reduces the bit rate of low-priority areas, further optimizing encoding efficiency. After spatial domain processing, dual preprocessing in the time and frequency domains is performed on the video data to generate multi-level preprocessed video data. Time domain preprocessing optimizes motion features between video frames, such as reducing inter-frame redundant data through adaptive motion compensation algorithms, and using a dynamic frame rate adjustment strategy to increase frame rate during intense motion changes and reduce frame rate during static scenes to optimize bandwidth utilization.Meanwhile, the frequency domain preprocessing analyzes the frequency components of the video data based on the discrete cosine transform or wavelet transform method, and adopts different compression strategies for different frequency components to improve the coding efficiency and perceptual quality. The joint preprocessing in the time domain and the frequency domain ensures that the video data has reached the optimal compression state before encoding. After completing the multi-level preprocessing, the data packet size and transmission unit organization method are dynamically adjusted according to the current status of the on-body WiFi network to finally form video data packets with priority labels. In this process, the network bandwidth, delay, packet loss rate and other parameters are continuously monitored, and the packet encapsulation strategy is adjusted in real time according to these state information. For example, in the case of sufficient bandwidth, larger data packets are used to reduce transmission overhead and improve data throughput, while in the case of limited bandwidth, smaller data packets are used to reduce the impact of packet loss and improve transmission stability. At the same time, the transmission scheduling strategy is dynamically adjusted for video data of different priorities, for example, data packets of high priority areas are sent first to ensure the integrity of the key target area, while data of low priority areas are delayed or degraded when the network condition is poor. Form video data packets with priority labels.

[0033] Step S400, the video data packets with priority labels are distributed to the multi-level transmission queue, and the data transmission is scheduled according to the current bandwidth status of the on-body WiFi to generate an ordered transmission stream.

[0034] Specifically, according to the importance of video data packets, timeliness and system resource occupation, all video data packets are classified and graded to establish a four-layer transmission priority queue including emergency level, high priority, medium priority and low priority. The emergency level queue is used to store critical video segments that need real-time transmission, such as abnormal activity detection, emergency alarm triggering or system high priority instructions, etc. This kind of data is highly sensitive to delay and must be transmitted with the lowest delay in network scheduling process. The high priority queue stores key frame data in important monitoring pictures, including video data of target detection area. The medium priority queue stores ordinary video stream data, such as background pictures of non-critical areas or data of low motion areas. The low priority queue is used to store secondary data for delayed transmission, such as offline stored backup data and low resolution video stream, etc. In the case of limited network bandwidth, this part of data is preferentially degraded or suspended transmission to ensure the stability of the high priority queue. Real-time bandwidth measurement is performed on the body WiFi network, and a network bandwidth prediction model is generated based on the measurement data to optimize bandwidth resource allocation. In the process of real-time bandwidth measurement, the current network throughput, packet loss rate, signal strength and interference factors are monitored, and the historical network state data is combined to predict the network bandwidth change trend in the future period of time by using time series analysis method. Based on this prediction model, the bandwidth allocation strategy of the four-layer transmission priority queue is dynamically adjusted, and the queue transmission quota table is calculated. The quota table clearly specifies the bandwidth share of different priority queues under the current network state, making the data flow allocation more reasonable. The emergency level and high priority queues can occupy a large proportion when the network bandwidth is sufficient, and can still maintain the minimum bandwidth guarantee when the bandwidth is limited. The medium priority and low priority queues dynamically adjust the transmission rate or even suspend transmission according to the bandwidth state to ensure the priority processing of core data. After obtaining the queue transmission quota table, the transmission protocol is optimized according to the round trip time of the body WiFi network to improve the efficiency and stability of data transmission. The measurement of round trip time evaluates the delay jitter of network transmission through periodic probe packets (such as TCP RTT or ICMP Ping), and adjusts the window size, congestion control strategy and data packet retransmission mechanism of the transmission protocol based on the bandwidth prediction model to minimize the packet loss rate and transmission delay. For network environments with high latency, adaptive bit rate control strategy is adopted to encode video stream hierarchically, and the encoding gear is dynamically switched according to the current network state to ensure the smoothness of the picture. In order to improve the reliability of high priority data packets, the hybrid automatic repeat request mechanism is adopted, that is, when data loss or damage is found, fast retransmission is requested first to ensure that emergency level and high priority data can be transmitted with the lowest error rate. After optimizing the transmission protocol, important data packets in the transmission control parameter set are processed for forward error correction to enhance the anti-interference ability of data transmission.The forward error correction technology adds redundant check information to the data packet, so that the receiving end can recover the complete data even if part of the data is lost. For emergency and high priority data packets, higher redundancy is used to improve data recovery capability, while for medium and low priority data packets, the redundancy is appropriately reduced to reduce bandwidth occupation. Through the forward error correction strategy, the system still maintains the transmission integrity of the key video data in the case of unstable personal WiFi signal or interference. After completing the forward error correction processing, the network anomaly monitoring is performed on the transmission data stream to ensure the stability and reliability of the entire transmission process. During the anomaly monitoring process, the key indicators such as data packet loss rate, round trip delay jitter, bandwidth utilization rate are continuously tracked, and the current network state is judged whether it is abnormal in combination with the network prediction model. When detecting the decrease of network bandwidth or the increase of packet loss rate, the data transmission strategy is automatically adjusted, such as reducing the transmission rate of medium and low priority queue, increasing the data retransmission times or adjusting the video encoding parameters, to adapt to the changes of network environment. According to the historical data analysis of the pattern of network anomaly occurrence, the future transmission scheduling strategy is optimized through machine learning method to reduce the influence of abnormal condition on system performance. Under the comprehensive action of all these optimization measures, the orderly transmission stream is generated, so that the camera remote control system can efficiently and stably transmit video data in the personal WiFi environment.

[0035] According to the metadata information in the ordered transmission stream, all received video data packets are reorganized, and different buffer sizes are set according to the different network conditions and task requirements of the cameras. For cameras with stable network and high data transmission rate, a smaller buffer is used to reduce delay and improve real-time performance. For cameras with limited network or low data transmission rate, the buffer is appropriately increased to reduce data jitter and ensure the continuity and availability of the picture. Through the buffer optimization strategy, the video streams of each camera are effectively balanced to form complete video stream data. Based on the recorded camera position information and viewing angle information in the distributed network topology, the spatial alignment and fusion of the camera pictures are performed to generate a panoramic monitoring view. Feature point matching is performed on the overlapping area, and common feature points between different camera viewing angles are identified through image key point detection algorithms such as SIFT or ORB, and the corresponding relationship of these feature points is calculated to accurately align the pictures of adjacent cameras. Based on the perspective transformation algorithm, the geometric relationship of different camera pictures is adjusted using the homography matrix, so that each picture can be seamlessly spliced into a complete panoramic view. In the body WiFi environment, due to the limited network bandwidth, the video quality of some areas may be reduced due to data compression or packet loss, so quality enhancement processing is performed on the low-quality areas in the panoramic monitoring view. The super-resolution reconstruction algorithm based on the ESRGAN network is used to restore the detail information of the low-resolution area, and the artifact removal filter is used to eliminate compression artifacts to improve the clarity of the picture. The ESRGAN network can recover more texture details in the super-resolution reconstruction process through the generative adversarial learning strategy, so that the picture quality of the low-quality area is close to the original high-resolution image, and the artifact removal filter further optimizes the edge area, making the enhanced picture more natural and smooth. Through this processing step, the overall quality of the panoramic monitoring view is effectively improved, and the monitoring picture still maintains high readability and high information density in a low-bandwidth environment. The quality-enhanced video image is subjected to target detection, multi-target tracking and behavior recognition calculation to monitor abnormal activities in the monitoring area in real time. In the target detection process, a lightweight deep learning model such as YOLO or EfficientDet is used to quickly detect pedestrians, vehicles and other key targets in the picture, and a multi-target tracking algorithm such as DeepSORT is used to track the motion trajectories of these targets. At the same time, based on the behavior recognition algorithm, the motion pattern of the target is analyzed to determine whether there is abnormal behavior. For example, through skeleton key point detection, it is identified whether the pedestrian has fallen, run, fight and other abnormal actions, or through trajectory analysis, it is detected whether the vehicle violates traffic rules, and then the target behavior analysis result is generated in real time. Based on the target behavior analysis result, a set of targeted camera control instructions is generated to optimize the shooting angle, focal length and other parameter configurations of the monitoring camera.According to the analysis result, the best camera adjustment strategy is automatically calculated, and complex control instructions are converted into a set of simplified control instruction data packets to reduce transmission cost and improve control response speed. For the area with abnormal activity, zoom-in instructions are sent to the relevant camera to enhance the clarity of the target, or the pan-tilt angle of the camera is adjusted to expand the monitoring range, so that important events can be accurately captured. After generating the control instruction data packet, a lightweight encryption communication protocol is used to encapsulate it to ensure the security and integrity of data transmission. In the body WiFi environment, the traditional encryption protocol will cause a large calculation overhead, and the lightweight encryption technology is used to reduce the encryption calculation burden while ensuring data security. In order to improve the flexibility of camera control, a progressive control strategy is adopted to decompose complex control instructions into a series of basic operation sequences, so that the camera can gradually execute the adjustment to avoid picture shaking or frame loss caused by drastic changes. Through the communication link established in the distributed network topology, the control instructions are transmitted to each camera node, so that the whole remote control process can run efficiently and stably in the body WiFi environment, thereby realizing accurate remote monitoring and intelligent camera management.

[0036] In the embodiment of the application, by constructing a distributed network topology structure and monitoring the body WiFi connection state in real time, the system can quickly perceive network changes and make corresponding adjustments, and still maintain stable video transmission and control response under network fluctuation conditions, significantly reducing the video stutter rate and control delay. The adaptive encoding parameter configuration based on the network bandwidth quality model is adopted, combined with the differentiated compression strategy of regional importance perception, so that the system can significantly reduce the overall data volume while ensuring the quality of key content. Through the multi-level transmission queue management and priority sorting mechanism, it is ensured that the control instructions can still be reliably transmitted under network limitation, and the progressive control strategy and real-time feedback mechanism are adopted to ensure the accuracy and real-time performance of remote control. The target detection technology based on deep learning is used to identify important areas in the video frame, and the differentiated encoding parameters and multi-level preprocessing operations are applied, combined with the super-resolution reconstruction algorithm and the anti-artifact filter processing network to deal with the low-quality area caused by network limitation, which significantly improves the subjective and objective quality of the monitoring video. Through the design of network exception handling strategy, adaptive segmented retransmission mechanism and forward error correction coding technology, the system can effectively deal with network abnormal conditions such as signal strength drop and packet loss rate surge, and ensure the continuous operation ability of the monitoring system in poor network environment. Through intelligent allocation of computing and processing tasks, local preprocessing is performed on the camera terminal, data aggregation is performed on the edge processing unit, and video synthesis is performed on the central processor to form a layered computing architecture, which effectively balances the computing load and energy consumption of each camera node and improves the overall operation efficiency of the system.

[0037] In a specific embodiment, the process of performing step S100 can specifically include the following steps:

[0038] Assigning a unique identification code to each camera node through the body WiFi module to obtain node identification data;

[0039] Periodically detecting the connection between the camera and the body WiFi module based on the node identification data to obtain connection state parameters;

[0040] According to the connection state parameters, the similarity of the camera nodes is calculated and dynamically grouped to form a camera node group;

[0041] For fixed cameras in the camera node group, record the geographic position coordinates and the angle of view information, and for mobile cameras, additionally record the real-time position and movement trajectory to generate a camera characteristic data set;

[0042] Transmit the camera characteristic data set to the edge processing unit deployed in the body WiFi access point to perform local data aggregation processing and produce preliminary processing results;

[0043] Using the preliminary processing results and the connection state parameters to construct a network resource status map, calculate the network connection quality, available bandwidth and delay characteristics between each camera node, and complete the establishment of the distributed network topology structure.

[0044] Specifically, each camera node is uniquely identified and coded by the Body WiFi module to ensure accurate identification and management of each camera node in the entire network environment. Each camera is assigned a unique identification code, which is generated based on the MAC address, physical device serial number, or hash calculation combined with the timestamp, ensuring the uniqueness of all devices in the network. At the same time, the basic information of the camera is recorded, such as device model, resolution, power consumption and supported encoding format, and these information is stored in the node identification database for subsequent network management and optimization calculation. After completing the camera identification code allocation, based on these node identification data, the connection between the camera and the Body WiFi module is periodically detected to obtain key connection state parameters. This process involves real-time monitoring of the camera's signal strength, throughput, data packet round-trip delay, packet loss rate and other key network indicators, and combining time series analysis methods to identify the dynamic changes in its connection quality. In order to ensure the timeliness of the detection data, the sliding window method is used to calculate the short-term average, so as to filter out the fluctuations caused by temporary interference, and the exponential weighted moving average method is used to calculate the long-term trend, so as to more accurately reflect the connection state of the camera node. After obtaining the connection state parameters of the camera, the similarity of the camera nodes is calculated, and based on the similarity, the camera node groups are formed. The core goal of similarity calculation is to identify cameras with similar characteristics and divide them into the same group in order to optimize data transmission and task collaboration in the future. Based on multiple dimensions, including geographical proximity, signal quality similarity, relative stability of data transmission rate and visual angle overlap, etc. Similarity calculation is achieved by the following formula:

[0045]

[0046] wherein, represents the similarity score between camera and camera , is the geographical distance, is the signal quality similarity, is the data transmission stability, is the visual angle overlap, The weight parameter is used to adjust the relative importance of each feature in the similarity calculation. According to the calculated similarity score, the hierarchical clustering or K-means clustering method is used to divide the cameras into different groups, so as to realize more efficient cooperation in distributed management and video data optimization. After forming the camera node group, the characteristic data of each camera is recorded. For fixed cameras, the geographic position coordinates, installation angle, focal length, coverage range and other information are recorded. For mobile cameras, the real-time position and movement trajectory are additionally recorded to support dynamic adjustment. The trajectory of the mobile camera is obtained by inertial sensor, GPS or base station signal triangulation, and combined with Kalman filtering or particle filtering algorithm for trajectory prediction to reduce data jitter. For example, assuming that a mobile camera is located on a smart patrol robot, the system analyzes its historical path data and predicts its future moving direction, so as to optimize the on-body WiFi network resource allocation and ensure that the camera always has stable network connection and data transmission capability during movement. After obtaining the camera characteristic data set, it is transmitted to the edge processing unit deployed in the on-body WiFi access point to perform local data aggregation processing and generate preliminary processing results. The edge processing unit is responsible for preprocessing and fusion of camera data to reduce the burden of cloud computing and improve data processing efficiency. In this process, the network status, geographic position information, data throughput and other information of the camera are summarized and analyzed, and a multi-level cache strategy is built based on the distributed computing framework to improve the real-time performance of data processing. For example, in the high bandwidth state, high resolution data is preferentially transmitted, while in the low bandwidth state, cache data or degradation transmission strategy is used to ensure the continuity of the monitoring picture. The edge processing unit uses machine learning method to predict the future network resource demand based on the historical data of the camera, and adjusts the data transmission plan in advance to optimize the overall system performance. The network resource status graph is constructed using the preliminary processing results and connection state parameters, and the network connection quality, available bandwidth and data transmission delay and other key characteristics between the camera nodes are calculated based on this to complete the establishment of the distributed network topology structure. The network resource status graph is a topology model used to describe the data transmission situation in the entire on-body WiFi monitoring system, in which each camera node is represented as a network node, and the communication link between the cameras is marked by bandwidth, delay and other parameters. For example, through the topology structure, it is judged whether there is a network bottleneck for a camera, and the data flow is dynamically adjusted to prevent excessive occupation of bandwidth resources. The topology structure is also used to optimize the transmission path of video data. For example, when multiple cameras need to upload data at the same time, the optimal upload order is determined through the network resource status graph, and an adaptive flow control mechanism is used to ensure that high-priority video data can be preferentially transmitted.

[0047] In a specific embodiment, the process of step S200 can specifically include the following steps:

[0048] A multi-thread parallel collection mechanism is designed for each camera node, and instantaneous uplink bandwidth, instantaneous downlink bandwidth, data packet round-trip delay and data packet loss rate are synchronously monitored to obtain original network state data;

[0049] The original network state data is marked with camera ID and encoded with time stamp to form a time sequence network state data set;

[0050] Based on the time sequence network state data set, an outlier filtering, smoothing processing and trend extraction operation are performed to obtain preprocessed network state data;

[0051] The preprocessed network state data is input into a random forest algorithm to analyze the video quality level under different bandwidth conditions, establish a functional relationship between bandwidth and encoding parameters, and generate a preliminary bandwidth quality mapping;

[0052] According to the complexity index of the camera collection scene, the preliminary bandwidth quality mapping is classified and refined, and a bandwidth demand difference model for static scenes, low motion scenes and high motion scenes is respectively constructed to form a scene adaptive mapping matrix;

[0053] The historical data and real-time data are combined, and the scene adaptive mapping matrix is given a time decay weight to generate a network bandwidth quality model;

[0054] The network bandwidth quality model is used to analyze the scene of the video content collected by the camera, and an optimal encoding parameter configuration table is calculated.

[0055] Specifically, a multi-threaded parallel collection mechanism is designed for each camera node to ensure real-time monitoring and synchronous calculation of key network parameters. In actual deployment, each camera node will simultaneously obtain the instantaneous uplink bandwidth, instantaneous downlink bandwidth, packet round-trip delay, and packet loss rate through independent threads, and upload these information to the central processing unit. The uplink and downlink bandwidths are measured by periodically sending and receiving fixed-size packets, while the round-trip delay is calculated using ICMP echo-based or TCP handshake delay-based measurement methods. The packet loss rate is obtained by calculating the proportion of lost packets within a time window. Due to the dynamic changes of the network environment, lock mechanism and asynchronous data queue are used in the multi-threaded collection process to ensure the real-time and consistency of data collection, thereby avoiding the problem of data lag caused by thread blocking. After the collection of raw network status data is completed, the data is marked with camera ID and encoded with timestamp to form a complete time-series network status data set. Camera ID marking ensures that each data point can accurately correspond to a specific camera, while timestamp encoding guarantees the time sequence integrity of the data for subsequent trend analysis. The time-series data set not only includes instantaneous network parameters, but also contains historical data, allowing the system to predict the trend of network status changes using time series analysis methods. The time-series network status data is subjected to outlier filtering, smoothing processing, and trend extraction to obtain network status data. The core goal of outlier filtering is to identify and eliminate extreme data points caused by transient interference or measurement errors, and the filtering methods include the IQR (Interquartile Range) method or the dynamic threshold method based on historical mean. Smoothing processing reduces short-term fluctuations through sliding window averaging or exponential weighted moving average methods, making the data more stable. Trend extraction uses time series decomposition techniques, such as the STL-based trend decomposition method, to identify long-term trends and periodic changes in the data, thereby improving the accuracy of subsequent prediction models. The preprocessed network status data is input into the random forest algorithm to analyze the video quality level under different bandwidth conditions and establish a functional relationship between bandwidth and encoding parameters. The random forest algorithm constructs multiple decision trees and generates the final prediction result based on the majority voting method, enabling effective learning of the impact of different network states on video quality. In this process, a large amount of historical data is used as the training set to ensure that the model can accurately predict the impact of video bit rate, frame rate, and GOP (Group of Pictures) parameters on picture quality under different network conditions. Based on this model, a preliminary bandwidth-quality mapping is constructed, and the optimal encoding parameter selection under different bandwidths is described using the following formula:

[0056]

[0057] where, represents the camera optimal encoding parameter set under the current bandwidth condition, This indicates the current available bandwidth for the camera. Indicates packet loss rate. This represents the round-trip time (RTD). This functional relationship is obtained by training a random forest model and used to dynamically adjust encoding parameters to ensure a balance between video quality and network adaptability. Different monitoring scenarios have varying degrees of complexity, therefore a single bandwidth quality mapping cannot be applied to all situations. Therefore, the initial bandwidth quality mapping is refined by classifying the scenarios captured by the cameras. Computer vision algorithms are used to analyze the video content and calculate the scene complexity index, which is weighted by indicators such as the average amplitude of motion vectors, edge density, and pixel change rate. Based on this, monitoring scenarios are divided into static scenes, low-motion scenes, and high-motion scenes, and different bandwidth requirement models are constructed for each category. For example, in static scenes, due to less image change, a higher GOP value and a lower bit rate are used to save bandwidth resources; while in high-motion scenes, the GOP value is reduced and the bit rate is increased to ensure the clarity of moving objects. A scene adaptive mapping matrix is ​​constructed to achieve encoding optimization for different scenarios. Based on this, historical and real-time data are combined, and time decay weights are assigned to the scene adaptive mapping matrix to generate the final network bandwidth quality model. The introduction of time decay weights ensures that the influence of historical data on current decisions decreases over time, thus guaranteeing that the model can adapt to the latest network environment. For example, assuming that the bandwidth of a certain camera has been continuously decreasing over the past month, the system automatically adjusts its time decay weight, making newer data have a greater impact on the model, thereby improving prediction accuracy. This process is modeled using an exponential decay function, for example:

[0058]

[0059] in, Indicates the current time The weight, The attenuation coefficient is... For time intervals. Larger ones The higher the value, the faster the weight decays, making the system more inclined to use the most up-to-date data for decision-making. Using a pre-constructed network bandwidth quality model, scene analysis is performed on the video content captured by the camera, and an optimal encoding parameter configuration table is calculated to ensure optimal control of video transmission quality under different bandwidth conditions.

[0060] In one specific embodiment, the process of performing scene analysis on the video content captured by the camera using a network bandwidth quality model and calculating the optimal encoding parameter configuration table can specifically include the following steps:

[0061] The video content collected by the camera is subjected to pixel difference calculation, edge density statistics and motion vector distribution analysis to obtain a scene change coefficient;

[0062] The video scene is classified into a low complexity scene, a medium complexity scene and a high complexity scene according to the scene change coefficient, and a reference encoding parameter table is established for the video scene to form a scene classification result;

[0063] The scene classification result is combined with a network bandwidth quality model to construct an encoding parameter search space;

[0064] Based on the encoding parameter search space, multi-objective optimization calculation is performed on three objective functions of a video quality index, a transmission delay and an energy consumption level to obtain an optimal encoding parameter combination;

[0065] The optimal encoding parameter combination is subjected to hierarchical division to form a hierarchical encoding parameter set, and a parameter smoothing control mechanism is introduced to the hierarchical encoding parameter set to generate an optimal encoding parameter configuration table.

[0066] Specifically, pixel difference calculation, edge density statistics, and motion vector distribution analysis are performed on the video content to accurately assess the scene change characteristics and calculate the scene change coefficient accordingly. Pixel difference calculation is used to measure the degree of change between video frames, and the system calculates the mean square error or structural similarity index between adjacent frames to determine the dynamic change of the picture content. Edge density statistics is based on the Sobel operator or Canny operator to extract edge information of the image and calculate the proportion of edge pixels per unit area. This indicator can reflect the complexity of the picture, such as the richness of static details such as text, texture, and buildings. Motion vector distribution analysis quantifies the motion trend of the video content by analyzing the motion vector field generated during video compression and encoding, including the direction of target motion, speed distribution, and regional motion consistency. Through these calculations, the scene change coefficient is obtained, which measures the dynamic complexity of the picture content. According to the numerical range of the scene change coefficient, the video scene is classified into low complexity scene, medium complexity scene, and high complexity scene, and according to the characteristics of different categories of scenes, a reference encoding parameter table is established to form the scene classification result. Low complexity scenes correspond to static scenes with small changes in pictures, such as indoor monitoring with fixed backgrounds or monitor screens. In this case, a larger GOP value is used to reduce inter-frame redundancy and reduce bit rate to optimize bandwidth occupancy. Medium complexity scenes contain low-speed moving targets, which require a balance between frame rate and bit rate. High complexity scenes involve a large number of fast-moving targets, so reducing GOP value, increasing bit rate, and optimizing intra-frame prediction are required to ensure the clarity of the motion area. Based on this classification, a reference encoding parameter table is established to enable different types of scenes to adopt the most suitable encoding strategy. After completing the scene classification, the scene classification result is combined with the network bandwidth quality model to construct the encoding parameter search space. The network bandwidth quality model provides information such as available bandwidth, packet loss rate, and delay of the current network environment, while the scene classification result defines the applicable encoding parameter range for different scenes. By combining the two, the search area of the encoding parameters is limited, thereby avoiding unnecessary optimization in unreasonable parameter ranges. The encoding parameter search space includes bit rate, frame rate, GOP, and quantization parameter (QP), and is dynamically adjusted according to the current network status. For example, when the bandwidth is sufficient, the search space is expanded to allow higher quality encoding, while when the bandwidth is limited, the search space is contracted to reduce unnecessary data transmission pressure, thereby improving network adaptability. Based on the encoding parameter search space, multi-objective optimization calculations are performed on the video quality index, transmission delay, and energy consumption level to determine the optimal combination of encoding parameters.The video quality index is used to measure the video clarity and subjective perception quality under different encoding parameters, including peak signal-to-noise ratio, structural similarity index, etc., while the transmission delay reflects the propagation speed of video data in the network, and the optimization goal is to reduce the data transmission time to improve the real-time performance of remote monitoring. The energy consumption level measures the resource consumption of encoding calculation, especially on embedded or mobile devices, and energy consumption optimization can effectively reduce the operating cost and heat dissipation pressure of the device. In the multi-objective optimization calculation process, a Pareto optimal solution search algorithm based on weight allocation is used to achieve the optimal balance between different objective functions, and finally determine an encoding parameter combination that can guarantee picture quality and stable operation under low delay and low energy consumption conditions. The optimization process is represented by the following formula:

[0067]

[0068] wherein, represents the optimal encoding parameter combination, is the encoding parameter has an impact on the transmission delay, is the energy consumption of encoding calculation, represents the video quality index, is a weight parameter for balancing the influence of the three. The optimization process ensures that the video encoding strategy can meet the actual needs in different application scenarios and dynamically adjusts according to environmental changes. After obtaining the optimal encoding parameter combination, hierarchical division is performed to form a hierarchical encoding parameter set, so that the encoding parameters can be flexibly adjusted under different network conditions to ensure the stability of video quality. For example, when the network condition is good, the high-quality encoding parameter layer is selected to provide the best video clarity, and when the network is limited, the encoding parameter layer with lower quality is automatically switched to reduce the data transmission pressure. In order to ensure smooth switching of encoding parameters between different levels, a parameter smoothing control mechanism is introduced, which makes the change of encoding parameters more gradual through weighted average, dynamic adjustment of quantization parameters, etc., so as to avoid picture quality jitter or lag caused by parameter mutation. For example, in the case of gradual decrease of bandwidth, the system will not immediately reduce the bit rate, but will first reduce the GOP value or adjust the quantization parameter, so that the picture quality can be transitioned in a more gentle manner, ensuring the coherence of visual experience. The optimal encoding parameter configuration table is generated, and the video encoding strategy is dynamically adapted under different bandwidth conditions.

[0069] In a specific embodiment, the process of step S300 can specifically include the following steps:

[0070] The video frames captured by the camera are input into a target detection neural network based on the improved MobileNet-SSD model to identify people, vehicles and abnormal activity areas, and generate a target position coordinate set;

[0071] Pixel-level weights are calculated for video frames based on the target location coordinate set to form an importance weight map;

[0072] Based on the importance weight graph, the quantization parameters in the optimal coding parameter configuration table are regionally adjusted to obtain the regional adaptive coding parameter table;

[0073] Based on the region adaptive coding parameter table, the video data is preprocessed in the spatial domain to obtain the spatially processed video data.

[0074] The video data processed in the spatial domain is preprocessed in both the temporal and frequency domains to generate multi-level preprocessed video data.

[0075] Based on the pre-processed video data and the current status of the portable WiFi network, the size of the data packets and the organization of the transmission units are dynamically adjusted to form video data packets with priority markings.

[0076] Specifically, video frames are input into a target detection neural network based on an improved MobileNet-SSD model to achieve accurate identification of key targets. Since cameras typically operate in resource-constrained environments, traditional deep target detection networks such as YOLO or Faster R-CNN have high computational costs, while MobileNet-SSD, due to its lightweight nature, is more suitable for low-power edge devices. To improve detection accuracy and robustness, the MobileNet-SSD model is improved, including the introduction of depthwise separable convolutions to reduce computational complexity, and the use of a feature pyramid network to enhance the detection capability for small targets. Furthermore, the size distribution of the prior boxes is optimized to better suit the detection needs of pedestrians, vehicles, and abnormal activity areas in real-world monitoring scenarios. The optimized model can efficiently identify targets in video frames with low computational cost and generate a target location coordinate set, which includes the target category, target center coordinates, target size information, and confidence score. Pixel-level weights are calculated on the video frames based on the target location coordinate set to form an importance weight map. The importance weight map is used to identify the relative importance of key and non-key regions in the image, ensuring that the image quality of key regions is prioritized in subsequent encoding processes. The calculation process is based on the following formula:

[0077]

[0078] in, Represents pixels The weight value, Representative target The importance score is calculated based on the target's confidence level, size, and category weight. For pixels distance to the target center, is a decay coefficient used to control the spatial diffusion range of the weight. Through this calculation, a smooth weight map is constructed, where the pixel weights in the key target region are higher, and the weights in the background region gradually decay. After generating the importance weight map, the quantization parameters in the optimal encoding parameter configuration table are adjusted regionally using this information to obtain a regionally adaptive encoding parameter table. In traditional video encoding, the quantization parameter is usually set globally, that is, the quantization precision of all pixel points is consistent, while under the regionally adaptive encoding strategy, the system assigns different quantization parameters to different regions based on the importance weight map. For example, in the high-weight region, the QP value is reduced to improve the encoding precision and reduce data loss, while in the low-weight region, the QP value is increased to reduce the bit rate overhead. In this process, an adaptive quantization mapping function is used to convert the weight value into a quantization parameter adjustment value, and the current network bandwidth state is combined to ensure that the overall code rate meets the network transmission requirements. After obtaining the regionally adaptive encoding parameter table, the video data is preprocessed in the spatial domain based on the table to optimize the video quality and reduce redundant data. The core goal of spatial domain preprocessing is to enhance the details of important regions while smoothing the background region to reduce the encoding burden. To achieve this goal, a multi-scale edge enhancement algorithm is used to perform high-frequency detail enhancement in key regions to improve the clarity of texture and edges, while in low-weight regions, adaptive blurring is performed to reduce the bit rate occupancy of visually insensitive regions. In addition, a block-level adaptive noise reduction algorithm is used to remove high-frequency noise in high-dynamic scenes to improve encoding efficiency. After completing the spatial domain preprocessing, the video data is preprocessed in the time domain and frequency domain to generate multi-level preprocessed video data. Time domain preprocessing mainly optimizes the motion characteristics in video sequences, using an adaptive inter-frame interpolation technology based on optical flow to insert additional intermediate frames in areas with strong motion to reduce inter-frame jumps, while reducing the frame rate in static areas to reduce data transmission costs. A time-adaptive inter-frame filter is also used to reduce motion blur and improve visual consistency on the time axis. Frequency domain preprocessing is based on discrete cosine transform or wavelet transform, which removes high-frequency noise and compresses low-importance frequency components to improve the compression efficiency of encoding. For example, in low-bandwidth mode, the strength of high-frequency filtering is increased to reduce the data volume, while in high-bandwidth mode, more high-frequency components are retained to enhance the picture details. After completing all preprocessing operations, the data packet size and transmission unit organization method are dynamically adjusted according to the current random WiFi network state to form video data packets with priority markers. Due to the large time variability of the random WiFi network, the bandwidth and latency will fluctuate sharply in a short time, so an adaptive data packet packaging strategy is used to optimize the transmission performance. In the case of sufficient bandwidth, large data packets are transmitted to reduce protocol overhead and improve throughput, while in the case of limited bandwidth or high packet loss rate, the data packets are split to reduce transmission loss and improve reliability.The multi-layer priority queue scheduling algorithm is used to allocate video data packets of different importance to different transmission priority queues, and dynamically adjust the data transmission order in combination with the current network load. For example, critical area data packets are preferentially transmitted to ensure the video quality of the target area, and background data packets are downgraded for transmission when the network is congested. Through a series of optimizations, efficient intelligent video encoding and transmission are finally achieved.

[0079] In a specific embodiment, the process of performing step S400 can specifically include the following steps:

[0080] According to the content importance, timeliness and system resource occupation of the video data packet, the video data packet with priority marking is classified and categorized, and four-layer transmission priority queues including emergency level, high priority, medium priority and low priority are established;

[0081] The real-time bandwidth of the on-body WiFi network is measured to generate a network bandwidth prediction model, and the four-layer transmission priority queue is calculated for bandwidth allocation based on the network bandwidth prediction model to obtain a queue transmission quota table;

[0082] According to the queue transmission quota table and the round-trip time delay of the on-body WiFi network, the transmission protocol is optimized to form a set of transmission control parameters;

[0083] The important data packets in the set of transmission control parameters are subjected to forward error correction processing to generate a transmission data stream, and the transmission data stream is subjected to network anomaly monitoring to produce an ordered transmission stream.

[0084] Specifically, based on the importance, timeliness, and system resource usage of video data packets, priority-tagged video data packets are classified into four levels: urgent, high-priority, medium-priority, and low-priority. Urgent-priority data packets include critical video segments related to security events, such as abnormal behavior detection, traffic accident monitoring, or video streams triggered by emergency alarms. This type of data requires minimal latency and high-quality transmission, and is therefore prioritized. High-priority data packets cover core video content from important monitoring areas, such as key scenes like facial recognition and license plate detection. While ensuring smooth transmission of urgent data, frame loss should be minimized. Medium-priority data packets are ordinary video streams or auxiliary monitoring footage, with certain real-time requirements but tolerating some latency and quality degradation. Low-priority data packets mainly contain background information, low-frame-rate recordings, or redundant data; when network resources are strained, this data is actively discarded or downgraded for transmission. After classifying the priority queues, real-time bandwidth measurements are performed on the portable Wi-Fi network, and a network bandwidth prediction model is built based on the measurement results to optimize the data stream transmission strategy. Real-time bandwidth measurement employs an adaptive probing algorithm, assessing the current network status by sending small-scale test packets and recording round-trip time, throughput, and packet loss rate. Simultaneously, combined with historical bandwidth data, time series analysis methods (such as ARIMA models or LSTM neural networks) are used to predict bandwidth change trends in the near future, thus forming a network bandwidth prediction model. Based on this prediction model, bandwidth resources are allocated more accurately, avoiding the impact of sudden bandwidth drops on high-priority packets. To rationally allocate bandwidth shares for different priority queues, a bandwidth allocation optimization formula is adopted:

[0085]

[0086] in, For queue The bandwidth allocation value, The priority weight of the queue. This indicates the current data load ratio of the queue. This represents the total available bandwidth. The total number of queues is defined by this optimization formula. The system dynamically adjusts bandwidth allocation based on the actual data flow requirements, ensuring stable bandwidth support for urgent and high-priority queues even with varying network loads, while automatically reducing allocation to low-priority queues when network resources are strained, thus guaranteeing stable transmission of core data. After calculating the queue transmission quota, the transmission protocol is optimized by considering the round-trip time (RTT) of the portable WiFi network to form a transmission control parameter set. RTT measurement employs a periodic probing mechanism, calculating the RTT value by sending timestamped data packets, and combining this with a bandwidth prediction model to adjust congestion control and rate control strategies in real time. For example, when network latency is high, the window size is appropriately reduced to decrease data loss, while the window is increased to improve throughput when bandwidth is sufficient. Simultaneously, a hybrid adaptive rate control strategy is adopted, allowing data packets from different priority queues to be transmitted at different rates. For instance, urgent data packets are sent at shorter intervals, while low-priority data packets are sent at longer intervals to reduce network resource consumption. An adaptive packet retransmission strategy based on latency jitter is implemented. When high jitter is detected, the redundancy of high-priority queue packets is increased to improve video stream continuity. While optimizing the transmission protocol, forward error correction is performed on critical packets within the transmission control parameter set to enhance data anti-interference capabilities. Forward error correction adds redundant information to packets, enabling the receiver to recover complete data even with partial data loss. A Reed-Solomon-based error correction coding strategy is adopted, adding appropriate redundant bits to urgent and high-priority packets to improve data recovery capabilities, while reducing redundant bits for low-priority packets to save bandwidth. To reduce additional computational overhead, the forward error correction ratio is dynamically adjusted based on real-time network conditions. For example, when the network packet loss rate is high, the system automatically increases the error correction ratio, while when the packet loss rate is low, the generation of redundant data is reduced to decrease additional transmission overhead. This process is described by the following optimization formula:

[0087]

[0088] in, Indicates data packet The number of error correction redundancy bits, Adjust the coefficient for priority. For data packet size, The current measured packet loss rate is the packet loss rate. Through this mechanism, the utilization efficiency of network resources is optimized while ensuring the integrity of important data. After completing the forward error correction processing, network anomaly monitoring is performed on the transmission data stream to ensure the stability and reliability of data transmission. The anomaly monitoring module continuously tracks key indicators such as packet loss rate, round-trip delay jitter, and bandwidth utilization, and combines a network bandwidth prediction model to determine whether the current network is in an abnormal state. When detecting a sudden decrease in bandwidth or a packet loss rate exceeding the threshold, automatically adjust the data stream scheduling strategy, such as reducing the transmission rate of low-priority queues, or increasing the transmission of duplicate packets in emergency-level data streams to improve data reliability. Based on historical anomaly data, analyze network failure patterns and use machine learning methods to optimize future scheduling strategies. For example, if a personal WiFi hotspot often experiences increased packet loss rates under high load conditions, the system can reduce its data stream load in advance to avoid network congestion. Through the above optimization measures, an orderly transmission stream is formed, enabling the camera remote control system to achieve efficient and stable data transmission in a personal WiFi environment.

[0089] In a specific embodiment, the personal WiFi-based camera remote control method further includes the following steps:

[0090] Reorganize the received video data packets according to the metadata information in the ordered transmission stream, set different buffer sizes for the video streams of different cameras, and form complete video streams;

[0091] Based on the camera position and viewing angle information recorded in the distributed network topology, perform feature point matching and perspective transformation on the overlapping areas in the complete video stream to generate a panoramic monitoring view;

[0092] For low-quality areas in the panoramic monitoring view due to network limitations, apply an ESRGAN network-based super-resolution reconstruction algorithm and a de-artifact filter to obtain quality-enhanced video images;

[0093] Perform target detection, multi-target tracking, and behavior recognition calculations on the quality-enhanced video images to monitor abnormal activities in real time and generate target behavior analysis results;

[0094] Generate targeted camera control instruction sets based on the target behavior analysis results and form simplified control instruction data packets;

[0095] Encapsulate the simplified control instruction data packets through a lightweight encryption communication protocol, decompose complex control instructions into basic operation sequences using a progressive control strategy, and transmit the control instructions back to each camera node through the communication links established in the distributed network topology.

[0096] Specifically, the received video data packets are reorganized according to the metadata information in the ordered transmission stream, and differentiated buffer sizes are set for video streams of different cameras to ensure complete video stream splicing and transmission stability. Each camera's video data packet carries metadata information during transmission, including camera ID, timestamp, frame number, resolution, bit rate, etc. The system sorts the data packets and reconstructs the lost data according to this information. Due to different network conditions of different cameras, there are differences in data packet arrival time and data integrity. Therefore, an adaptive buffering strategy is adopted, with smaller buffer size for high-bandwidth stable cameras to reduce latency, and larger buffer size for low-bandwidth or high-packet-loss cameras to reduce jitter and improve video stream continuity. In this way, the video streams of all cameras are synchronized and integrated to form a complete monitoring picture. Based on the camera position and angle information recorded in the distributed network topology, feature point matching and perspective transformation are performed on the overlapping areas in the video picture to generate a panoramic monitoring view. Since there is some overlap in the shooting area of multiple cameras, SIFT (Scale-Invariant Feature Transform) or ORB (Oriented FAST and Rotated BRIEF) algorithm is used to detect key points in the image during panoramic splicing, and feature matching points between adjacent camera pictures are calculated to determine the overlapping area of picture splicing. Based on the RANSAC (Random Sample Consensus) algorithm, false matching points are removed, and a homography matrix is calculated to realize perspective transformation, so that the pictures of different cameras can be seamlessly spliced into a complete panoramic view. Due to the bandwidth limitation of the body WiFi, the video data transmitted by some cameras has low resolution or compression artifacts, resulting in a decrease in the quality of some areas in the panoramic monitoring view. Therefore, for low-quality areas, an ESRGAN network-based super-resolution reconstruction algorithm is applied, combined with a de-artifact filter to improve picture quality. The ESRGAN network uses a generative adversarial learning framework to restore high-frequency details in the video and reduce blurring and artifacts caused by encoding compression. The de-artifact filter uses a non-local mean and bilateral filtering method to further optimize edge sharpness and color transition, making the picture more natural. After obtaining the quality-enhanced video image, target detection, multi-target tracking, and behavior recognition calculations are performed to monitor abnormal activities in real time and generate target behavior analysis results. The improved YOLO or EfficientDet model is used for target detection to identify pedestrians, vehicles, and abandoned objects in the picture, and the OpenPose pose estimation algorithm is used to analyze pedestrian behavior and detect abnormal situations such as falling, running, and fighting. At the same time, the multi-target tracking algorithm is used to track the motion trajectories of multiple targets based on Kalman filtering and ReID (Re-Identification) technology, so that different targets can be accurately distinguished in complex scenes.Based on the target behavior analysis results, targeted camera control instruction sets are generated, and a simplified control instruction data packet is formed to optimize the camera's view angle adjustment, zoom, and tracking functions. To reduce the bandwidth occupancy of control signaling, a hierarchical control strategy is adopted, and adjustment instructions are sent to the camera only when necessary. For example, when the system detects abnormal behavior, it does not immediately adjust all cameras, but evaluates whether the current camera's view angle can effectively capture the target. If the current camera's coverage is sufficient, only the zoom parameter is adjusted, otherwise other cameras are mobilized to track the target. The generation of control instruction sets adopts a dynamic scheduling algorithm to calculate the optimal camera switching path, reducing unnecessary angle adjustment and improving monitoring stability. After generating the control instructions, the simplified control instruction data packet is packaged through a lightweight encryption communication protocol, and a progressive control strategy is adopted to decompose complex control instructions into basic operation sequences. Through the communication links established in the distributed network topology, the control instructions are transmitted back to each camera node. The lightweight encryption communication protocol uses AES-GCM (Authenticated Encryption Mode) to ensure data security in low-bandwidth environments while avoiding excessive computational overhead. The progressive control strategy allows the camera to gradually adjust its view angle and parameters after receiving the instructions, reducing jitter caused by sudden adjustments. Through the distributed network topology, the system efficiently transmits optimized control instructions to camera nodes, enabling the entire remote monitoring system to operate efficiently in low-bandwidth environments and ensuring the readability and stability of the monitoring images in complex scenarios.

[0097] The method further comprises the following steps between the step of assigning the video data packets marked with priority to the multi-level transmission queue and the step of receiving the ordered transmission stream and performing multi-camera video synthesis: a step of adaptively managing the state of the camera nodes, comprising: performing real-time state monitoring on the camera nodes in the distributed network topology, collecting node state parameters including battery power level, chip temperature, processor load, storage space occupancy and on-body WiFi signal strength, and generating a camera node health state diagram; performing time series analysis on the node state parameters based on the camera node health state diagram, identifying abnormal patterns such as rapid power decline, abnormal temperature rise and signal dramatic fluctuation by using threshold detection and change rate calculation, and forming node risk warning data; for the low-power camera nodes in the node risk warning data, starting a hierarchical energy consumption control mechanism, dynamically adjusting the processor working frequency, WiFi transmission power and encoding calculation complexity according to the remaining power percentage, and establishing a mapping relationship table between power and function degradation; re-evaluating the priority of the data packets from the abnormal nodes in the ordered transmission stream, assigning transmission priority channels for key data of power-critical or network-very-unstable nodes, and downgrading or temporarily suspending the transmission of unnecessary data, to obtain a node state-aware transmission scheduling strategy; calculating the emergency working parameters of the abnormal nodes using the node state-aware transmission scheduling strategy, including the minimum acceptable video quality, the maximum allowed encoding delay and the shortest battery endurance time, transmitting the working parameters to the central processor through a lightweight state synchronization protocol to form a node emergency response scheme; re-planning the resource allocation of the central processor based on the node emergency response scheme, reducing the sending frequency of control instructions to abnormal nodes, simplifying the complexity of control instructions, adjusting the data request strategy, and generating a resource allocation optimization table considering node state; decomposing the resource allocation optimization table into processing instructions for various node state abnormalities, including power emergency mode instructions, network limit mode instructions and high-temperature protection mode instructions, and sending them to related nodes through independent high-reliability control channels to generate emergency working mode switching trigger signals; after the abnormal node state recovers, performing a hierarchical recovery process based on the preset recovery threshold conditions, gradually improving the node function and performance level, restoring the normal working parameters, and reporting the state changes to the central processor to update the node availability information in the distributed network topology, and completing the node state adaptive management cycle.

[0098] The method further comprises the steps of protecting the control instruction and the video data from security risks between the two steps of generating the camera control instruction and transmitting the camera control instruction back to the camera node through the distributed network topology, including: performing multi-dimensional security risk assessment on the simplified control instruction data packet, analyzing the instruction type, the influence range and the execution authority, assigning 128-bit, 192-bit and 256-bit three-level encryption strength to control instructions of different security levels, and establishing a control instruction security classification table; performing lightweight encryption processing on the control instruction based on the control instruction security classification table, generating an instruction signature using an elliptic curve cryptography algorithm, deriving an instruction key in combination with a camera node unique identification code and a time stamp, and forming an anti-tampering encrypted control instruction; adding a variable-length challenge-response authentication code to the encrypted control instruction, generating a random challenge sequence through the on-body WiFi authentication node, requiring the receiving party to return a specific response sequence, building a two-way authentication channel, and ensuring the legality of the source of the control instruction; establishing an instruction transmission anomaly detection mechanism for the two-way authentication channel, statistically analyzing the instruction transmission frequency, the instruction content similarity and the instruction complexity, detecting and marking suspicious DDoS attacks and instruction injection behaviors, and generating security threat warning information; dynamically adjusting the network defense strategy using the security threat warning information, starting a hierarchical defense mechanism when an attack behavior is detected, including temporarily blocking suspicious sources, limiting the instruction receiving frequency and forcing secondary authentication, forming an adaptive security defense configuration; implementing content integrity protection on the video data stream collected by the camera, dividing the original video frame into data blocks of a fixed size, calculating the hash value of each data block and chaining them, building a video data block chain structure, and generating an anti-tampering video verification fingerprint; transmitting the video verification fingerprint through an independent encrypted channel and separating it from the video data, recalculating the hash value of the received video data at the central processor end and comparing it with the verification fingerprint, detecting possible video tampering behaviors, and outputting the video data integrity verification result; based on the video data integrity verification result and the adaptive security defense configuration, establishing a security trust score system between the camera node and the central processor, dynamically adjusting the encryption strength and verification frequency of data transmission, generating a secure communication trust model, and completing the security protection process of the control instruction and the video data.

[0099] Please refer to Figure 2 , Figure 2 The structure schematic block diagram of the camera remote control system 200 based on the on-body WiFi provided by the embodiment of the application is shown in Figure 2 The camera remote control system 200 based on the on-body WiFi comprises:

[0100] The acquisition module 210 is configured to acquire the position information, the coverage range and the signal strength of the camera node, and construct a distributed network topology.

[0101] The scene analysis module 220 is configured to perform scene analysis on the video content collected by the camera according to the distributed network topology, and calculate an optimal encoding parameter configuration table.

[0102] The compression processing module 230 is configured to divide the video frame into different priority areas according to the optimal encoding parameter configuration table, perform differential compression processing, and form a video data packet with a priority label.

[0103] The data transmission module 240 is configured to distribute the video data packet with the priority label to a multi-level transmission queue, schedule data transmission according to the current bandwidth state of the personal WiFi, and generate an ordered transmission stream.

[0104] Through the cooperation of the above components, by constructing a distributed network topology and monitoring the personal WiFi connection state in real time, the system can quickly perceive network changes and make corresponding adjustments, and still maintain stable video transmission and control response under network fluctuations, significantly reducing the video stutter rate and control delay. The adaptive encoding parameter configuration based on the network bandwidth quality model, combined with the differential compression strategy based on the importance perception of the region, enables the system to significantly reduce the overall data volume while ensuring the quality of key content. Through multi-level transmission queue management and priority sorting mechanism, the control instruction can still be reliably transmitted under network limitations. At the same time, the progressive control strategy and real-time feedback mechanism ensure the accuracy and real-time performance of remote control. Using the target detection technology based on deep learning to identify important areas in the video frame, and applying differential encoding parameters and multi-level preprocessing operations, combined with the super-resolution reconstruction algorithm and the de-artifact filter processing network, the system significantly improves the subjective and objective quality of the monitoring video. Through the design of network exception handling strategy, adaptive segmented retransmission mechanism and forward error correction coding technology, the system can effectively deal with network anomalies such as signal strength drop and packet loss rate surge, ensuring the continuous operation of the monitoring system in poor network environment. By intelligently allocating computing and processing tasks, performing local preprocessing on the camera terminal, data aggregation on the edge processing unit, and video synthesis on the central processor, a layered computing architecture is formed, effectively balancing the computing load and energy consumption of each camera node, and improving the overall operation efficiency of the system.

[0105] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, system and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.

[0106] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the entire or part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0107] The above description and the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for remotely controlling a camera based on portable WiFi, characterized in that, include: The process involves collecting location information, coverage area, and signal strength of camera nodes to construct a distributed network topology. Specifically, this includes: assigning unique identifiers to each camera node using a portable Wi-Fi module to obtain node identification data; periodically detecting the connection between the camera and the portable Wi-Fi module based on the node identification data to obtain connection status parameters; calculating similarity and dynamically grouping camera nodes according to the connection status parameters to form camera node groups; recording the geographical coordinates and viewing angle information of fixed cameras in the camera node groups, and additionally recording the real-time location and movement trajectory of mobile cameras to generate a camera characteristic dataset; transmitting the camera characteristic dataset to an edge processing unit deployed at the portable Wi-Fi access point for local data aggregation processing to generate preliminary processing results; and using the preliminary processing results and the connection status parameters to construct a network resource status map, calculating the network connection quality, available bandwidth, and latency characteristics between each camera node to complete the establishment of the distributed network topology. Based on the distributed network topology, scene analysis is performed on the video content captured by the cameras to calculate the optimal encoding parameter configuration table. Specifically, this includes: designing a multi-threaded parallel acquisition mechanism for each camera node to synchronously monitor instantaneous uplink bandwidth, instantaneous downlink bandwidth, packet round-trip latency, and packet loss rate to obtain raw network state data; encoding the raw network state data with camera IDs and timestamps to form a time-series network state dataset; performing outlier filtering, smoothing, and trend extraction operations based on the time-series network state dataset to obtain preprocessed network state data; inputting the preprocessed network state data into a random forest algorithm to analyze the video quality level under different bandwidth conditions, establishing a functional relationship between bandwidth and encoding parameters, and generating a preliminary bandwidth quality mapping; classifying and refining the preliminary bandwidth quality mapping according to the complexity index of the camera-captured scene, and constructing bandwidth requirements for static scenes, low-motion scenes, and high-motion scenes respectively. A difference model is derived to form a scene adaptive mapping matrix. Historical data and real-time data are combined, and time decay weights are assigned to the scene adaptive mapping matrix to generate a network bandwidth quality model. Pixel difference calculation, edge density statistics, and motion vector distribution analysis are performed on the video content captured by the camera to obtain scene change coefficients. Based on the scene change coefficients, video scenes are classified into low-complexity, medium-complexity, and high-complexity scenes, and a baseline coding parameter table is established for each video scene to form a scene classification result. The scene classification result is combined with the network bandwidth quality model to construct a coding parameter search space. Based on the coding parameter search space, multi-objective optimization calculations are performed on three objective functions: video quality index, transmission delay, and energy consumption level, to obtain the optimal coding parameter combination. The optimal coding parameter combination is hierarchically divided to form a hierarchical coding parameter set, and a parameter smoothing control mechanism is introduced into the hierarchical coding parameter set to generate an optimal coding parameter configuration table. According to the optimal encoding parameter configuration table, video frames are divided into different priority regions, and differentiated compression processing is performed to form video data packets with priority markings. Specifically, this includes: inputting video frames captured by the camera into a target detection neural network based on an improved MobileNet-SSD model to identify people, vehicles, and abnormal activity areas, generating a target location coordinate set; calculating pixel-level weights for video frames based on the target location coordinate set to form an importance weight map; adjusting the quantization parameters in the optimal encoding parameter configuration table regionally according to the importance weight map to obtain a regional adaptive encoding parameter table; performing spatial domain preprocessing on the video data based on the regional adaptive encoding parameter table to obtain spatially processed video data; performing dual temporal and frequency domain preprocessing on the spatially processed video data to generate multi-level preprocessed video data; and dynamically adjusting the data packet size and transmission unit organization method according to the multi-level preprocessed video data and the current portable WiFi network status to form video data packets with priority markings. The priority-tagged video data packets are allocated to multi-level transmission queues, and data transmission is scheduled according to the current bandwidth status of the portable WiFi to generate an ordered transmission stream. Specifically, this includes: classifying the priority-tagged video data packets according to their content importance, timeliness, and system resource usage, establishing a four-level transmission priority queue containing urgent, high-priority, medium-priority, and low-priority levels; performing real-time bandwidth measurement on the portable WiFi network, generating a network bandwidth prediction model, and calculating bandwidth allocation for the four-level transmission priority queues based on the network bandwidth prediction model to obtain a queue transmission quota table; optimizing the transmission protocol based on the queue transmission quota table and the round-trip latency of the portable WiFi network to form a transmission control parameter set; performing forward error correction processing on important data packets in the transmission control parameter set to generate a transmission data stream, and monitoring network anomalies in the transmission data stream to generate an ordered transmission stream; and processing the received video data according to the metadata information in the ordered transmission stream. The data packets are reassembled, and differentiated buffer sizes are set for the video streams from different cameras to form a complete video stream. Based on the camera position and viewing angle information recorded in the distributed network topology, feature point matching and perspective transformation are performed on the overlapping areas in the complete video stream to generate a panoramic monitoring view. For low-quality areas in the panoramic monitoring view caused by network limitations, an ESRGAN-based super-resolution reconstruction algorithm and artifact removal filter are applied to obtain enhanced video images. Target detection, multi-target tracking, and behavior recognition calculations are performed on the enhanced video images to monitor abnormal activities in real time and generate target behavior analysis results. Based on the target behavior analysis results, a targeted set of camera control instructions is generated to form a simplified control instruction data packet. The simplified control instruction data packet is encapsulated using a lightweight encrypted communication protocol, and a progressive control strategy is used to decompose complex control instructions into basic operation sequences. The control instructions are then transmitted back to each camera node through the communication links established in the distributed network topology.

2. A remote control system for a camera based on portable WiFi, characterized in that, For executing the portable WiFi-based remote camera control method as described in claim 1, the portable WiFi-based remote camera control system includes: The acquisition module is used to collect the location information, coverage area, and signal strength of camera nodes to build a distributed network topology. The scene analysis module is used to perform scene analysis on the video content captured by the camera based on the distributed network topology and calculate the optimal encoding parameter configuration table. The compression processing module is used to divide video frames into different priority regions according to the optimal encoding parameter configuration table, perform differentiated compression processing, and form video data packets with priority tags. The data transmission module is used to allocate the video data packets with priority tags to a multi-level transmission queue, schedule data transmission according to the current bandwidth status of the portable WiFi, and generate an ordered transmission stream.

Citation Information

Patent Citations

  • Multi-path data packet scheduling method based on real-time video structure driving

    CN118400317A

  • Mobile terminal remote control monitoring system and method

    CN119135847A

  • Monitoring video transmission method and system based on wireless communication

    CN119421008A