Camera remote control method and system based on portable WiFi

By building a distributed network topology in the camera remote control system and real-time monitoring of portable WiFi connection status, adaptive encoding parameter configuration and differentiated compression strategy are adopted to solve the problems of video lag and control delay in traditional systems in portable WiFi network environment, and stable and efficient video transmission and control response are achieved.

CN120091222AActive Publication Date: 2025-06-03SHENZHEN NEW SAIBO TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510307176.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-16
Publication Date
2025-06-03
Estimated Expiration
2045-03-16

AI Technical Summary

Technical Problem

Traditional camera remote control systems have problems such as video lag, control delay and loss of key information in portable WiFi network environments, and lack effective priority management mechanisms.

Method used

By constructing a distributed network topology, the portable WiFi connection status is monitored in real time, and the adaptive encoding parameter configuration based on the network bandwidth quality model and the differentiated compression strategy of regional importance perception are used to allocate priority tagged video data packets to the multi-level transmission queue, and data transmission is scheduled according to the current bandwidth status.

Benefits of technology

Maintaining stable video transmission and control response under network fluctuations significantly reduces video lag rate and control delay, ensuring the quality of key content while greatly reducing the overall data volume, and ensuring reliable transmission of control instructions under network constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091222A_ABST
    Figure CN120091222A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of camera remote control, and discloses a camera remote control method and system based on portable WiFi. The method comprises the following steps: acquiring position information, coverage range and signal intensity of camera nodes, and constructing a distributed network topology structure; according to the distributed network topology structure, scene analysis is carried out on video content collected by a camera, and an optimal coding parameter configuration table is calculated; dividing the video frame into different priority regions according to the optimal coding parameter configuration table, and executing differential compression processing to form a video data packet with a priority mark; according to the method, video data packets with priority marks are distributed to a multi-level transmission queue, data transmission is scheduled according to the current bandwidth state of the portable WiFi, and an ordered transmission stream is generated, so that the computing load and energy consumption of each camera node are effectively balanced, stable video transmission and control response can still be kept under the condition of network fluctuation, and the video transmission efficiency is improved. And the video lagging rate and the control delay are obviously reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of camera remote control, and particularly to a method and system for camera remote control based on portable Wi-Fi. Background Art

[0002] With the rapid popularization of security monitoring and remote video applications, distributed multi-camera monitoring systems have been widely studied and applied. Such systems usually consist of multiple physically separated cameras connected to a central processor through a network to achieve joint video processing and all-round monitoring. However, traditional wired network connections have problems such as complex wiring, high installation costs, and poor deployment flexibility. As a portable wireless network solution, portable Wi-Fi provides new possibilities for camera remote control. However, as a mobile network environment, the network status of portable Wi-Fi is easily affected by factors such as the environment, distance, and obstacles, resulting in problems such as large bandwidth fluctuations, unstable connections, and high packet loss rates. These uncertainties pose severe challenges to high-quality video transmission and precise control.

[0003] Existing camera remote control systems mostly adopt transmission strategies with fixed parameters and cannot effectively cope with the dynamic changes of the portable Wi-Fi network environment. When the network condition deteriorates, the system either maintains high-bitrate transmission, resulting in serious video stuttering and control delays, or blindly reduces the video quality at the expense of the monitoring effect. In addition, traditional solutions often evenly distribute limited bandwidth resources to all video content, ignoring the importance differences of different video region contents, resulting in the loss of key information or the problem that non-critical regions occupy too much bandwidth resources. At the same time, the conventional camera control instruction transmission lacks a priority management mechanism. In the case of network congestion, control instructions may be delayed or lost, seriously affecting the real-time performance and reliability of remote control. Summary of the Invention

[0004] The present invention provides a method and system for camera remote control based on portable Wi-Fi. The present invention effectively balances the computing load and energy consumption of each camera node, and can still maintain stable video transmission and control response under network fluctuations, significantly reducing the video stuttering rate and control delay.

[0005] In a first aspect, the present invention provides a method for camera remote control based on portable Wi-Fi. The method for camera remote control based on portable Wi-Fi includes:

[0006] Collect the position information, coverage range, and signal strength of camera nodes to construct a distributed network topology structure;

[0007] According to the distributed network topology structure, perform scene analysis on the video content collected by the camera, and calculate an optimal coding parameter configuration table;

[0008] Divide the video frames into different priority regions according to the optimal coding parameter configuration table, perform differential compression processing, and form video data packets with priority tags;

[0009] Allocate the video data packets with priority tags to a multi-level transmission queue, schedule data transmission according to the current bandwidth status of the portable Wi-Fi, and generate an ordered transmission stream.

[0010] In a second aspect, the present invention provides a remote control system for a camera based on a portable Wi-Fi. The remote control system for a camera based on a portable Wi-Fi includes:

[0011] An acquisition module, configured to acquire the location information, coverage range, and signal strength of the camera node, and construct a distributed network topology;

[0012] A scene analysis module, configured to perform scene analysis on the video content acquired by the camera according to the distributed network topology, and calculate an optimal coding parameter configuration table;

[0013] A compression processing module, configured to divide the video frames into different priority regions according to the optimal coding parameter configuration table, perform differential compression processing, and form video data packets with priority tags;

[0014] A data transmission module, configured to allocate the video data packets with priority tags to a multi-level transmission queue, schedule data transmission according to the current bandwidth status of the portable Wi-Fi, and generate an ordered transmission stream.

[0015] In the technical solution provided by the present invention, by constructing a distributed network topology and real-time monitoring the connection status of the portable WiFi, the system can quickly perceive network changes and make corresponding adjustments, and still maintain stable video transmission and control response under network fluctuation conditions, significantly reducing the video freezing rate and control latency. By adopting an adaptive coding parameter configuration based on the network bandwidth quality model and combining a differential compression strategy with regional importance perception, the system can greatly reduce the overall data volume while ensuring the quality of key content. Through a multi-level transmission queue management and priority sorting mechanism, it ensures the reliable transmission of control instructions even under network constraints. At the same time, by adopting a progressive control strategy and a real-time feedback mechanism, it guarantees the accuracy and real-time nature of remote control. Using object detection technology based on deep learning to identify important regions in video frames, applying differential coding parameters and multi-level preprocessing operations, and combining a super-resolution reconstruction algorithm and an anti-artifact filter to process low-quality regions caused by network constraints, significantly improves the subjective and objective quality of surveillance videos. By designing a network anomaly handling strategy, an adaptive segmented retransmission mechanism, and a forward error correction coding technology, the system can effectively handle network anomalies such as a sudden drop in signal strength and a sudden increase in packet loss rate, ensuring the continuous operation ability of the surveillance system in a harsh network environment. By intelligently allocating computing and processing tasks, performing local preprocessing at the camera terminal, aggregating data at the edge processing unit, and performing video synthesis at the central processor, a hierarchical computing architecture is formed, effectively balancing the computing load and energy consumption of each camera node and improving the overall operation efficiency of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following described drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0017] Figure 1 It is a schematic flowchart of the method for remotely controlling a camera based on a portable WiFi provided by an embodiment of the present application;

[0018] Figure 2 It is a schematic block diagram of the structure of the camera remote control system based on a portable WiFi provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0020] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged, so the actual execution order may change based on the actual situation.

[0021] It should also be understood that the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0022] It should be further understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0023] The following will describe in detail some embodiments of this application with reference to the accompanying drawings. Without conflict, the features in the following embodiments and the embodiments can be combined with each other.

[0024] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for remotely controlling a camera based on a portable Wi-Fi provided by an embodiment of this application. As Figure 1 shown, the method for remotely controlling a camera based on a portable Wi-Fi provided by an embodiment of this application includes steps S100 to S400.

[0025] Step S100: Collect the location information, coverage range, and signal strength of the camera nodes, and construct a distributed network topology structure;

[0026] It can be understood that the execution subject of the present invention can be a remote control system for a camera based on a portable Wi-Fi, or a terminal or a server. Specifically, it is not limited here. An embodiment of the present invention will be described by taking the server as the execution subject as an example.

[0027] Specifically, the unique identification codes are assigned to each camera node through the portable Wi-Fi module to ensure that all cameras in the network can be accurately distinguished, and the node identification data is obtained. Due to the dynamic changes in the portable Wi-Fi network environment, based on this node identification data, the connection between the camera and the portable Wi-Fi module is periodically detected to obtain the connection status parameters of the camera node. This parameter includes key indicators such as connection stability, signal strength, data throughput, and packet loss rate, which can effectively reflect the current communication status of the camera. The similarity calculation of the camera nodes is performed according to the connection status parameters, and the dynamic grouping method is used to classify the camera nodes to form camera node groups. The similarity calculation is carried out based on multiple dimensions, such as signal strength, location information, view overlap, network connection stability, etc. By comprehensively analyzing these factors, camera groups with similar characteristics are divided. In the monitoring system, different cameras cover adjacent areas or need to cooperate for shooting due to scene requirements. Dynamic grouping helps to reduce data redundancy and improve the overall monitoring efficiency and data transmission efficiency of the system. After the camera nodes are grouped, the feature information of fixed cameras and mobile cameras is recorded respectively. For fixed cameras, key information such as their geographical location coordinates, installation angles, and view ranges are recorded to accurately locate the monitoring coverage and visible areas of each camera when constructing the network topology. For mobile cameras, their real-time location coordinates are recorded, and their movement trajectory data, including movement speed, direction change, and path characteristics, are additionally collected to dynamically adjust the data transmission strategy in the portable Wi-Fi environment to ensure the stable transmission of the mobile camera's images and adapt to the changing network environment. In this way, a complete camera feature dataset containing the features of fixed cameras and mobile cameras is generated, covering the spatial distribution, network connection status, and motion state of the cameras. The camera feature dataset is transmitted to the edge processing unit deployed at the portable Wi-Fi access point to perform local data aggregation processing. The remote control of the camera in the portable Wi-Fi environment uses edge computing technology to reduce data transmission latency and improve real-time processing capabilities. After receiving the camera feature dataset, the edge processing unit performs preprocessing operations such as data filtering, format conversion, and anomaly detection, and conducts preliminary data aggregation calculations according to the spatial distribution and connection status parameters of the cameras to generate preliminary processing results. A network resource status map is constructed using the preliminary processing results and connection status parameters, comprehensively considering the network connection quality, available bandwidth, and latency characteristics between camera nodes to accurately reflect the availability and transmission capabilities of the current portable Wi-Fi network environment. Based on the connection status parameters of the cameras, the signal strength and data transmission rate between nodes are analyzed to calculate the network connection quality between cameras. By continuously monitoring the sending rate and packet loss rate of camera data packets, the available bandwidth of each node is evaluated, and a dynamic bandwidth allocation model is established for the network load changes in different time periods.Meanwhile, by measuring the transmission delay of the camera video stream, calculating the data transmission time delay between nodes, and combining with the network topology structure, the data transmission path is optimized to reduce the impact of delay on remote control. Through the above steps, the establishment of the distributed network topology structure is completed.

[0028] Step S200: According to the distributed network topology structure, perform scene analysis on the video content collected by the camera, and calculate the optimal coding parameter configuration table;

[0029] Specifically, a multi-threaded parallel acquisition mechanism is designed for each camera node. This mechanism can simultaneously monitor the instantaneous uplink bandwidth, instantaneous downlink bandwidth, data packet round-trip delay, and data packet loss rate, so as to obtain the original network status data. The introduction of the multi-threaded architecture ensures that the acquisition of different network parameters can be carried out synchronously, avoiding the problem of data inconsistency caused by time sequence misalignment, while improving the acquisition efficiency, so that the system can respond to dynamic network changes in the portable WiFi environment in real time. The original network status data is marked with camera ID and timestamp encoded to form a time series network status data set. Each data record contains the corresponding camera node identification, acquisition time information, and various network performance parameters. Due to the large volatility and instability of the portable WiFi environment, outlier filtering, smoothing, and trend extraction are performed on the time series network status data set to obtain more stable pre-processed network status data. In the process of outlier filtering, statistical methods are used to eliminate abnormal data points caused by sudden interference, and smoothing is used to reduce the impact of violent fluctuations in instantaneous bandwidth to ensure the stability of data trends. Based on the time series analysis method, the change trend of network status data is extracted to establish a more representative network feature model. The preprocessed network status data is input into the random forest algorithm to analyze the video quality level under different bandwidth conditions and establish the functional relationship between bandwidth and encoding parameters. The random forest algorithm plays a key role in data feature learning and mapping modeling in this process. By constructing multiple decision trees and combining a voting mechanism, it can effectively analyze the impact of bandwidth fluctuations on video encoding quality while avoiding the overfitting problem caused by a single model. Based on this calculation result, a preliminary bandwidth quality map is generated, which describes how the encoding parameters of the camera video data should be adjusted under different network conditions to ensure the optimal video quality and bandwidth utilization in the portable WiFi environment. The preliminary bandwidth quality map is classified and refined according to the complexity index of the camera acquisition scene. According to the motion characteristics of the video content, the monitoring scenes are divided into static scenes, low-motion scenes, and high-motion scenes, and the bandwidth demand difference models corresponding to these scenes are constructed respectively. Static scenes refer to monitoring areas with no obvious object movement, such as fixed facility monitoring, and their encoding parameters use a lower bit rate to save bandwidth; low-motion scenes cover areas where people move slowly or the environment changes slightly, such as ordinary indoor monitoring, so the encoding bit rate is appropriately increased to ensure detail clarity; high-motion scenes include targets with violent movements, such as traffic monitoring or sports scenes. Such scenes require high instantaneous response capabilities for video encoding, so a higher bit rate is used and the inter-frame prediction strategy is optimized to reduce motion artifacts and improve picture smoothness. Through classification and refinement processing, a scene adaptive mapping matrix is ​​formed. Combine historical data with real-time data, and assign time decay weights to the scene adaptive mapping matrix to build a more stable network bandwidth quality model.The introduction of time decay weights enables the system to fully utilize the network performance trends accumulated over the long term, while ensuring that recent data has a higher impact on the dynamic adjustment of coding parameters, thus achieving better adaptability in the portable WiFi environment with frequent bandwidth fluctuations. Using this network bandwidth quality model, scene analysis is performed on the video content collected by the camera, and the optimal coding parameter configuration table is comprehensively calculated to ensure efficient and stable video transmission under different network conditions, optimizing the overall performance of camera remote control.

[0030] Calculate the pixel differences, edge density statistics, and motion vector distribution analysis for the video content captured by the camera to accurately quantify the change characteristics of the video frame and obtain the scene change coefficient. Pixel difference calculation is used to measure the brightness and color change degree between adjacent frames, thereby reflecting the dynamic change characteristics of the video content; edge density statistics is based on the edge detection algorithm to calculate the distribution density of edge pixels in the frame to evaluate the detail complexity of the frame; and motion vector distribution analysis is to identify the main motion trend in the video scene by analyzing the motion vectors generated by inter-frame prediction during the video coding process. By integrating these three calculation processes, effectively evaluate the dynamic complexity of the camera monitoring area and determine the scene change coefficient. Classify the video scene into low-complexity scenes, medium-complexity scenes, and high-complexity scenes according to the scene change coefficient, and establish a benchmark coding parameter table for different categories of scenes to form the scene classification result. Among them, the low-complexity scene corresponds to a static or less-changing monitoring frame, such as a fixed background or a scene with less moving people. The coding parameters of such a scene use a lower bit rate and a larger inter-frame interval to reduce bandwidth consumption. The medium-complexity scene includes monitoring areas with certain motion but relatively stable change rates. Such scenes need to balance the bit rate, frame rate, and coding complexity to take into account both the picture quality and transmission efficiency. The high-complexity scene usually involves scenes with violent motion or rapid changes. At this time, the coding strategy uses a higher bit rate and a short inter-frame interval, and optimizes the motion compensation mechanism to ensure the clarity and smoothness of the picture. Based on this classification result, assign an initial coding parameter table for different categories of scenes to provide a reasonable coding benchmark. Combine the scene classification result with the network bandwidth quality model to construct a coding parameter search space. The network bandwidth quality model provides feedback information on the real-time bandwidth status, while the scene classification result defines the coding parameter range applicable to different scenes. After combining the two, limit the search area of the coding parameters to avoid ineffective optimization within unsuitable coding parameter ranges. The coding parameter search space covers multiple coding dimensions such as bit rate, frame rate, GOP (group frame structure), quantization parameter, etc., and ensures that the search range meets the requirements of network resource availability and scene characteristics. After constructing the coding parameter search space, based on this space, perform multi-objective optimization calculations on three objective functions: video quality index, transmission delay, and energy consumption level, to determine the optimal coding parameter combination. Among them, the video quality index is used to measure the subjective and objective quality of the video under different coding parameters, including indicators such as peak signal-to-noise ratio, structural similarity index, etc., to ensure the clarity and stability of the video frame after coding compression; the transmission delay reflects the transmission delay of video data in the portable Wi-Fi network environment. The optimization of this objective function aims to reduce the delay jitter of data transmission to improve the real-time performance of remote monitoring; the energy consumption level is used to measure the computational cost and power consumption of coding operations, especially in mobile devices or low-power scenarios. The optimization of this objective function can effectively reduce the device load and energy consumption.Adopt a multi-objective optimization algorithm to comprehensively calculate the optimal solutions of these three objective functions under different combinations of coding parameters, and finally determine an optimal combination of coding parameters that can achieve the best balance among video quality, transmission delay, and energy consumption. Hierarchical division is performed on the optimal combination of coding parameters to form a hierarchical coding parameter set, improving the adaptability of coding so that it can dynamically adjust according to different network environments. For example, in the case of sufficient network bandwidth, select a high-quality coding parameter layer to ensure the best video clarity; while in the case of limited network bandwidth, downgrade to a lower-quality coding parameter layer to reduce the data transmission pressure and improve the fluency. To ensure smooth switching of coding parameters between different levels, a parameter smoothing control mechanism is introduced into the hierarchical coding parameter set. This mechanism avoids picture quality jitter or stuttering caused by parameter mutations during network fluctuations through methods such as time-weighted averaging and gradual adjustment of coding parameters. After hierarchical division and parameter smoothing control, an optimal coding parameter configuration table is generated, enabling it to achieve optimal video coding control under different network conditions and video scenarios, thus ensuring the efficiency and stability of the camera remote control system based on the portable Wi-Fi in complex environments.

[0031] Step S300: Divide the video frames into different priority regions according to the optimal coding parameter configuration table, perform differential compression processing, and form video data packets with priority tags;

[0032] Specifically, the video frames captured by the camera are input into the target detection neural network based on the improved MobileNet-SSD model to accurately identify the key objects in the picture. The MobileNet-SSD model is used in embedded devices and edge computing scenarios due to its lightweight architecture and efficient target detection capabilities. The improved version of the model further optimizes feature extraction and computational efficiency on this basis, so that it can still maintain high detection accuracy in a resource-constrained portable WiFi environment. The target detection neural network is used to identify people, vehicles, and abnormal activity areas in the picture, and generates a target position coordinate set during the detection process to accurately mark the spatial distribution of key targets in the video frame. After the target detection is completed, the video frame is pixel-level weighted based on the target position coordinate set to form an importance weight map. The calculation process combines the detection confidence of the target area, the spatial position distribution of the target, and the motion trajectory information of the target to determine the importance weight value of each pixel. For the detected people, vehicles, and abnormal activity areas, a higher weight value is assigned to ensure that these areas maintain a high image quality in the subsequent encoding process, while for the background or unimportant areas, a lower weight value is assigned to reduce the transmission pressure of redundant data under limited bandwidth conditions. The generation of importance weight map enables the system to perform regional coding optimization during video compression to allocate network bandwidth resources more reasonably. The quantization parameters in the optimal coding parameter configuration table are regionally adjusted according to the importance weight map to obtain the regional adaptive coding parameter table. The regional adjustment of quantization parameters involves optimization of multiple coding levels, such as using a lower quantization step size for high-weight areas to reduce data loss and improve picture clarity, and using a higher quantization step size for low-weight areas to further reduce data volume. Through the regional adaptive coding strategy, the visual quality of key target areas is ensured without significantly increasing the data load, and the effectiveness of remote monitoring is improved. The video data is preprocessed in the spatial domain based on the regional adaptive coding parameter table to obtain spatially optimized video data. Spatial domain preprocessing mainly involves operations such as noise reduction, edge enhancement, and background smoothing to optimize the detail level of the picture before encoding, so that it is more suitable for the coding strategy in the portable WiFi environment. Among them, noise reduction processing helps to reduce the coding redundancy caused by high-frequency noise, edge enhancement ensures that the outline of key targets is clearer, and background smoothing reduces the bit rate of low-priority areas, thereby further optimizing the coding efficiency. After the spatial domain processing is completed, the video data is preprocessed in both the time domain and the frequency domain to generate multi-level preprocessed video data. The time domain preprocessing optimizes the motion characteristics between video frames, such as reducing redundant data between frames through an adaptive motion compensation algorithm, and adopting a dynamic frame rate adjustment strategy to increase the frame rate when the motion changes dramatically and reduce the frame rate in static scenes to optimize bandwidth utilization.Meanwhile, frequency-domain preprocessing is based on discrete cosine transform or wavelet transform methods to analyze the frequency components of video data and adopt different compression strategies for different frequency components to improve coding efficiency and perceptual quality. The joint preprocessing in the time domain and the frequency domain ensures that the video data has reached the optimal compression state before encoding. After completing multi-level preprocessing, the packet size and the organization method of transmission units are dynamically adjusted according to the current portable Wi-Fi network status to finally form video data packets with priority tags. During this process, parameters such as network bandwidth, latency, and packet loss rate are continuously monitored, and the packet encapsulation strategy is adjusted in real time according to this status information. For example, when the bandwidth is sufficient, larger data packets are used to reduce transmission overhead and improve data throughput, while when the bandwidth is limited, smaller data packets are used to reduce the impact of packet loss and improve transmission stability. At the same time, the transmission scheduling strategy is dynamically adjusted for video data with different priorities. For example, data packets in high-priority areas are sent first to ensure the integrity of the pictures in key target areas, while data in low-priority areas are delayed or degraded when the network condition is poor. Form video data packets with priority tags.

[0033] Step S400: Allocate the video data packets with priority tags to a multi-level transmission queue, schedule data transmission according to the current bandwidth status of the portable Wi-Fi, and generate an ordered transmission stream.

[0034] Specifically, all video data packets are classified and categorized based on the content importance, timeliness, and system resource occupancy of the video data packets, and a four-layer transmission priority queue including the emergency level, high priority level, medium priority level, and low priority level is established. The emergency-level queue is used to store critical video segments that need to be transmitted in real time, such as data for abnormal activity detection, emergency alarm triggering, or system high-priority instructions. This type of data is highly sensitive to latency and must be ensured to be transmitted with the lowest latency during network scheduling; the high-priority queue stores the key frame data in important monitoring images, including the video data in the target detection area; the medium-priority queue stores ordinary video stream data, such as the background images in non-critical areas or data in low-motion areas; and the low-priority queue is used to store secondary data for delayed transmission, such as backup data for offline storage, low-resolution video streams, etc. In the case of limited network bandwidth, this part of the data is preferentially degraded or suspended to ensure the stability of the high-priority queue. Real-time bandwidth measurement is performed on the portable Wi-Fi network, and a network bandwidth prediction model is generated based on the measurement data to optimize the bandwidth resource allocation. During the real-time bandwidth measurement process, the current network throughput, packet loss rate, signal strength, and interference factors are monitored, and combined with historical network state data, a time series analysis method is used to predict the network bandwidth change trend in the next period of time. Based on this prediction model, the bandwidth allocation strategy of the four-layer transmission priority queue is dynamically adjusted, and a queue transmission quota table is calculated through calculation. This quota table clearly stipulates the bandwidth share used by different priority queues in the current network state, making the data traffic allocation more reasonable. The emergency-level and high-priority queues can occupy a larger proportion when the network bandwidth is sufficient, and still be able to maintain the lowest bandwidth guarantee when the bandwidth is limited, while the medium-priority and low-priority queues dynamically adjust the transmission rate or even suspend transmission according to the bandwidth status to ensure the priority processing of core data. After obtaining the queue transmission quota table, the transmission protocol is optimized according to the round-trip latency of the portable Wi-Fi network to improve the efficiency and stability of data transmission. The round-trip latency is measured by periodically probing data packets (such as TCP RTT or ICMP Ping) to evaluate the latency jitter of network transmission, and combined with the bandwidth prediction model to adjust the window size, congestion control strategy, and data packet retransmission mechanism of the transmission protocol to minimize the packet loss rate and transmission delay. For a network environment with high latency, an adaptive bitrate control strategy is adopted to perform hierarchical encoding on the video stream and dynamically switch the encoding level according to the current network state to ensure the smoothness of the video. To improve the reliability of high-priority data packets, a hybrid automatic repeat request mechanism is adopted, that is, when data loss or damage is detected, a fast retransmission is preferentially requested, so as to ensure that the emergency-level and high-priority data can be transmitted with the lowest error rate. After optimizing the transmission protocol, forward error correction processing is performed on the important data packets in the transmission control parameter set to enhance the anti-interference ability of data transmission.Forward error correction technology enables the receiving end to recover complete data even when some data is lost by attaching redundant check information to data packets. For emergency-level and high-priority data packets, a higher redundancy level is adopted to improve data recovery ability, while for medium-priority and low-priority data packets, the redundancy level is appropriately reduced to minimize bandwidth occupancy. Through the forward error correction strategy, the system maintains the transmission integrity of critical video data even when the portable Wi-Fi signal is unstable or there is interference. After forward error correction processing is completed, network anomaly monitoring is performed on the transmitted data stream to ensure the stability and reliability of the entire transmission process. During the anomaly monitoring process, key metrics such as packet loss rate, round-trip delay jitter, and bandwidth utilization are continuously tracked, and the current network status is judged for anomalies in combination with a network prediction model. When a decrease in network bandwidth or an increase in the packet loss rate is detected, the data transmission strategy is automatically adjusted, such as reducing the transmission rate of the medium- and low-priority queues, increasing the number of data retransmissions, or adjusting video encoding parameters, to adapt to changes in the network environment. The pattern of network anomalies is analyzed based on historical data, and the future transmission scheduling strategy is optimized through machine learning methods to reduce the impact of abnormal conditions on system performance. With the combined effect of all these optimization measures, an ordered transmission stream is generated, enabling the camera remote control system to transmit video data efficiently and stably in a portable Wi-Fi environment.

[0035] Recombine all received video data packets according to the metadata information in the ordered transmission stream, and set different buffer sizes based on the different network conditions and task requirements of the cameras. For cameras with stable networks and high data transmission rates, use smaller buffers to reduce latency and improve real-time performance. For cameras with limited networks or low data transmission rates, appropriately increase the buffer size to reduce data jitter and ensure the continuity and usability of the video. Through the buffer optimization strategy, effectively balance the video streams of each camera to form complete video stream data. Based on the camera position information and viewing angle information recorded in the distributed network topology, spatially align and fuse the images of each camera to generate a panoramic surveillance view. Perform feature point matching on the overlapping areas, identify the common feature points between different camera views through image key point detection algorithms (such as SIFT or ORB), and calculate the corresponding relationships of these feature points to accurately align the images of adjacent cameras. Based on the perspective transformation algorithm, use the homography matrix to adjust the geometric relationships of different camera images so that the images can be seamlessly stitched into a complete panoramic view. In the portable Wi-Fi environment, due to limited network bandwidth, the video quality in some areas will decline due to data compression or packet loss. Therefore, perform quality enhancement processing on the low-quality areas in the panoramic surveillance view. Adopt a super-resolution reconstruction algorithm based on the ESRGAN network to restore the detailed information in the low-resolution areas, and use an artifact removal filter to eliminate compression artifacts to improve the clarity of the image. The ESRGAN network can restore more texture details during the super-resolution reconstruction process through the generative adversarial learning strategy, making the image quality in the low-quality areas close to the original high-resolution image. At the same time, the artifact removal filter further optimizes the edge areas, making the enhanced image more natural and smooth. Through this processing step, effectively improve the overall quality of the panoramic surveillance view, so that the surveillance image still maintains high readability and high information density in a low-bandwidth environment. Perform object detection, multi-object tracking, and behavior recognition calculations on the quality-enhanced video images to monitor abnormal activities in the surveillance area in real time. During the object detection process, use lightweight deep learning models, such as YOLO or EfficientDet, to quickly detect pedestrians, vehicles, and other key objects in the image, and track the movement trajectories of these objects through multi-object tracking algorithms (such as DeepSORT). At the same time, based on the behavior recognition algorithm, analyze the action patterns of the objects to determine whether there are abnormal behaviors. For example, identify whether pedestrians have abnormal actions such as falling, running, or fighting through skeleton key point detection, or detect whether vehicles violate traffic rules through trajectory analysis, and then generate target behavior analysis results in real time. Based on the target behavior analysis results, generate a targeted camera control instruction set to optimize the shooting angle, focal length, and other parameter configurations of the surveillance camera.Automatically calculate the optimal camera adjustment strategy based on the analysis results, and convert complex control instructions into a set of streamlined control instruction data packets to reduce transmission costs and improve control response speed. For areas with abnormal activities, send zoom instructions to relevant cameras to enhance the clarity of the target, or adjust the pan-tilt angle of the cameras to expand the monitoring range, so as to ensure that important events can be accurately captured. After generating the control instruction data packets, use a lightweight encryption communication protocol to encapsulate them to ensure the security and integrity of data transmission. In the portable Wi-Fi environment, traditional encryption protocols will cause large computational overheads. Adopt lightweight encryption technology to reduce the encryption calculation burden while ensuring data security. To improve the flexibility of camera control, adopt a progressive control strategy, decompose complex control instructions into a series of basic operation sequences, so that the cameras can perform adjustments step by step to avoid picture jitter or frame loss caused by drastic changes. Through the communication links established in the distributed network topology, transmit the control instructions to each camera node, enabling the entire remote control process to operate efficiently and stably in the portable Wi-Fi environment, thus achieving precise remote monitoring and intelligent camera management.

[0036] In the embodiments of the present invention, by constructing a distributed network topology and real-time monitoring the connection status of the portable Wi-Fi, the system can quickly perceive network changes and make corresponding adjustments, and still maintain stable video transmission and control response under network fluctuation conditions, significantly reducing the video freezing rate and control latency. Adopt an adaptive coding parameter configuration based on the network bandwidth quality model, combined with a differential compression strategy that perceives the importance of regions, enabling the system to greatly reduce the overall data volume while ensuring the quality of key content. Through a multi-level transmission queue management and priority sorting mechanism, ensure that control instructions can still be reliably transmitted under network constraints. At the same time, adopt a progressive control strategy and a real-time feedback mechanism to ensure the accuracy and real-time nature of remote control. Use object detection technology based on deep learning to identify important regions in video frames, and apply differential coding parameters and multi-level preprocessing operations. Combine super-resolution reconstruction algorithms and artifact removal filters to process low-quality regions caused by network constraints, significantly improving the subjective and objective quality of the monitoring video. By designing network anomaly handling strategies, an adaptive segmented retransmission mechanism, and forward error correction coding technology, the system can effectively handle network anomalies such as sudden drops in signal strength and sudden increases in packet loss rate, ensuring the continuous operation ability of the monitoring system in a harsh network environment. By intelligently allocating computing and processing tasks, perform local preprocessing at the camera terminals, aggregate data at the edge processing units, and perform video synthesis at the central processor to form a hierarchical computing architecture, effectively balancing the computing load and energy consumption of each camera node and improving the overall operating efficiency of the system.

[0037] In a specific embodiment, the process of executing step S100 may specifically include the following steps:

[0038] Assign unique identification codes to each camera node through the portable Wi-Fi module to obtain node identification data;

[0039] Periodically detect the connection between the camera and the portable Wi-Fi module based on the node identification data to obtain connection status parameters;

[0040] Calculate the similarity and dynamically group the camera nodes according to the connection status parameters to form a camera node group;

[0041] Record the geographical location coordinates and viewing angle information for the fixed cameras in the camera node group, and additionally record the real-time position and movement trajectory for the mobile cameras to generate a camera feature data set;

[0042] Transmit the camera feature data set to the edge processing unit deployed at the portable Wi-Fi access point to perform local data aggregation processing and generate a preliminary processing result;

[0043] Use the preliminary processing result and the connection status parameters to construct a network resource status map, calculate the network connection quality, available bandwidth, and latency characteristics between each camera node, and complete the establishment of the distributed network topology.

[0044] Specifically, the unique identification code assignment is carried out for each camera node through the portable Wi-Fi module to ensure the accurate identification and management of each camera node in the entire network environment. A unique identification code is assigned to each camera, and this identification code is generated based on the MAC address, physical device serial number, or hash calculation combined with the timestamp to ensure the uniqueness of all devices in the network. At the same time, record the basic information of the camera, such as device model, resolution, power consumption, and supported encoding formats, and store this information in the node identification database for subsequent network management and optimization calculations. After completing the camera identification code assignment, based on these node identification data, the connection between the camera and the portable Wi-Fi module is periodically detected to obtain key connection status parameters. This process involves real-time monitoring of key network metrics such as the signal strength, throughput, round-trip delay of data packets, and packet loss rate of the camera, and using time series analysis methods to identify the dynamic changes in its connection quality. To ensure the timeliness of the detection data, the sliding window method is used to calculate the short-term average value to filter out the fluctuations caused by short-term interference, and the exponentially weighted moving average method is used to calculate the long-term trend to more accurately reflect the connection status of the camera node. After obtaining the connection status parameters of the camera, the similarity calculation is performed on the camera nodes, and dynamic grouping is carried out based on the similarity to form camera node groups. The core objective of the similarity calculation is to identify cameras with similar characteristics and divide them into the same group for subsequent optimization of data transmission and task collaboration. The calculation is performed based on multiple dimensions, including geographical proximity, signal quality similarity, relative stability of data transmission rate, and view overlap degree, etc. The similarity calculation is implemented through the following formula:

[0045]

[0046] Among them, represents the similarity score between camera and camera ; is the geographical distance, is the signal quality similarity, is the data transmission stability, is the view overlap degree, is a weight parameter used to adjust the relative importance of various features in similarity calculation. According to the calculated similarity scores, hierarchical clustering or K-means clustering methods are used to divide the cameras into different groups for more efficient collaboration in distributed management and video data optimization. After forming the camera node groups, the characteristic data of each camera is recorded. For fixed cameras, information such as their geographical location coordinates, installation angles, focal lengths, and coverage ranges is recorded, while for mobile cameras, their real-time positions and movement trajectories are additionally recorded to support dynamic adjustment. The movement trajectories of mobile cameras are obtained through methods such as inertial sensors, GPS, or base station signal triangulation, and trajectory prediction is performed in combination with algorithms such as Kalman filtering or particle filtering to reduce data jitter. For example, assume a mobile camera is on an intelligent patrol robot. Then the system analyzes its historical path data and predicts its future movement direction to optimize the allocation of portable Wi-Fi network resources and ensure that the camera always has stable network connection and data transmission capabilities during movement. After obtaining the camera characteristic data set, it is transmitted to the edge processing unit deployed at the portable Wi-Fi access point to perform local data aggregation processing and generate preliminary processing results. The edge processing unit is responsible for preprocessing and fusing camera data to reduce the computing burden on the cloud and improve data processing efficiency. During this process, the network status, geographical location information, data throughput, etc. of the cameras are summarized and analyzed, and a multi-level caching strategy is constructed based on the distributed computing framework to improve the real-time performance of data processing. For example, in a high-bandwidth state, high-resolution data is preferentially transmitted, while in a low-bandwidth state, cached data or a degraded transmission strategy is used to ensure the continuity of the monitoring video. The edge processing unit uses machine learning methods to predict the future network resource requirements based on the historical data of the cameras and adjusts the data transmission plan in advance to optimize the overall system performance. A network resource status map is constructed using the preliminary processing results and connection status parameters, and based on this, key characteristics such as the network connection quality, available bandwidth, and data transmission delay between each camera node are calculated to complete the establishment of the distributed network topology. The network resource status map is a topological model used to describe the data transmission situation within the entire portable Wi-Fi monitoring system. Among them, each camera node is represented as a network node, and the communication links between the cameras are labeled with parameters such as bandwidth and delay. For example, through this topological structure, it is judged whether there is a network bottleneck for a certain camera, and the data flow is dynamically adjusted to prevent excessive occupation of bandwidth resources. This topological structure is also used to optimize the transmission path of video data. For example, when multiple cameras need to upload data simultaneously, the optimal upload order is judged through the network resource status map, and an adaptive flow control mechanism is adopted to ensure that high-priority video data can be preferentially transmitted.

[0047] In a specific embodiment, the process of executing step S200 may specifically include the following steps:

[0048] Design a multi-threaded parallel acquisition mechanism for each camera node to synchronously monitor the instantaneous uplink bandwidth, instantaneous downlink bandwidth, packet round-trip delay, and packet loss rate, and obtain the original network status data;

[0049] Mark the camera ID and encode the timestamp for the original network status data to form a time-series network status data set;

[0050] Perform outlier filtering, smoothing processing, and trend extraction operations based on the time-series network status data set to obtain the preprocessed network status data;

[0051] Input the preprocessed network status data into the random forest algorithm, analyze the video quality level under different bandwidth conditions, establish the functional relationship between the bandwidth and encoding parameters, and generate a preliminary bandwidth-quality mapping;

[0052] Classify and refine the preliminary bandwidth-quality mapping according to the complexity index of the camera acquisition scenario, and respectively construct the bandwidth demand difference models for static scenarios, low-motion scenarios, and high-motion scenarios to form a scenario-adaptive mapping matrix;

[0053] Combine historical data with real-time data, assign a time decay weight to the scenario-adaptive mapping matrix to generate a network bandwidth-quality model;

[0054] Use the network bandwidth-quality model to perform scenario analysis on the video content collected by the camera and calculate the optimal encoding parameter configuration table.

[0055] Specifically, a multi-threaded parallel acquisition mechanism is designed for each camera node to ensure real-time monitoring and synchronous calculation of key network parameters. In actual deployment, each camera node uses independent threads to simultaneously obtain the instantaneous upstream bandwidth, instantaneous downstream bandwidth, packet round-trip delay, and packet loss rate, and uploads this information to the central processing unit. The upstream and downstream bandwidths are measured by periodically sending and receiving packets of a fixed size, while the round-trip delay is calculated using measurement methods based on ICMP echo or TCP handshake latency. The packet loss rate is obtained by calculating the proportion of lost packets within a time window. Due to the relatively drastic dynamic changes in the network environment, a lock mechanism and an asynchronous data queue are used during multi-threaded acquisition to ensure the real-time and consistency of data acquisition, thereby avoiding data lag problems caused by thread blocking. After the acquisition of the original network state data, camera ID marking and timestamp encoding are performed on the data to form a complete time-series network state dataset. The camera ID marking is used to ensure that each data point can be accurately corresponded to a specific camera, while the timestamp encoding guarantees the temporal integrity of the data for subsequent trend analysis. The time-series dataset not only includes instantaneous network parameters but also historical data, enabling the system to use time series analysis methods to predict the changing trends of the network state. Outlier filtering, smoothing processing, and trend extraction are performed on the time-series network state data to obtain the network state data. The core objective of outlier filtering is to identify and eliminate extreme data points caused by instantaneous interference or measurement errors, and the filtering methods include the IQR (Interquartile Range) method or the dynamic threshold method based on historical means. Smoothing processing reduces short-term fluctuations through methods such as moving window averaging or exponentially weighted moving average, making the data more stable. Trend extraction uses time series decomposition techniques, such as the STL-based trend decomposition method, to identify long-term trends and periodic changes in the data, thereby improving the accuracy of subsequent prediction models. The preprocessed network state data is input into the random forest algorithm to analyze the video quality levels under different bandwidth conditions and establish a functional relationship between the bandwidth and encoding parameters. The random forest algorithm generates the final prediction result by constructing multiple decision trees and based on the majority voting method, thereby being able to effectively learn the impact of different network states on video quality. During this process, a large amount of historical data is used as the training set to ensure that the model can accurately predict the impact of video bit rate, frame rate, and GOP (Group of Pictures) parameters on the picture quality under different network conditions. Based on this model, a preliminary bandwidth-quality mapping is constructed, and the optimal encoding parameter selection under different bandwidths is described using the following formula:

[0056]

[0057] where represents the set of optimal encoding parameters of camera under the current bandwidth condition, Represents the current available bandwidth of the camera, represents the packet loss rate, and represents the round-trip delay. This functional relationship is obtained through training by a random forest model and is used to dynamically adjust encoding parameters to ensure a balance between video quality and network adaptability. Since the complexities of different monitoring scenarios vary, a single bandwidth-quality mapping cannot be applicable to all cases. Therefore, according to the complexity index of the camera acquisition scenario, the preliminary bandwidth-quality mapping is classified and refined. Computer vision algorithms are used to perform scene analysis on the video content, and the complexity index of the scene is calculated. This index is obtained by weighted calculation of indicators such as the average magnitude of motion vectors, edge density, and pixel change rate. Based on this, the monitoring scenarios are divided into static scenarios, low-motion scenarios, and high-motion scenarios, and different bandwidth demand models are constructed for each category. For example, in a static scenario, since the picture changes little, a higher GOP value and a lower bitrate are adopted to save bandwidth resources; while in a high-motion scenario, the GOP value is reduced and the bitrate is increased to ensure the clarity of moving objects. A scene-adaptive mapping matrix is constructed to achieve encoding optimization for different scenarios. On this basis, historical data and real-time data are combined, and a time decay weight is assigned to the scene-adaptive mapping matrix to generate the final network bandwidth-quality model. The introduction of the time decay weight causes the influence of historical data on the current decision to decrease over time, thereby ensuring that the model can adapt to the latest network environment. For example, assume that the bandwidth of a certain camera has been continuously decreasing in the past month, then the system automatically adjusts its time decay weight, making the influence of newer data on the model greater, thereby improving the accuracy of prediction. This process is modeled by an exponential decay function, for example:

[0058]

[0059] where, represents the weight at the current time of, is the decay coefficient, is the time interval. A larger value means faster weight decay, thus making the system more inclined to use the latest data for decision-making. Using the constructed network bandwidth-quality model, scene analysis is performed on the video content collected by the camera, and an optimal encoding parameter configuration table is calculated to ensure optimal control of the video transmission quality under different bandwidth conditions.

[0060] In a specific embodiment, the process of performing scene analysis on the video content collected by the camera using the network bandwidth-quality model and calculating the optimal encoding parameter configuration table may specifically include the following steps:

[0061] Calculate the pixel difference, edge density statistics, and motion vector distribution analysis of the video content collected by the camera to obtain the scene change coefficient;

[0062] Classify the video scene into low-complexity scene, medium-complexity scene, and high-complexity scene according to the scene change coefficient, and establish a benchmark coding parameter table for the video scene to form the scene classification result;

[0063] Combine the scene classification result with the network bandwidth quality model to construct the coding parameter search space;

[0064] Based on the coding parameter search space, perform multi-objective optimization calculations on three objective functions of video quality index, transmission delay, and energy consumption level to obtain the optimal coding parameter combination;

[0065] Perform hierarchical division on the optimal coding parameter combination to form a hierarchical coding parameter set, and introduce a parameter smoothing control mechanism into the hierarchical coding parameter set to generate the optimal coding parameter configuration table.

[0066] Specifically, pixel difference calculation, edge density statistics, and motion vector distribution analysis are performed on the video content to accurately evaluate the change characteristics of the scene, and the scene change coefficient is calculated accordingly. Pixel difference calculation is used to measure the degree of change between video frames. The system calculates the mean square error or structural similarity index between adjacent frames to judge the dynamic change of the picture content. Edge density statistics is based on the Sobel operator or Canny operator to extract the edge information of the image and calculate the proportion of edge pixels per unit area. This indicator can reflect the complexity in the picture, such as the richness of static details like text, texture, and buildings. Motion vector distribution analysis quantifies the motion trend of the video content by analyzing the motion vector field generated during the video compression encoding process, including the motion direction of the target, speed distribution, and regional motion consistency. Through these calculations, the scene change coefficient is comprehensively obtained to measure the dynamic complexity of the picture content. According to the numerical range of the scene change coefficient, the video scene is classified into low-complexity scenes, medium-complexity scenes, and high-complexity scenes, and a benchmark encoding parameter table is established according to the characteristics of different types of scenes to form the scene classification result. Low-complexity scenes correspond to static scenes with little change in the picture, such as indoor surveillance or monitor screens with a fixed background. In this case, a larger GOP value is used to reduce inter-frame redundancy, while reducing the bit rate and optimizing bandwidth occupancy. Medium-complexity scenes contain slowly moving targets. Such scenes have certain requirements for frame rate and bit rate, so a balance needs to be achieved between picture quality and bandwidth. High-complexity scenes involve a large number of fast-moving targets. At this time, the GOP value is reduced, the bit rate is increased, and intra-frame prediction is optimized to ensure the clarity of the motion area. Based on this classification, a benchmark encoding parameter table is established so that different types of scenes can adopt the most suitable encoding strategy. After completing the scene classification, the scene classification result is combined with the network bandwidth quality model to construct a search space for encoding parameters. The network bandwidth quality model provides information such as the available bandwidth, packet loss rate, and latency of the current network environment, while the scene classification result defines the applicable encoding parameter range for different scenes. Combining the two limits the search area of encoding parameters, thus avoiding ineffective optimization within unreasonable parameter ranges. The search space for encoding parameters includes multiple encoding dimensions such as bit rate, frame rate, GOP, and quantization parameter (QP), and is dynamically adjusted according to the current network conditions. For example, when the bandwidth is sufficient, the search space is expanded to allow higher-quality encoding, while when the bandwidth is limited, the search space is contracted to reduce unnecessary data transmission pressure, thereby improving network adaptability. Based on the search space for encoding parameters, multi-objective optimization calculations are performed on three objective functions: video quality index, transmission delay, and energy consumption level to determine the optimal combination of encoding parameters.The video quality index is used to measure the video clarity and subjective perceived quality under different encoding parameters, including indicators such as peak signal-to-noise ratio and structural similarity index. The transmission delay reflects the propagation speed of video data in the network, and its optimization goal is to reduce the data transmission time to improve the real-time performance of remote monitoring. The energy consumption level measures the resource consumption of encoding calculations, especially on embedded or mobile devices. Energy consumption optimization can effectively reduce the operating cost and heat dissipation pressure of the device. During the multi-objective optimization calculation process, a Pareto optimal solution search algorithm based on weight allocation is adopted to achieve an optimal balance between different objective functions, and finally, a combination of encoding parameters that can ensure both picture quality and stable operation under low latency and low energy consumption conditions is determined. This optimization process is represented by the following formula:

[0067]

[0068] Among them, represents the optimal combination of encoding parameters, is the influence of the encoding parameter on the transmission delay, is the energy consumption of the encoding calculation, represents the video quality index, is the weight parameter used to balance the influence of the three. This optimization process ensures that in different application scenarios, the video encoding strategy can meet the actual needs and be dynamically adjusted according to environmental changes. After obtaining the optimal combination of encoding parameters, hierarchical division is performed to form a hierarchical encoding parameter set, enabling the encoding parameters to be flexibly adjusted under different network states to ensure the stability of video quality. For example, when the network condition is good, a high-quality encoding parameter layer is selected to provide the best video clarity, while when the network is limited, it automatically switches to a lower-quality encoding parameter layer to reduce the data transmission pressure. To ensure smooth switching of encoding parameters between different levels, a parameter smoothing control mechanism is introduced. This mechanism makes the change of encoding parameters more gradual through methods such as weighted average and dynamic adjustment of quantization parameters, thus avoiding picture quality jitter or freezing caused by parameter mutations. For example, when the bandwidth gradually decreases, the system does not immediately reduce the bit rate but first reduces the GOP value or adjusts the quantization parameter, enabling the picture quality to transition in a more gentle manner and ensuring the coherence of the visual experience. Generate an optimal encoding parameter configuration table and dynamically adapt the video encoding strategy under different bandwidth conditions.

[0069] In a specific embodiment, the process of executing step S300 may specifically include the following steps:

[0070] Input the video frames collected by the camera into the object detection neural network based on the improved MobileNet-SSD model to identify people, vehicles, and abnormal activity areas, and generate a set of target position coordinates;

[0071] Perform pixel-level weight calculation on the video frames based on the target position coordinate set to form an importance weight map;

[0072] Perform regional adjustment on the quantization parameters in the optimal coding parameter configuration table according to the importance weight map to obtain a region-adaptive coding parameter table;

[0073] Perform spatial domain preprocessing on the video data based on the region-adaptive coding parameter table to obtain the video data after spatial domain processing;

[0074] Perform dual preprocessing in the time domain and frequency domain on the video data after spatial domain processing to generate the video data after multi-level preprocessing;

[0075] Dynamically adjust the packet size and transmission unit organization method according to the video data after multi-level preprocessing and the current portable Wi-Fi network status to form video data packets with priority tags.

[0076] Specifically, input the video frames into the target detection neural network based on the improved MobileNet-SSD model to achieve accurate recognition of key targets. Since cameras usually operate in resource-constrained environments, traditional deep target detection networks such as YOLO or Faster R-CNN have relatively high computational overhead, while MobileNet-SSD is more suitable for low-power edge devices due to its lightweight characteristics. To improve the detection accuracy and robustness, the MobileNet-SSD model is improved, including introducing depthwise separable convolutions to reduce computational complexity, and at the same time using a feature pyramid network to enhance the detection ability for small targets. And by optimizing the size distribution of the prior boxes, it is made more adaptable to the detection requirements of pedestrians, vehicles, and abnormal activity areas in actual monitoring scenarios. The optimized model can efficiently identify the targets in the video frames at low computational cost and generate a target position coordinate set, which includes target category, target center coordinates, target size information, and confidence scores. Perform pixel-level weight calculation on the video frames based on the target position coordinate set to form an importance weight map. The importance weight map is used to identify the relative importance of key and non-key regions in the picture, ensuring that the image quality of key regions can be preferentially guaranteed in the subsequent encoding process. The calculation process is based on the following formula:

[0077]

[0078] where, represents the weight value of pixel point ; represents the importance score of target , and this score is calculated based on the confidence, size, and category weight of the target, is pixel point Distance to the target center, is an attenuation coefficient used to control the spatial diffusion range of the weight. Through this calculation method, a smooth weight map is constructed, in which the pixel weights of the key target area are higher, while the weights of the background area gradually decay. After the importance weight map is generated, the information is used to regionally adjust the quantization parameters in the optimal coding parameter configuration table to obtain a regional adaptive coding parameter table. In the traditional video encoding process, the quantization parameters are usually set globally, that is, the quantization accuracy of all pixels is consistent, while under the regional adaptive coding strategy, the system assigns different quantization parameters to different regions based on the importance weight map. For example, in high-weight areas, the QP value is reduced to improve coding accuracy and reduce data loss, while in low-weight areas, the QP value is increased to reduce bit rate overhead. In this process, an adaptive quantization mapping function is used to convert the weight value into a quantization parameter adjustment value, and combined with the current network bandwidth status to ensure that the overall bit rate meets the network transmission requirements. After obtaining the regional adaptive coding parameter table, the video data is preprocessed in the spatial domain based on the table to optimize the video quality and reduce redundant data. The core goal of spatial domain preprocessing is to enhance the details of important areas while smoothing the background areas to reduce the coding burden. To achieve this goal, a multi-scale edge enhancement algorithm is used to perform high-frequency detail enhancement in key areas to improve the clarity of textures and edges, while in low-weight areas, adaptive blur processing is performed to reduce the bit rate occupation of visually insensitive areas. In addition, a block-level adaptive noise reduction algorithm is used to remove high-frequency noise in high-dynamic scenes to improve coding efficiency. After completing spatial domain preprocessing, the video data is double-preprocessed in the time domain and frequency domain to generate multi-level preprocessed video data. Temporal domain preprocessing mainly optimizes the motion characteristics in the video sequence. Adaptive inter-frame interpolation technology based on optical flow is used to insert additional intermediate frames in areas with intense motion to reduce inter-frame jumps, while reducing the frame rate in static areas to reduce data transmission costs. A temporal adaptive inter-frame filter is used to reduce motion blur and improve visual consistency on the timeline. Frequency domain preprocessing is based on discrete cosine transform or wavelet transform, which improves the compression efficiency of encoding by removing high-frequency noise and compressing low-importance frequency components. For example, in low-bandwidth mode, the intensity of high-frequency filtering is increased to reduce the amount of data, while in high-bandwidth mode, more high-frequency components are retained to enhance picture details. After completing all preprocessing operations, the packet size and the organization of transmission units are dynamically adjusted according to the current portable WiFi network status to form a video packet with priority marking. Since the portable WiFi network has a large time-variability, the bandwidth and latency will fluctuate dramatically in a short period of time. An adaptive packet encapsulation strategy is adopted to optimize transmission performance. When the bandwidth is sufficient, large data packets are used for transmission to reduce protocol overhead and improve throughput. When the bandwidth is limited or the packet loss rate is high, the data packets are split to reduce transmission loss and improve reliability.Using a multi-level priority queue scheduling algorithm, video data packets of different importance levels are assigned to different transmission priority queues, and in combination with the current network load situation, the data transmission order is dynamically adjusted. For example, packets in key areas are sent first to ensure the video quality of the target area, while background packets are downgraded for transmission during network congestion. Through this series of optimizations, efficient intelligent video coding and transmission are ultimately achieved.

[0079] In a specific embodiment, the process of executing step S400 may specifically include the following steps:

[0080] Classify the video data packets with priority tags according to the content importance, timeliness, and system resource occupancy of the video data packets, and establish a four-layer transmission priority queue including the emergency level, high priority level, medium priority level, and low priority level;

[0081] Perform real-time bandwidth measurement on the portable Wi-Fi network to generate a network bandwidth prediction model, and calculate the bandwidth allocation for the four-layer transmission priority queue based on the network bandwidth prediction model to obtain a queue transmission quota table;

[0082] Optimize the transmission protocol according to the queue transmission quota table and the round-trip delay of the portable Wi-Fi network to form a set of transmission control parameters;

[0083] Perform forward error correction processing on the important data packets in the set of transmission control parameters to generate a transmission data stream, and monitor the network anomalies of the transmission data stream to generate an ordered transmission stream.

[0084] Specifically, according to the content importance, timeliness, and system resource occupancy of video data packets, the video data packets with priority tags are classified hierarchically to establish a four - layer transmission priority queue including emergency level, high priority level, medium priority level, and low priority level. Emergency - level data packets include key video segments related to security events, such as video streams triggered by abnormal behavior detection, traffic accident monitoring, or emergency alerts. This type of data requires the lowest latency and ensures high - quality transmission, so it is scheduled first. High - priority data packets cover the core video content of important monitoring areas, such as key images for face recognition, license plate detection, etc. On the premise of ensuring the smooth transmission of emergency data, frame loss should be minimized as much as possible. Medium - priority data packets are ordinary video streams or auxiliary monitoring images, with certain real - time requirements but tolerating a certain degree of delay and quality degradation. Low - priority data packets mainly contain background information, low - frame - rate records, or redundant data. When network resources are scarce, these data are actively discarded or transmitted at a reduced level. After completing the classification of the priority queue, real - time bandwidth measurement is performed on the portable WiFi network, and a network bandwidth prediction model is constructed based on the measurement results to optimize the transmission strategy of data streams. Real - time bandwidth measurement uses an adaptive detection algorithm to evaluate the current network status by sending small - scale test data packets and recording the round - trip time, throughput, and packet loss rate. At the same time, combined with historical bandwidth data, time - series analysis methods (such as ARIMA model or LSTM neural network) are used to predict the bandwidth change trend in the short term in the future, thus forming a network bandwidth prediction model. Based on this prediction model, bandwidth resources are allocated more precisely to avoid the impact of sudden bandwidth drops on high - priority data packets. To reasonably allocate the bandwidth shares of different priority queues, the bandwidth allocation optimization formula is adopted:

[0085]

[0086] Among them, is the bandwidth allocation value for queue , is the priority weight of the queue, represents the current data load ratio of this queue, is the current total available bandwidth, is the total number of queues. Through this optimized formula, the system dynamically adjusts the bandwidth allocation according to the actual requirements of the data stream, enabling the emergency-level and high-priority queues to still obtain stable bandwidth support when the network load changes, while the low-priority queue automatically reduces the allocation when network resources are scarce, ensuring the stable transmission of core data. After calculating the queue transmission quota, combined with the round-trip delay of the portable WiFi network, the transmission protocol is optimized to form a set of transmission control parameters. The measurement of the round-trip delay adopts a periodic detection mechanism, calculates the RTT value by sending timestamp packets, and combines with the bandwidth prediction model to adjust the congestion control and rate control strategies in real time. For example, in the case of high network delay, the window size is appropriately reduced to reduce the data loss rate, while when the bandwidth is relatively sufficient, the window is increased to improve the throughput. At the same time, a hybrid adaptive rate control strategy is adopted, enabling the data packets of different priority queues to use different transmission rates. For example, the emergency-level data packets are sent at shorter intervals, while the low-priority data packets use longer intervals to reduce the occupancy of network resources. And an adaptive data packet retransmission strategy is calculated based on the delay jitter. When a high jitter situation is detected, the redundancy of the data packets in the high-priority queue is increased to improve the continuity of the video stream. While optimizing the transmission protocol, forward error correction is performed on the important data packets in the set of transmission control parameters to enhance the anti-interference ability of the data. The forward error correction technology adds redundant information to the data packets, enabling the receiving end to still recover the complete data even in the case of partial data loss. An error correction coding strategy based on Reed-Solomon codes is adopted, adding appropriate redundant bits to the emergency-level and high-priority data packets to improve the data recovery ability, while for the low-priority data packets, the redundant bits are reduced to save bandwidth. To reduce the additional computational overhead, the forward error correction ratio is dynamically adjusted based on the real-time network conditions. For example, when the network packet loss rate is high, the system will automatically increase the error correction ratio, while when the packet loss rate is low, the generation of redundant data is reduced to reduce the additional transmission overhead. This process is described by the following optimized formula:

[0087]

[0088] where represents the number of error correction redundant bits of the data packet ; is the priority adjustment coefficient is the data packet size is the currently measured packet loss rate. Through this mechanism, while ensuring the integrity of important data, the utilization efficiency of network resources is optimized. After forward error correction processing is completed, network anomaly monitoring is performed on the transmitted data stream to ensure the stability and reliability of data transmission. The anomaly monitoring module continuously tracks key metrics such as packet loss rate, round-trip delay jitter, and bandwidth utilization, and combines with a network bandwidth prediction model to determine whether the current network is in an abnormal state. When a sudden drop in bandwidth or a packet loss rate exceeding the threshold is detected, the data stream scheduling strategy is automatically adjusted, such as reducing the transmission rate of the medium and low priority queues, or increasing the transmission of duplicate packets in the emergency-level data stream to improve data reliability. And based on historical anomaly data analysis, the network failure mode is analyzed, and machine learning methods are used to optimize future scheduling strategies. For example, if a certain portable Wi-Fi hotspot often has an increased packet loss rate under high load, the system can reduce its data stream load in advance to avoid network congestion. Through the above optimization measures, an ordered transmission stream is formed, enabling the camera remote control system to achieve efficient and stable data transmission in a portable Wi-Fi environment.

[0089] In a specific embodiment, the method for remotely controlling a camera based on a portable Wi-Fi further includes the following steps:

[0090] Recombine the received video data packets according to the metadata information in the ordered transmission stream, set different buffer sizes for the video streams of different cameras to form a complete video stream;

[0091] Based on the camera positions and perspective information recorded in the distributed network topology, perform feature point matching and perspective transformation on the overlapping regions in the complete video stream to generate a panoramic surveillance view;

[0092] For the low-quality regions in the panoramic surveillance view caused by network limitations, apply a super-resolution reconstruction algorithm based on the ESRGAN network and an artifact removal filter to obtain a video image with enhanced quality;

[0093] Perform object detection, multi-object tracking, and behavior recognition calculations on the video image with enhanced quality, monitor abnormal activities in real time, and generate target behavior analysis results;

[0094] Generate a targeted camera control instruction set according to the target behavior analysis results to form a streamlined control instruction data packet;

[0095] Encapsulate the streamlined control instruction data packet through a lightweight encryption communication protocol, decompose complex control instructions into basic operation sequences using a progressive control strategy, and transmit the control instructions back to each camera node through the communication links established in the distributed network topology.

[0096] Specifically, the received video data packets are recombined according to the metadata information in the ordered transmission stream, and different buffer sizes are set for the video streams of different cameras to ensure complete video stream stitching and transmission stability. Each camera's video data packets carry metadata information during transmission, including camera ID, timestamp, frame sequence number, resolution, bit rate, etc. The system sorts the data packets and reconstructs the lost data based on this information. Since different cameras are in different network conditions, there are differences in the arrival time and data integrity of their data packets. Therefore, an adaptive buffering strategy is adopted. A smaller buffer is set for cameras with high-bandwidth and stable connections to reduce latency, while for cameras with low bandwidth or high packet loss rates, a larger buffer is used to reduce jitter and improve the continuity of the video stream. In this way, the video streams of all cameras are synchronously integrated to form a complete surveillance picture. Based on the camera position and viewing angle information recorded in the distributed network topology, feature point matching and perspective transformation are performed on the overlapping areas in the video frames to generate a panoramic surveillance view. Since there is a certain overlap in the shooting areas of multiple cameras, during panoramic stitching, the SIFT (Scale-Invariant Feature Transform) or ORB (Oriented FAST and Rotated BRIEF) algorithm is used to detect the key points in the images, and the feature matching points between adjacent camera frames are calculated to determine the overlapping areas for frame stitching. Based on the RANSAC (Random Sample Consensus) algorithm, the incorrect matching points are removed, and the homography matrix is calculated to achieve perspective transformation, enabling the frames of different cameras to be seamlessly stitched into a complete panoramic view. Due to the bandwidth limitation of the portable Wi-Fi, the video data transmitted by some cameras has a relatively low resolution or compression artifacts, resulting in a decrease in the picture quality of some areas in the panoramic surveillance view. Therefore, for low-quality areas, a super-resolution reconstruction algorithm based on the ESRGAN network is applied, combined with an artifact removal filter, to improve the picture quality. The ESRGAN network can restore the high-frequency details in the video and reduce the blur and artifacts caused by encoding compression through a generative adversarial learning framework. At the same time, the artifact removal filter uses methods based on non-local means and bilateral filtering to further optimize the edge sharpness and color transition, making the picture more natural. After obtaining the video images with enhanced quality, object detection, multi-object tracking, and behavior recognition calculations are performed to monitor abnormal activities in real time and generate target behavior analysis results. Object detection uses improved YOLO or EfficientDet models to identify pedestrians, vehicles, left-behind objects, etc. in the frames, and deep learning-based pose estimation algorithms (such as OpenPose) are used to analyze pedestrian behavior and detect abnormal situations such as falls, running, and fighting. At the same time, using a multi-object tracking algorithm based on Kalman filtering and ReID (Re-Identification) technology, the movement trajectories of multiple targets are tracked, enabling accurate differentiation of different targets in complex scenarios.Based on the analysis results of the target behavior, generate a targeted camera control instruction set and form a streamlined control instruction data packet to optimize the camera's perspective adjustment, zoom, and tracking functions. To reduce the bandwidth occupancy of control signaling, a hierarchical control strategy is adopted, and adjustment instructions are sent to the camera only when necessary. For example, when the system detects an abnormal behavior, it does not immediately adjust all cameras. Instead, it evaluates whether the current camera's perspective can effectively capture the target. If the coverage range of the current camera is sufficient, only the zoom parameter is adjusted. Otherwise, other cameras are mobilized to track the target. The generation of the control instruction set uses a dynamic scheduling algorithm to calculate the optimal camera switching path to reduce unnecessary angle adjustments and improve the stability of monitoring. After generating the control instructions, the streamlined control instruction data packet is encapsulated through a lightweight encryption communication protocol, and a progressive control strategy is adopted to decompose complex control instructions into basic operation sequences. Through the communication links established in the distributed network topology, the control instructions are sent back to each camera node. The lightweight encryption communication protocol uses AES-GCM (Authenticated Encryption Mode) to ensure the security of data in a low-bandwidth environment while avoiding excessive computational overhead. The progressive control strategy allows the camera to gradually adjust its perspective and parameters after receiving the instructions to reduce jitter caused by sudden adjustments. Through the distributed network topology, the system efficiently transmits the optimized control instructions to the camera nodes, enabling the entire remote monitoring system to operate efficiently in a low-bandwidth environment and ensuring the readability and stability of the monitoring screen in complex scenarios.

[0097] Among them, between the steps of allocating video data packets with priority tags to a multi-level transmission queue, receiving an ordered transmission stream, and performing multi-camera video synthesis, there is also a step of adaptively managing the status of camera nodes, including: performing real-time status monitoring on camera nodes in a distributed network topology, collecting node status parameters including battery power level, chip temperature, processor load, storage space occupancy rate, and portable WiFi signal strength, and generating a health status map of camera nodes; performing time series analysis on the node status parameters based on the health status map of camera nodes, using threshold detection and rate-of-change calculation to identify abnormal patterns such as rapid battery power decline, abnormal temperature rise, and severe signal fluctuations, and forming node risk warning data; for low-battery camera nodes in the node risk warning data, starting a hierarchical energy consumption control mechanism, dynamically adjusting the processor operating frequency, WiFi transmission power, and encoding calculation complexity according to the remaining battery percentage, and establishing a mapping relationship table between battery power and function degradation; re-evaluating the priorities of data packets from abnormal nodes in the ordered transmission stream, allocating transmission priority channels for critical data of nodes with critical battery power or extremely unstable networks, while degrading or delaying the transmission of non-essential data to obtain a transmission scheduling strategy for node status perception; calculating emergency working parameters of abnormal nodes using the transmission scheduling strategy for node status perception, including the minimum acceptable video quality, the maximum allowable encoding delay, and the shortest battery life, and transmitting the working parameters to the central processor through a lightweight status synchronization protocol to form a node emergency response plan; re-planning the resource allocation of the central processor based on the node emergency response plan, reducing the sending frequency of control instructions to abnormal nodes, simplifying the complexity of control instructions, and adjusting the data request strategy to generate an optimized resource allocation table considering node status; decomposing the optimized resource allocation table into processing instructions for various types of node status abnormalities, including power emergency mode instructions, network limit mode instructions, and high-temperature protection mode instructions, and sending them to relevant nodes through an independent highly reliable control channel to generate an emergency working mode switching trigger signal; after the status of abnormal nodes is restored, performing a hierarchical recovery process based on preset recovery threshold conditions, gradually improving the node function and performance level, restoring normal working parameters, and reporting the status change to the central processor at the same time, updating the node availability information in the distributed network topology to complete the node status adaptive management loop.

[0098] Among them, between the steps of generating the camera control instruction and transmitting the camera control instruction back to the camera node through the distributed network topology, there is also a step of securing the control instruction and video data, including: performing multi-dimensional security risk assessment on the streamlined control instruction data packet, analyzing the instruction type, influence scope, and execution permission, assigning three levels of encryption strength of 128 bits, 192 bits, and 256 bits to control instructions of different security levels, and establishing a control instruction security classification table; performing lightweight encryption processing on the control instruction based on the control instruction security classification table, generating an instruction signature using the elliptic curve cryptography algorithm, deriving an instruction key by combining the unique identification code and timestamp of the camera node, and forming a tamper-proof encrypted control instruction; adding a variable-length challenge-response authentication code to the encrypted control instruction, generating a random challenge sequence through the portable Wi-Fi authentication node, requiring the receiving party to return a specific response sequence, and constructing a two-way authentication channel to ensure the legitimacy of the source of the control instruction; establishing an instruction transmission anomaly detection mechanism for the two-way authentication channel, statistically analyzing the instruction transmission frequency, instruction content similarity, and instruction complexity, detecting and marking suspicious DDoS attacks and instruction injection behaviors, and generating security threat warning information; dynamically adjusting the network defense strategy using the security threat warning information, and activating a hierarchical defense mechanism when detecting an attack behavior, including temporarily blocking the suspicious source, restricting the instruction receiving frequency, and forcing secondary authentication, to form an adaptive security defense configuration; implementing content integrity protection for the video data stream collected by the camera, dividing the original video frame into data blocks of a fixed size, calculating the hash value of each data block and connecting them in a chain, constructing a video data blockchain structure, and generating a tamper-proof video verification fingerprint; transmitting the video verification fingerprint separately from the video data through an independent encrypted channel, recalculating the hash value of the received video data at the central processing unit end and comparing it with the verification fingerprint to detect possible video tampering behaviors, and outputting the video data integrity verification result; establishing a security trust score system between the camera node and the central processing unit based on the video data integrity verification result and the adaptive security defense configuration, dynamically adjusting the encryption strength and verification frequency of data transmission, and generating a secure communication trust model to complete the security protection process of the control instruction and video data.

[0099] Please refer to Figure 2 , Figure 2 which is a schematic block diagram of the structure of the camera remote control system 200 based on portable Wi-Fi provided by an embodiment of the present application. As Figure 2 shown, the camera remote control system 200 based on portable Wi-Fi includes:

[0100] An acquisition module 210, configured to acquire the location information, coverage range, and signal strength of the camera node, and construct a distributed network topology;

[0101] A scene analysis module 220, configured to perform scene analysis on video content collected by a camera according to a distributed network topology structure, and calculate an optimal encoding parameter configuration table;

[0102] A compression processing module 230, configured to divide video frames into different priority regions according to the optimal encoding parameter configuration table, perform differential compression processing, and form video data packets with priority tags;

[0103] A data transmission module 240, configured to allocate video data packets with priority tags to a multi-level transmission queue, schedule data transmission according to the current bandwidth state of the portable Wi-Fi, and generate an ordered transmission stream.

[0104] Through the collaborative cooperation of the above-mentioned various components, by constructing a distributed network topology structure and real-time monitoring of the connection status of the portable Wi-Fi, the system can quickly sense network changes and make corresponding adjustments, and still maintain stable video transmission and control response under network fluctuation conditions, significantly reducing the video freezing rate and control delay. By adopting an adaptive encoding parameter configuration based on a network bandwidth quality model and combining a differential compression strategy with regional importance awareness, the system can greatly reduce the overall data volume while ensuring the quality of key content. Through a multi-level transmission queue management and priority sorting mechanism, it ensures that control instructions can still be reliably transmitted under network constraints. At the same time, by adopting a progressive control strategy and a real-time feedback mechanism, it guarantees the accuracy and real-time nature of remote control. By using a deep learning-based object detection technology to identify important regions in video frames, and applying differential encoding parameters and multi-level preprocessing operations, combined with a super-resolution reconstruction algorithm and an artifact removal filter to process low-quality regions caused by network constraints, it significantly improves the subjective and objective quality of surveillance videos. By designing a network anomaly handling strategy, an adaptive segmented retransmission mechanism, and a forward error correction coding technology, the system can effectively handle network anomalies such as a sudden drop in signal strength and a sudden increase in packet loss rate, ensuring the continuous operation ability of the surveillance system in a harsh network environment. By intelligently allocating computing and processing tasks, performing local preprocessing at the camera terminal, aggregating data at the edge processing unit, and performing video synthesis at the central processor, a hierarchical computing architecture is formed, effectively balancing the computing load and energy consumption of each camera node and improving the overall operation efficiency of the system.

[0105] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, system, and unit can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0106] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0107] As described above, the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of this application.

Claims

1. A camera remote control method based on portable WiFi, characterized in that: include: Collect the location information, coverage and signal strength of camera nodes to build a distributed network topology; According to the distributed network topology, scene analysis is performed on the video content collected by the camera to calculate the optimal encoding parameter configuration table; Dividing the video frame into different priority areas according to the optimal encoding parameter configuration table, performing differential compression processing, and forming a video data packet with a priority mark; The video data packets with priority marks are allocated to the multi-level transmission queues, and the data transmission is scheduled according to the current bandwidth status of the portable WiFi to generate an ordered transmission stream.

2. The camera remote control method based on portable WiFi according to claim 1 is characterized in that: The collecting of the location information, coverage and signal strength of the camera nodes and the construction of a distributed network topology structure include: The portable WiFi module is used to assign unique identification codes to each camera node to obtain node identification data; Based on the node identification data, periodically detect the connection between the camera and the portable WiFi module to obtain connection status parameters; Performing similarity calculation and dynamic grouping of camera nodes according to the connection state parameters to form a camera node group; Recording geographic location coordinates and viewing angle information for fixed cameras in the camera node group, and additionally recording real-time location and movement trajectory for mobile cameras, to generate a camera characteristic data set; Transmitting the camera characteristic data set to an edge processing unit deployed at a portable WiFi access point, performing local data aggregation processing, and generating preliminary processing results; The preliminary processing results and the connection status parameters are used to construct a network resource status diagram, calculate the network connection quality, available bandwidth and delay characteristics between each camera node, and complete the establishment of a distributed network topology structure.

3. The camera remote control method based on portable WiFi according to claim 1 is characterized in that: The method of performing scene analysis on the video content collected by the camera according to the distributed network topology structure and calculating the optimal encoding parameter configuration table includes: A multi-threaded parallel acquisition mechanism is designed for each camera node to synchronously monitor the instantaneous uplink bandwidth, instantaneous downlink bandwidth, data packet round-trip delay, and data packet loss rate to obtain the original network status data; Performing camera ID tagging and timestamp encoding on the raw network status data to form a time series network status data set; Perform outlier filtering, smoothing and trend extraction operations based on the time series network status data set to obtain pre-processed network status data; Inputting the preprocessed network status data into a random forest algorithm, analyzing the video quality level under different bandwidth conditions, establishing a functional relationship between bandwidth and encoding parameters, and generating a preliminary bandwidth quality map; The preliminary bandwidth quality mapping is classified and refined according to the complexity index of the camera acquisition scene, and bandwidth requirement difference models of static scenes, low-motion scenes and high-motion scenes are respectively constructed to form a scene adaptive mapping matrix; Combining historical data with real-time data, assigning time decay weights to the scenario adaptive mapping matrix, and generating a network bandwidth quality model; The network bandwidth quality model is used to perform scene analysis on the video content collected by the camera and calculate the optimal encoding parameter configuration table.

4. The camera remote control method based on portable WiFi according to claim 3 is characterized in that: The method of using the network bandwidth quality model to perform scene analysis on the video content collected by the camera and calculating the optimal encoding parameter configuration table includes: Perform pixel difference calculation, edge density statistics and motion vector distribution analysis on the video content captured by the camera to obtain the scene change coefficient; Classifying the video scene into a low-complexity scene, a medium-complexity scene and a high-complexity scene according to the scene change coefficient, and establishing a reference encoding parameter table for the video scene to form a scene classification result; Combining the scene classification result with the network bandwidth quality model to construct a coding parameter search space; Based on the coding parameter search space, multi-objective optimization calculation is performed on three objective functions of video quality index, transmission delay and energy consumption level to obtain an optimal coding parameter combination; The optimal coding parameter combination is hierarchically divided to form a hierarchical coding parameter set, and a parameter smoothing control mechanism is introduced into the hierarchical coding parameter set to generate an optimal coding parameter configuration table.

5. The camera remote control method based on portable WiFi according to claim 1, characterized in that: The step of dividing the video frames into different priority areas according to the optimal encoding parameter configuration table, performing differential compression processing, and forming a video data packet with a priority mark includes: The video frames captured by the camera are input into the target detection neural network based on the improved MobileNet-SSD model to identify people, vehicles and abnormal activity areas, and generate a target position coordinate set; Performing pixel-level weight calculation on the video frame based on the target position coordinate set to form an importance weight map; Performing regional adjustment on the quantization parameters in the optimal coding parameter configuration table according to the importance weight map to obtain a regional adaptive coding parameter table; Performing spatial domain preprocessing on the video data based on the region adaptive encoding parameter table to obtain spatial domain processed video data; Performing dual preprocessing in the time domain and the frequency domain on the video data processed in the spatial domain to generate multi-level preprocessed video data; The data packet size and the transmission unit organization mode are dynamically adjusted according to the multi-level pre-processed video data and the current portable WiFi network status to form a video data packet with a priority mark.

6. The camera remote control method based on portable WiFi according to claim 1, characterized in that: The step of allocating the priority-marked video data packets to a multi-level transmission queue, scheduling data transmission according to the current bandwidth status of the portable WiFi, and generating an ordered transmission stream includes: Classify the video data packets with priority marks according to the importance of the content, timeliness and system resource occupation of the video data packets, and establish a four-layer transmission priority queue including emergency level, high priority, medium priority and low priority; Performing real-time bandwidth measurement on the portable WiFi network, generating a network bandwidth prediction model, and performing bandwidth allocation calculation on the four-layer transmission priority queue based on the network bandwidth prediction model to obtain a queue transmission quota table; The transmission protocol is optimized according to the queue transmission quota table and the round-trip delay of the portable WiFi network to form a transmission control parameter set; Forward error correction processing is performed on important data packets in the transmission control parameter set to generate a transmission data stream, and network anomaly monitoring is performed on the transmission data stream to generate an ordered transmission stream.

7. The camera remote control method based on portable WiFi according to claim 1, characterized in that: The camera remote control method based on portable WiFi also includes: Reorganize the received video data packets according to the metadata information in the ordered transmission stream, set differentiated buffer sizes for the video streams of different cameras, and form a complete video stream; Based on the camera position and viewing angle information recorded in the distributed network topology structure, feature point matching and perspective transformation are performed on the overlapping areas in the complete video stream to generate a panoramic monitoring view; For low-quality areas in the panoramic monitoring view caused by network limitations, a super-resolution reconstruction algorithm based on the ESRGAN network and a de-artifacting filter are applied to obtain a quality-enhanced video image; Performing target detection, multi-target tracking and behavior recognition calculations on the quality-enhanced video images, monitoring abnormal activities in real time, and generating target behavior analysis results; Generate a targeted camera control instruction set according to the target behavior analysis result to form a simplified control instruction data packet; The simplified control instruction data packet is encapsulated through a lightweight encrypted communication protocol, and a progressive control strategy is adopted to decompose the complex control instruction into a basic operation sequence, and the control instruction is transmitted back to each camera node through the communication link established in the distributed network topology structure.

8. A camera remote control system based on portable WiFi, characterized in that: Used to execute the camera remote control method based on portable WiFi as described in any one of claims 1 to 7, the camera remote control system based on portable WiFi comprises: The acquisition module is used to collect the location information, coverage and signal strength of camera nodes and build a distributed network topology structure; A scene analysis module, used to perform scene analysis on the video content collected by the camera according to the distributed network topology structure, and calculate the optimal encoding parameter configuration table; A compression processing module, used to divide the video frame into different priority areas according to the optimal encoding parameter configuration table, perform differential compression processing, and form a video data packet with a priority mark; The data transmission module is used to distribute the video data packets with priority marks to the multi-level transmission queue, schedule the data transmission according to the current bandwidth status of the portable WiFi, and generate an ordered transmission stream.

Citation Information

Patent Citations

  • Multi-path data packet scheduling method based on real-time video structure driving

    CN118400317A

  • Mobile terminal remote control monitoring system and method

    CN119135847A

  • Distributed intelligent video image scheduling system

    CN119155522A

  • Monitoring video transmission method and system based on wireless communication

    CN119421008A

  • Data transmission method, system and device based on portable WiFi network

    CN119485573A

Cited By

  • Lens picture quality analysis method based on video image

    CN120298416A

  • Wireless optimization transmission method based on smart home

    CN120416937A

  • Data transmission method and device based on accelerator card

    CN120762895A

  • Intelligent network control method for low-delay video return and related equipment

    CN120835184A

  • Intelligent network control methods and related equipment for low-latency video backhaul

    CN120835184B