Video stream transmission method, device, electronic device and storage medium
By generating logical addresses and establishing mapping relationships, the limitations of traditional static IP configuration in dynamic network environments are resolved, automatic updating of mapping tables is achieved, manual intervention and error risks are reduced, and system reliability and maintenance efficiency are improved.
Patent Information
- Application Number
- CN202510634969.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The static IP configuration of traditional edge gateways requires manual intervention in a dynamic network environment, resulting in high costs and high error risks. This is especially cumbersome and error-prone in large-scale deployments and mixed heterogeneous device scenarios.
By parsing the device metadata of the network camera device, generating a logical address, and establishing a mapping relationship between the IP address and the logical address, the mapping table is automatically updated to achieve dynamic IP transparent access and reduce manual intervention.
It reduces the cost and error risk of manual configuration, improves system reliability and maintenance efficiency, and is suitable for large-scale deployment and dynamic network environments.
Smart Images

Figure CN120151317B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of video stream processing, and in particular to a video stream transmission method, device, electronic device and storage medium. Background Art
[0002] With the advancement of video surveillance technology, edge gateways play a crucial role in video streaming. Traditional edge gateways use static Internet Protocol (IP) configurations, requiring manual configuration of the IP addresses of the Internet Protocol Camera (IPC) and the upstream platform. Furthermore, the edge gateway must manually adapt to the upstream platform's requirements, such as specifying the encoding format and encryption. Once adapted, the corresponding video stream is pushed to the upstream platform.
[0003] The connection between the edge gateway and the IPC and upper-level platform primarily relies on static IP configuration. Each time a connection is established, the edge gateway generates a new static IP. This static IP also changes when the dynamic environment changes or the edge gateway restarts. Traditional solutions require manual configuration of the changed static IP address on the IPC and upper-level platform. Furthermore, in large-scale surveillance scenarios (such as smart cities deploying thousands of cameras), manually maintaining dynamic IP addresses is extremely costly.
[0004] When connecting an IPC, you need to configure its IP address, subnet mask, and other parameters one by one on the edge gateway, which is cumbersome and prone to errors. In scenarios where heterogeneous devices are mixed (such as cameras from different brands with different protocols), the manual adaptation workload increases exponentially. Summary of the Invention
[0005] The technical problem to be solved by the embodiments of the present application is to provide a video stream transmission method, device, electronic device and storage medium to solve the limitations of traditional static IP configuration in a dynamic network environment. When the IP address of the edge gateway changes, the system will automatically update the mapping table without manual intervention. This can significantly improve the reliability and maintenance efficiency of the system while reducing the cost of manually maintaining dynamic IP and the workload of manual configuration. It not only reduces the risk of errors caused by manual configuration, but also is suitable for large-scale deployment and dynamic network environments.
[0006] In a first aspect, an embodiment of the present application provides a video stream transmission method, the method comprising:
[0007] Parsing a registration request sent by a network camera device to obtain device metadata of the network camera device;
[0008] Processing the device metadata to obtain a logical address corresponding to the network camera device;
[0009] Establishing a mapping relationship between a local IP address and the logical address, and establishing a connection between the network camera device and the upper platform based on the logical address;
[0010] Obtaining a video stream to be transmitted sent by the network camera device through the logical address;
[0011] The video stream to be transmitted is transmitted to the upper-level platform via the logical address.
[0012] In a second aspect, an embodiment of the present application provides a video stream transmission device, the device comprising:
[0013] A metadata acquisition module, configured to parse a registration request sent by a network camera device and obtain device metadata of the network camera device;
[0014] A logical address acquisition module, configured to process the device metadata to obtain a logical address corresponding to the network camera device;
[0015] A connection establishment module, configured to establish a mapping relationship between a local IP address and the logical address, and to establish a connection between the network camera device and a higher-level platform based on the logical address;
[0016] A video stream acquisition module, configured to acquire the video stream to be transmitted sent by the network camera device via the logical address;
[0017] The video stream transmission module is used to transmit the video stream to be transmitted to the upper-level platform through the logical address.
[0018] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0019] A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements any one of the above-mentioned video stream transmission methods when executing the program.
[0020] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute any of the above-mentioned video stream transmission methods.
[0021] Compared with the prior art, the embodiments of the present application have the following advantages:
[0022] In an embodiment of the present application, the device metadata of the network camera device is obtained by parsing the registration request sent by the network camera device. The device metadata is processed to obtain the logical address corresponding to the network camera device. A mapping relationship between the local IP address and the logical address is established, and a connection is established between the network camera device and the upper platform based on the logical address. The video stream to be transmitted sent by the network camera device through the logical address is obtained. The video stream to be transmitted is transmitted to the upper platform through the logical address. The embodiment of the present application solves the limitations of traditional static IP configuration in a dynamic network environment through dynamic IP transparent access and virtual logical address mapping technology. When the edge side gateway IP address changes, the system will automatically update the mapping table without manual intervention. It can reduce the cost of manual maintenance of dynamic IP and the workload of manual configuration while significantly improving the reliability and maintenance efficiency of the system. It not only reduces the risk of errors caused by manual configuration, but also is suitable for large-scale deployment and dynamic network environments.
[0023] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0025] Figure 1 A flowchart of a video streaming transmission method provided in an embodiment of the present application;
[0026] Figure 2 A schematic diagram of a model processing flow provided in an embodiment of the present application;
[0027] Figure 3 A schematic structural diagram of a video stream transmission device provided in an embodiment of the present application;
[0028] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0030] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0031] The terms used in the embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "an", "the" and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.
[0032] Reference Figure 1 , shows a flow chart of the steps of a video stream transmission method provided by an embodiment of the present application. Figure 1 As shown, the video stream transmission method may include: step 101, step 102, step 103, step 104 and step 105.
[0033] Step 101: Parse a registration request sent by a network camera device to obtain device metadata of the network camera device.
[0034] The embodiments of the present application can be applied to an edge gateway, that is, the execution entity is the edge gateway.
[0035] Edge Gateway: An intelligent device deployed at the edge of the network, serving as a bridge between terminal devices and the cloud or central data center, with functions such as data processing, protocol conversion, and security protection.
[0036] Network cameras can be devices such as network cameras. Internet Protocol Cameras (IPCs) are a new generation of cameras that combine traditional cameras with network technology. They can transmit video, audio, alarm, and control signals over the network and are managed by a network monitoring host (NVR (Network Video Recorder) or monitoring management platform).
[0037] Device metadata is a collection of data that describes the characteristics and attributes of a network camera. It contains detailed information about the device, helping the system identify, manage, configure, and interact with other devices or systems. In this example, device metadata may include data such as the camera brand, protocol type (such as RTSP (Real-Time Streaming Protocol)), resolution, and geographic location.
[0038] In order to be able to access the monitoring system and be managed by the upper-level platform, when the network camera device meets the set conditions (such as the network camera device is powered on for the first time. Or the network parameters of the network camera device (such as IP address, subnet mask, etc.) change, or switch from one network environment to another. Or when the network camera device is restored to factory settings, its configuration information is reset, etc.), it can send a registration request to the edge gateway and send its own device metadata to the edge gateway. Specifically, the network camera device can send a data message for registration to the edge gateway, and carry device metadata, etc. in the data message.
[0039] After receiving the registration request sent by the network camera device, the edge gateway can parse the registration request and obtain the device metadata of the network camera device.
[0040] After parsing the registration request sent by the network camera device to obtain device metadata, step 102 is executed.
[0041] Step 102: Process the device metadata to obtain a logical address corresponding to the network camera device.
[0042] After parsing the registration request sent by the network camera device to obtain the device metadata, the device metadata can be processed to obtain the logical address corresponding to the network camera device. In this embodiment, the network topology of the network camera device can be determined based on the geographic location information of the network camera device and the communication relationship between the network camera device, the edge gateway, and the upper-level platform. The logical address is generated based on the network topology, the device brand information, and the protocol feature information of the network camera device. This implementation process can be described in detail in conjunction with the specific implementation method below.
[0043] In a specific implementation of the present application, the above step 102 may include:
[0044] Sub-step A1: Determine the network topology information of the network camera device based on the geographic location information and the communication relationship between the network camera device, the edge gateway and the upper-level platform.
[0045] In this embodiment, the device metadata may include: device brand information, protocol feature information, and geographic location information.
[0046] After parsing the network camera device's device metadata, the network topology can be determined based on the device's geographic location and its communication relationship with the edge gateway and higher-level platform. Geographic location information helps clarify the device's physical location distribution, while communication relationships reveal the data transmission paths and connection methods between the device and the edge gateway and higher-level platform. By combining these two pieces of information, the topology of the network camera device within the entire network can be constructed, clearly demonstrating the interconnections between devices and the flow of data.
[0047] After determining the network topology information of the network camera device based on the geographic location information and the communication relationship between the network camera device, the edge gateway and the upper-level platform, sub-step A2 is executed.
[0048] Sub-step A2: Generate the logical address according to the protocol feature information, the device brand information and the network topology information.
[0049] After obtaining the network topology information for a network camera device, a logical address corresponding to the network camera device can be generated based on the protocol feature information, device brand information, and network topology information. Specifically, the protocol feature information reflects the characteristics of the communication protocol used by the device. Different protocols have different address encoding rules or identification methods. Device brand information can be associated with specific address allocation strategies or naming conventions. Combining this information with the network topology information can generate a unique logical address for the network camera device. This address not only identifies the device's location on the network, but also reflects some of the device's characteristics and affiliation.
[0050] The embodiment of the present application generates a logical address of the network camera device according to the network topology diagram, device brand and protocol feature information of the network camera device for subsequent connection establishment. It is not affected by the dynamic changes of the IP address of the edge gateway and can adapt to the dynamic network environment.
[0051] In this embodiment, a logical address can be generated by a neural network model. The implementation process can be described in detail in conjunction with the following specific implementation methods.
[0052] In a specific implementation of the present application, the above sub-step A1 may include: sub-step B1, sub-step B2 and sub-step B3.
[0053] Sub-step B1: Input the protocol feature information, the device brand information and the network topology information into a pre-trained logical address generation model, wherein the logical address generation model includes: a convolutional neural network and a graph neural network.
[0054] In this embodiment, the training process of the logical address generation model (including the convolutional neural network and the graph neural network) can refer to the following steps:
[0055] 1. Data Collection and Preprocessing: Collect a large amount of sample data containing protocol feature information, device brand information, network topology information, and corresponding correct logical addresses (semantic addresses). Preprocess this data, such as encoding the protocol feature information and device brand information into numerical data, representing the network topology information as appropriate graph data, and normalizing the data to meet the model's input requirements.
[0056] 2. Model initialization: Initialize the logical address generation model, including the parameters of each layer of the Convolutional Neural Network (CNN) and Graph Neural Network (GNN). Random initialization is usually used, but some pre-trained weights can also be used to initialize some layers to speed up training.
[0057] 3. Forward propagation: The preprocessed sample data is fed into the model. First, a convolutional neural network extracts features from the protocol, device brand, and network topology information to generate protocol, brand, and topology feature vectors. These feature vectors are then fed into a graph neural network for further processing to generate a semantic address.
[0058] 4. Calculate loss: Compare the semantic address generated by the model with the real logical address in the sample data, and use an appropriate loss function (such as mean square error loss function, cross entropy loss function, etc.) to calculate the difference between the two and obtain the loss value.
[0059] 5. Back propagation and parameter update: Based on the loss value, the gradient of the parameters of each layer of the model is calculated through the back propagation algorithm, and then an optimization algorithm (such as stochastic gradient descent) is used to update the model parameters according to the gradient to reduce the loss value.
[0060] 6. Iterative training: Repeat steps 3 to 5, and continuously adjust the model parameters until the loss value converges to a smaller value or reaches the preset number of training rounds. At this time, the model training is considered complete.
[0061] After obtaining the protocol feature information, device brand information, and network topology information of the network camera device, the protocol feature information, device brand information, and network topology information may be input into a pre-trained logical address generation model.
[0062] Sub-step B2: Call the convolutional neural network to process the protocol feature information, the device brand information and the network topology map information to obtain the protocol feature vector of the protocol feature information, the brand feature vector of the device brand information and the topology feature vector of the network topology map information.
[0063] After the protocol feature information, device brand information and network topology information are input into the pre-trained logical address generation model, a convolutional neural network can be called to process the protocol feature information, device brand information and network topology information to obtain the protocol feature vector of the protocol feature information, the brand feature vector of the device brand information and the topology feature vector of the network topology information.
[0064] Specifically, the protocol feature information is processed by a convolutional neural network through a series of convolutional layers, pooling layers, and activation functions. The convolution kernels in the convolutional layers slide over the protocol feature data to extract local features. The pooling layers compress and reduce the dimensionality of the extracted features, reducing the data volume while retaining important features. The activation function introduces nonlinear factors to enhance the model's expressiveness. Through these operations, the protocol feature information is converted into a protocol feature vector, which contains the key features of the protocol.
[0065] The processing process for device brand information can be as follows: Device brand information undergoes a similar convolutional neural network processing process. Because device brand information is typically categorical data, it requires encoding before input into the model, such as using one-hot encoding or word embedding. After multiple layers of convolutional neural network processing, a brand feature vector representing the characteristics of the device brand is obtained.
[0066] The network topology graph processing process can be as follows: The network topology graph is represented as graph-structured data, where nodes represent network elements such as devices and gateways, and edges represent the connections between them. This graph data is then fed into a convolutional neural network, where features are extracted from the graph structure through operations such as a specially designed graph convolution layer. The graph convolution layer updates the feature representation of a node based on its neighbor information, thereby capturing the structural features of the network topology graph. After processing by the convolutional neural network, a topological feature vector is obtained for the network topology graph, which reflects the topological structure of the network.
[0067] Sub-step B3: Call the graph neural network to process the protocol feature vector, the brand feature vector and the topology feature vector to obtain a semantic address, and use the semantic address as the logical address.
[0068] After obtaining the protocol feature vector, brand feature vector, and topology feature vector, a graph neural network can be called to process the protocol feature vector, brand feature vector, and topology feature vector to obtain a semantic address, and the semantic address can be used as a logical address. Specifically, the graph neural network can be used to fuse and further process these feature vectors through a series of graph neural network layers. In the graph neural network layer, nodes exchange information through a message passing mechanism, so that the relationship between different feature vectors and network topology information can be comprehensively considered. Through the processing of multi-layer graph neural networks, these feature vectors are fused into a semantic address. This address can accurately reflect the various characteristics of the network camera device and its semantic information such as its location and relationship in the network.
[0069] The embodiments of this application utilize convolutional neural networks to automatically extract key features from protocol signatures, device brand information, and network topology information, eliminating the need for manual feature extraction methods and significantly improving the efficiency and accuracy of feature extraction. For complex network topology information, graph convolutional neural networks can effectively capture its structural features, providing strong support for generating accurate logical addresses.
[0070] After the device metadata is processed to obtain the logical address corresponding to the network camera device, step 103 is executed.
[0071] Step 103: Establish a mapping relationship between the local IP address and the logical address, and establish a connection between the network camera device and the upper-level platform based on the logical address.
[0072] The upper level platform is a central system for centralized management, monitoring and processing of network camera device data. In this example, the upper level platform can be one platform or multiple platforms, and this embodiment does not limit this.
[0073] After processing the device metadata to obtain the logical address corresponding to the network camera device, a mapping relationship between the local IP address and the logical address can be established. Specifically, a mapping table can be generated in advance, and the mapping relationship between the IP address and the logical address can be saved in the mapping table for management. In this example, the mapping relationship can be stored in a high-performance memory database. When the IP address of the edge gateway changes, the system automatically updates the mapping table without manual intervention, thereby significantly improving the reliability and maintenance efficiency of the system.
[0074] At the same time, a connection between the network camera device and the upper platform can be established based on the logical address. Specifically, the logical address can be configured on both the network camera device side and the upper platform side to establish a connection between the network camera device-edge side gateway-upper platform.
[0075] After the connection between the network camera device and the upper-level platform is established based on the logical address, step 104 is executed.
[0076] Step 104: Acquire the video stream to be transmitted sent by the network camera device via the logical address.
[0077] After establishing a connection between the network camera device and the upper-level platform based on the logical address, the video stream sent by the network camera via the logical address can be retrieved. Specifically, after capturing the image, the network camera encodes and compresses the raw video data to reduce data transmission volume. Common encoding formats include H.264 and H.265, which effectively reduce bandwidth usage while maintaining video quality. The encoded video data is encapsulated into packets in a specific format, such as RTP (Real-Time Transport Protocol) packets. In addition to the video data, these packets also contain information such as a timestamp and sequence number, ensuring correct video decoding and playback at the receiving end. The device also enters the logical address in the destination address field of the packet to clarify the data transmission direction.
[0078] After the video stream to be transmitted sent by the network camera device via the logical address is acquired, step 105 is executed.
[0079] Step 105: Transmit the video stream to be transmitted to the upper-level platform via the logical address.
[0080] After obtaining the video stream to be transmitted sent by the network camera device through the logical address, the video stream to be transmitted can be transmitted to the upper-level platform through the logical address. Specifically, after successfully obtaining and confirming that the video stream data is correct, the edge-side gateway will re-encapsulate the video stream data packet. In order to adapt to the communication requirements with the upper-level platform, the gateway may adjust the format of the data packet or add some additional header information, such as routing information, priority identification, etc. At the same time, the gateway will fill in the logical address of the upper-level platform in the destination address field of the data packet to plan the transmission path of the video stream. The edge-side gateway sends the encapsulated video stream data packet through the network. The data packet is transmitted in the network according to the routing rules and finally reaches the upper-level platform.
[0081] The embodiments of the present application solve the limitations of traditional static IP configuration in dynamic network environments through dynamic IP transparent access and virtual logical address mapping technology. When the IP address of the edge gateway changes, the system will automatically update the mapping table without manual intervention. This can significantly improve the reliability and maintenance efficiency of the system while reducing the cost of manual maintenance of dynamic IP and the workload of manual configuration. It not only reduces the risk of errors caused by manual configuration, but also is suitable for large-scale deployment and dynamic network environments.
[0082] In this embodiment, the neural network model can also be used to predict when the IP address of the edge gateway will change in the future, so as to ensure that the transmission of the video stream remains stable and efficient even when the IP address of the edge gateway changes. This implementation process can be described in detail in conjunction with the following specific implementation methods.
[0083] In a specific implementation of the present application, the method may further include: steps C1 to C7.
[0084] Step C1: Obtain local network status data.
[0085] In this embodiment, the network status data of the edge gateway can be obtained in real time through a network monitoring tool or an API (Application Program Interface) interface. For example, SNMP (Simple Network Management Protocol) is used to obtain real-time bandwidth, latency, packet loss rate, IP change history data, subnet topology (i.e., the connection architecture and layout of network devices within the subnet), and other information from the edge gateway. Data collection is performed at fixed time intervals (such as 1 minute) to ensure that the acquired data can reflect changes in the network status in a timely manner.
[0086] After obtaining the local network status data, execute step C2.
[0087] Step C2: inputting the network status data and the device metadata into a pre-trained dynamic IP prediction model, wherein the dynamic IP prediction model includes a long short-term memory network and a dynamic IP prediction network.
[0088] A dynamic IP prediction model is a model used to predict changes in a device's IP address in the future. The training process for a dynamic IP prediction model can include the following steps:
[0089] 1. Data Collection: Collect a large amount of historical network status data and device metadata. This historical network status data covers network parameters of the edge gateway at different points in the past, such as bandwidth utilization, network latency, and packet loss rate. This data can reflect how network status changes over time. Device metadata includes information about the edge gateway's hardware configuration, the number of connected devices, and device types. By collecting data over a long period of time and across multiple scenarios, we build a rich and representative dataset, providing a foundation for model training.
[0090] 2. Data preprocessing: Collected data is cleaned to remove noise and outliers, such as erroneous network latency data caused by equipment failure. The data is then normalized to map data of varying ranges and magnitudes to the same interval. For example, bandwidth utilization (0-100%) and network latency (0-1000ms) are normalized to the interval [0, 1] to help the model better learn data characteristics. Furthermore, the data is divided into training, validation, and test sets to ensure that the model fully learns data patterns during training and is objectively evaluated during the validation and testing phases.
[0091] 3. Model Building and Training: Build a dynamic IP prediction model that includes a long short-term memory (LSTM) network and a dynamic IP prediction network. The LSTM network effectively processes data with time series characteristics. Through memory cells and gating mechanisms, it captures the changing trends and long-term dependencies of network status data over time. The dynamic IP prediction network uses the features extracted by the LSTM and combines them with device characteristics to construct a complex mapping relationship to predict the probability of IP changes.
[0092] The model is trained using the training set, employing an appropriate loss function, such as the cross-entropy loss function (suitable for probabilistic prediction tasks). The model parameters are continuously adjusted through the backpropagation algorithm to ensure that the model's predicted IP change probability matches the actual IP change as closely as possible. During training, the model is evaluated using the validation set to monitor the training results and prevent overfitting or underfitting. Training is terminated when the model reaches optimal performance on the validation set.
[0093] 4. Model Evaluation and Optimization: Use the test set to conduct a final evaluation of the trained model, calculating metrics such as accuracy, recall, and F1 value to comprehensively assess the model's predictive performance. If the model performance does not meet expectations, analyze the model's prediction results to identify any issues, such as inaccurate predictions of IP changes under certain network conditions. Targeted adjustments to the model structure, additional training data, or optimization of training parameters can be made to further improve model performance.
[0094] After acquiring local network status data, the network status data and device metadata of the network camera device can be input into the pre-trained dynamic IP prediction model. Specifically, the acquired network status data and device metadata can be organized and encoded according to the format required by the model, such as by converting them into tensor form, and then input into the model through the model's input layer.
[0095] Step C3: Call the long short-term memory network to process the network status data and the device metadata to obtain a timing feature vector corresponding to the network status data and a device feature vector of the device metadata, wherein the timing feature vector is used to characterize the timing characteristics of the network status of the edge gateway.
[0096] Within the model, network status data is input into the LSTM network. Using its memory cells and gating mechanism, the LSTM network processes the time series of network status data step by step, learning long-term dependencies and changing trends in the data. It then outputs a time series feature vector corresponding to the network status data. Simultaneously, device metadata is extracted and converted to generate a device feature vector. This time series feature vector represents the time series characteristics of the edge gateway's network status.
[0097] Step C4: calling the dynamic IP prediction network to process the time series feature vector and the device feature vector to obtain an IP change probability, where the IP change probability represents the probability that the IP address of the edge gateway changes within a future time window.
[0098] Then, the time series feature vector and device feature vector output by the LSTM network can be input into the dynamic IP prediction network. The dynamic IP prediction network performs nonlinear transformation and combination on the feature vector through multiple layers of neurons and activation functions, learns the complex mapping relationship between features and IP change probability, and finally outputs the IP change probability. This IP change probability can be used to represent the probability that the IP address of the edge gateway will change in the future time window.
[0099] Step C5: According to the IP change probability, the target time when the IP address changes is predicted.
[0100] After obtaining the IP change probability, the target time for the IP address change can be predicted based on the IP change probability. Specifically, a probability threshold can be pre-set. When the IP change probability exceeds this threshold, the target time for the IP address change is predicted based on the time window and the current time. For example, if the time window is 1 hour and the current time is 10:00, and the IP change probability exceeds the threshold at 10:15 and continues to rise, the predicted target time is 11:00, and so on.
[0101] After the target time when the IP address changes is predicted based on the IP change probability, step C6 is executed.
[0102] Step C6: When the target time is reached, the target IP address that has changed locally is obtained.
[0103] After predicting the target time for the IP address change based on the IP change probability, the target IP address that has changed locally can be obtained when the target time is reached. Specifically, a timer can be set in the edge gateway. When the time reaches the target time, the current IP address of the edge gateway, i.e., the target IP address, can be obtained through the network configuration interface or device management tool.
[0104] Step C7: Update the mapping relationship between the IP address and the logical address to the mapping relationship between the target IP address and the logical address.
[0105] After obtaining the changed target IP address locally, the corresponding logical address record can be found in the local mapping table, and the original IP address can be replaced with the target IP address to complete the mapping update. At the same time, the relevant network devices and systems are notified of the latest mapping relationship to ensure normal data transmission.
[0106] The embodiments of the present application ensure that after an IP address change, the connection between the network camera device and the upper-level platform based on the logical address can be quickly restored to normal, maintaining the continuity of services such as video streaming and avoiding communication failures caused by untimely mapping updates. At the same time, by utilizing a pre-trained model, the patterns and features learned by the model from historical data are fully utilized to process and analyze current data, avoiding the time and resource consumption caused by repeated model training.
[0107] In this embodiment, when transmitting a video stream to a higher-level platform, a transcoding strategy of the higher-level platform can be generated according to the requirements of the higher-level platform to process the video stream before transmitting it. This implementation process can be described in detail in conjunction with the following specific implementation methods.
[0108] In a specific implementation of the present application, the above step 105 may include: sub-steps D1 to D4.
[0109] Sub-step D1: Obtain the video stream demand information corresponding to the upper-level platform and the local current network load information.
[0110] In this embodiment, the video stream demand information corresponding to the upper-level platform can be obtained in advance. Specifically, the upper-level platform can actively send the video stream demand information to the edge-side gateway through a pre-set communication protocol. This information contains a variety of key parameters, such as the desired video resolution (such as 1080P, 4K), frame rate (25fps, 30fps, etc.), encoding format (H.264, H.265), and special requirements for video content (such as whether encryption, watermarking, etc. are required). After receiving the request sent by the upper-level platform, the edge-side gateway parses the data and extracts the video stream demand information.
[0111] Furthermore, when generating the video stream transcoding strategy for the upper-level platform, local network load information can be obtained. In this example, obtaining local network load information relies on various network monitoring methods. First, the built-in monitoring functions of network devices (such as routers and switches) are utilized to obtain real-time network bandwidth utilization, network latency, packet loss rate, and other data through the SNMP protocol. Second, the edge gateway itself collects statistics on the number of connected network cameras and data transmission rates, combining this data to assess the current network load. For example, if the amount of data uploaded by the network camera per unit time is large and bandwidth utilization is nearing saturation, it indicates that the network load is high.
[0112] Sub-step D2: Generate a video stream transcoding strategy corresponding to the upper-level platform according to the video stream demand information and the network load information.
[0113] After obtaining the video stream demand information and the network load information, a video stream transcoding strategy corresponding to the upper-level platform can be generated according to the video stream demand information and the network load information.
[0114] In a specific implementation of the present application, the acquired video stream demand information and network load information can be integrated and analyzed. For example, if the upper-level platform requires a high-resolution, high-frame-rate video stream, but the current network load is high and the bandwidth is insufficient, then the video stream demand needs to be appropriately adjusted. Based on the analysis results, a corresponding video stream transcoding strategy is formulated. If the network load is too high, the video resolution and frame rate may be reduced, and a more efficient encoding format may be selected (such as converting H.264 to H.265, which has a smaller amount of video data after H.265 encoding at the same image quality) to reduce the amount of video stream data and reduce network transmission pressure; if the network load is low and there is sufficient bandwidth support, the video stream can be transcoded according to the needs of the upper-level platform, and the video can even be optimized based on meeting the needs, such as improving the image quality of the video. In addition, a dynamic transcoding strategy can be formulated based on the fluctuation of the network load to automatically adjust the video stream parameters when the network load changes.
[0115] In another specific implementation of the present application, the video stream demand information and network load can be processed by a pre-trained strategy generation model to generate a video stream transcoding strategy. This implementation process can be described in detail in conjunction with the following specific implementation methods.
[0116] In another specific implementation of the present application, the above sub-step D2 may include: sub-step E1 to sub-step E3.
[0117] Sub-step E1: inputting the video stream demand information and the network load information into a policy generation model, wherein the policy generation model includes: a feature encoding network and a policy generation network.
[0118] In this embodiment, the strategy generation model refers to a model used to generate a video stream transcoding strategy. The model training process may include the following steps:
[0119] 1. Data Collection and Annotation: Collect a large amount of historical video stream demand information and network load information, covering different scenarios, time periods, and various network conditions. Video stream demand information includes parameters such as resolution, frame rate, and encoding format, while network load information covers metrics such as bandwidth utilization, network latency, and packet loss rate. At the same time, annotate the collected data and, based on actual application scenarios, identify the optimal video stream transcoding strategies for different data combinations, such as reducing resolution or switching encoding formats. This creates a accurately annotated dataset, providing supervisory signals for model training.
[0120] 2. Data preprocessing: The collected raw data is cleaned to remove noise and outliers, such as incorrect bandwidth data or malformed demand information caused by equipment failure. The data is then normalized and standardized, converting data of varying types and magnitudes into a unified format and range to better meet the model's input requirements. The dataset is then divided into training, validation, and test sets to ensure the model fully learns the data patterns during training and is objectively evaluated during the validation and testing phases.
[0121] 3. Model Building and Initialization: Build a policy generation model consisting of a feature encoding network and a policy generation network. The feature encoding network can use a convolutional neural network (CNN) or Transformer architecture to extract features related to video stream demand and network load. The policy generation network can be built based on a recurrent neural network (RNN) or a multi-layer perceptron (MLP) to generate a video stream transcoding policy based on the encoded features. Initialize the model parameters using random initialization or transfer parameters from a pre-trained model to lay the foundation for model training.
[0122] 4. Model training and optimization: Use the training set to train the model, selecting an appropriate loss function, such as cross-entropy loss or mean squared error loss (depending on the specific form of policy generation). Use the backpropagation algorithm to continuously adjust the model parameters to minimize the error between the model's predicted video stream transcoding policy and the annotated true policy. During training, use the validation set to evaluate the model, monitor its training results, and prevent overfitting or underfitting. If the model's performance on the validation set no longer improves, optimize it by adjusting the learning rate, adding regularization terms, or changing the model structure until the model achieves optimal performance on the validation set.
[0123] 5. Model Evaluation and Deployment: The trained model is evaluated using a test set, calculating metrics such as accuracy, recall, and F1 score to comprehensively assess its performance. If the model meets performance requirements, it is deployed to the edge gateway, enabling it to process new video stream demand information and network load information in real time and generate appropriate video stream transcoding strategies.
[0124] After obtaining the video stream demand information of the upper-level platform and the local current network load information, the video stream demand information and the network load information can be input into the strategy generation model.
[0125] Sub-step E2: calling the feature coding network to perform feature coding processing on the video stream demand information and the network load information respectively, to obtain a demand feature vector corresponding to the video stream demand information and a load feature vector corresponding to the network load information.
[0126] First, a feature coding network may be called to perform feature coding processing on the video stream demand information and the network load information respectively, so as to obtain a demand feature vector corresponding to the video stream demand information and a load feature vector corresponding to the network load information.
[0127] In the feature encoding network, video stream demand information and network load information are processed through multiple layers of neurons and activation functions. For example, using a CNN architecture, the data undergoes operations such as convolutional and pooling layers to extract key features such as resolution and frame rate, ultimately generating a demand feature vector. Similarly, network operations transform network load information such as bandwidth utilization and network latency into a load feature vector. During the encoding process, the network automatically learns the feature representations in the data, highlighting key information that has a significant impact on policy generation.
[0128] Sub-step E3: calling the strategy generation network to process the demand feature vector and the load feature vector to obtain the video stream transcoding strategy.
[0129] After obtaining the demand feature vector and the load feature vector, the policy generation network can be called to process the demand feature vector and the load feature vector to obtain the video stream transcoding strategy. Specifically, after the policy generation network receives the demand feature vector and the load feature vector output by the feature encoding network, it learns the mapping relationship between the feature vector and the video stream transcoding strategy through nonlinear transformation and combination of multiple layers of neurons. Based on the input feature vector, the network will output a set of parameters or decisions. These parameters or decisions correspond to specific video stream transcoding strategies, such as adjusting the video resolution, selecting the encoding format, etc. For example, if the load feature vector shows that the network bandwidth is insufficient, and the demand feature vector requires high-resolution video, the policy generation network may output a transcoding strategy that reduces the resolution.
[0130] This embodiment of the application generates a targeted video stream transcoding strategy based on feature vectors, fully considering the actual video stream requirements and network load, and achieving a dynamic balance between video stream quality and network transmission performance. Compared with traditional fixed strategies or simple rule-based judgments, the strategy generated by this model is more intelligent and flexible, able to adapt to complex and changing network environments and diverse video stream requirements, effectively improving the efficiency and quality of video stream transmission.
[0131] Sub-step D3: Processing the video stream to be transmitted based on the video stream transcoding strategy to obtain a processed video stream.
[0132] After obtaining the video stream transcoding strategy, the transmitted video stream can be processed based on the strategy to produce a processed video stream. Specifically, the edge gateway invokes the video transcoding software or hardware module to process the transmitted video stream according to the generated transcoding strategy. For example, if the transcoding strategy requires reducing the video resolution from 4K to 1080P, the transcoding module will scale each frame of the video. If the encoding format needs to be changed, the transcoding module will re-encode the original video stream according to the new encoding standard. During the transcoding process, the audio portion of the video is also synchronized to ensure audio and video consistency. Furthermore, the transcoding module monitors the transcoding progress and processing results in real time. If any anomalies (such as transcoding failure or severe video quality degradation) are detected, they will be promptly reported to the edge gateway for processing.
[0133] Sub-step D4: transmitting the processed video stream to the upper-level platform.
[0134] After obtaining the processed video stream, it can be transmitted to the upper-level platform. Specifically, the edge gateway can encapsulate the processed video stream according to the communication protocol agreed with the upper-level platform, add the necessary header information (such as source address, destination address, video stream format identifier, etc.), and then send the encapsulated video stream data packet to the upper-level platform via the network.
[0135] The video stream transcoding strategy generated by the embodiment of this application can fully consider the network load and achieve a balance between video stream quality and network transmission performance while meeting the video streaming requirements of the upper-level platform. This not only avoids network congestion caused by excessive pursuit of video quality, but also prevents low video quality due to network limitations, thereby improving the overall efficiency of video streaming and user experience.
[0136] In this embodiment, different types of video streams can also be automatically identified and classified, and heterogeneous video streams can be divided into multiple channels for independent push according to the specific needs of the upper-level platform, thereby improving the efficiency and flexibility of video stream push. This implementation process can be described in detail in conjunction with the following specific implementation methods.
[0137] In a specific implementation of the present application, the above sub-step D3 may include: sub-step F1 and step F2.
[0138] Sub-step F1: establishing a data transmission channel with each of the upper-level platforms.
[0139] In this embodiment, when there are multiple upper-level platforms, a data transmission channel can be established with each upper-level platform.
[0140] Sub-step F2: Process the video streams to be transmitted separately according to the video stream transcoding strategy corresponding to each of the upper-level platforms to obtain the processed video streams corresponding to each of the upper-level platforms. The processed video streams contain the channel identifier of the corresponding data transmission channel. The channel identifier is used to indicate that the processed video streams corresponding to each of the upper-level platforms are transmitted to the corresponding upper-level platform through the corresponding data transmission channel.
[0141] Furthermore, the video streams to be transmitted can be processed separately according to the video stream transcoding policy corresponding to each upper-level platform, resulting in a processed video stream corresponding to each upper-level platform. The processed video streams contain the channel identifier of the corresponding data transmission channel, which indicates that the processed video stream corresponding to each upper-level platform should be transmitted to the corresponding upper-level platform via the corresponding data transmission channel. Specifically, the edge gateway retrieves the corresponding video stream transcoding policy from the policy repository based on the identifier of each upper-level platform. For example, if upper-level platform A requires a video resolution of 1080P and an encoding format of H.265, the edge gateway, after obtaining this policy, invokes the video transcoding module to process the video stream to be transmitted according to this policy. The transcoding module decodes the video stream, adjusts parameters (such as reducing resolution), and re-encodes it, converting the original video stream into the format and parameters required by upper-level platform A. After completing the video stream transcoding, the channel identifier of the corresponding upper-level platform data transmission channel is added to each processed video stream. The channel identifier can be a unique identifier or a specific number associated with the channel. By inserting a channel identifier into the header or specific field of the video stream data packet, it is clearly indicated that the processed video stream should be transmitted to the corresponding upper-level platform through the corresponding channel. For example, if processed video stream A corresponds to the data transmission channel of upper-level platform A, the identifier of channel A, "Channel-A-001", is added to its packet header. This allows network devices to accurately route it to the corresponding channel based on the identifier during transmission.
[0142] The embodiment of the present application can realize the push of multiple video streams through one channel, meet the different needs of multiple upper-level platforms, realize the push of heterogeneous data streams, and improve the flexibility of video stream push.
[0143] Next, combine Figure 2 A detailed description of the dynamic IP adaptation module, AI learning model, and multi-channel heterogeneous stream push process integrated in the edge gateway (in this example, the above multiple models can be integrated into a model with multiple functions).
[0144] 1. AI (Artificial Intelligence) learning model deployment
[0145] (1) Data collection and preprocessing:
[0146] 1. Enter data as shown in the following table:
[0147] Table 1 Input data
[0148]
[0149] As shown in Table 1, input data includes IPC device metadata, network status data, upper-level platform requirements, security logs, and historical performance data. IPC device metadata includes camera brand, protocol type, resolution, and geographic location, which can be used for logical address generation and device classification (device classification facilitates IPC device management by the edge gateway). Network status data can include real-time bandwidth, latency, packet loss rate, IP change history, and subnet topology (the distribution of various network devices (such as IPCs, edge gateways, and upper-level platforms) within a subnet and the connections between them), which can be used for dynamic IP prediction. Upper-level platform requirements can include information such as video format (H.264 / H.265) and encryption, which can be used for adaptive policy generation. Security logs can include abnormal traffic records (DDoS attacks, port scans), packet size distribution, and other log information, which can be used to train security protection models. Historical performance data can include transcoding time, bandwidth utilization, service interruption records, and mapping table update delays, which can be used for feedback optimization of the model.
[0150] 2. Data Preprocessing: 1) Data Cleaning: Eliminate invalid data (e.g., camera unresponsiveness records) and handle missing values (interpolation or deletion of incomplete samples). 2) Feature Computation: a) Categorical Data Encoding: Convert protocol type, device brand, and other data to one-hot encoding (one-hot encoding). Each category corresponds to a dimension, and each category is mapped to a unique binary vector. Assuming the categorical data has k categories, the one-hot encoding of each sample is a binary vector of length k.
[0151] b) Time series feature extraction: IP change frequency and network load fluctuation period (i.e., the mean, variance, and trend of the network load). IP change frequency is calculated by counting the number of IP changes within the time window T. IP change probability = number of IP changes / T.
[0152] 3) Data normalization: Normalize the calculated feature data and map the data to a specific interval of [0, 1] to facilitate subsequent feature extraction and model training.
[0153] (2) Model architecture design (such as Figure 2 shown).
[0154] The AI model needs to handle dynamic IP configuration, logical address allocation, policy generation, and other tasks simultaneously, using a layered architecture:
[0155] Data input layer: Receives data from multiple sources (device metadata, network status, platform requirements).
[0156] Shared feature layer: 1) Convolutional neural network: Extracts device protocol features, such as the camera's protocol type. 2) Long short-term memory network (LSTM): Captures temporal dependencies (such as IP change cycles).
[0157] Task-specific layer: 1) Dynamic IP prediction branch: The fully connected network (FCN) outputs the probability of IP change; 2) Logical address generation branch: The graph neural network (GNN) models device topology relationships and outputs semantic addresses; 3) Strategy generation branch: The reinforcement learning (RL) agent outputs the upper-level platform transcoding strategy.
[0158] Downstream processing: The updated mapping table is fed into the video stream processing engine. The security response module intercepts abnormal traffic.
[0159] Feedback loop strategy: The performance data of the video stream processing engine is fed back to the strategy generation branch to continuously optimize the model.
[0160] (3) Model training:
[0161] A. Supervised Learning:
[0162] 1) Dynamic IP prediction:
[0163] Input: device metadata + network status time series data;
[0164] The output is: IP change probability (0-1), which indicates the possibility of IP address change within a certain time window in the future. After predicting the IP change, the mapping table update is triggered in advance to reduce service interruption.
[0165] 2) Logical address generation:
[0166] Input: Device topology (nodes are IPCs, edge gateways, and upper-level platforms, and edges are communication relationships).
[0167] Output: Semantic address, including device function, location, and other information.
[0168] B. Reinforcement Learning (Strategy Generation):
[0169] Generate the corresponding policy based on the following:
[0170] a. Environment: A dynamic IP network simulator simulates dynamic network changes such as bandwidth fluctuations and random IP changes, providing a test platform close to the real network environment for model training.
[0171] b. State space: Network load: The current network traffic load, such as bandwidth usage. Platform requirements: The requirements of the upper-level platform for video streaming, such as video format and encryption.
[0172] c. Action Space: Transcoding parameters: including video stream resolution, bit rate and other parameters, which affect the quality of the video stream. Encryption switch (enable / disable): determines whether to encrypt the video stream.
[0173] d. Algorithm: Proximal policy optimization algorithms are used to balance exploration (exploring new strategies) and exploitation (using known effective strategies). Proximal policy optimization algorithms optimize strategies to maximize long-term rewards while maintaining the stability of policy updates.
[0174] C. Branch network for deploying security detection in the model.
[0175] Specifically, the training process can be as follows: 1. Preprocess the security log data, such as data cleaning, standardization, and data labeling. 2. Model Architecture Design: A convolutional neural network is used to process the visual features of the video stream, and a long short-term memory network (LSTM) is used to process the time series features of the video stream. In addition, a multi-layer perceptron is added to receive features extracted from the base network layers (i.e., the convolutional neural network and LSTM) as well as the security log data as input. 3. Training Data Partitioning: The preprocessed data is divided into a training set, a validation set, and a test set. Typically, the training set is used for model training, the validation set is used to adjust the model's hyperparameters and evaluate model performance, and the test set is used to ultimately evaluate the model's generalization ability. 4. Model Training: The model is trained using the training set, and the model parameters are adjusted by minimizing a loss function. The loss function can be a cross-entropy loss function, which measures the difference between the model's predictions and the ground-truth annotations. During training, the security detection branch learns how to identify video streams with security risks based on the security log data. The base model learns to extract features from the video streams and combines these features with the input of the security detection branch. Use the validation set to evaluate the performance of the model and calculate indicators such as accuracy, recall, and F1 value to determine the effectiveness of the model in identifying security risk video stream data. Based on the evaluation results, adjust the model's hyperparameters, such as the learning rate, number of network layers, and number of neurons, to optimize the model's performance.
[0176] D. Model integration and optimization:
[0177] Data collection: During subsequent operations, the edge gateway will continue to collect relevant data to help the model better adapt to device performance changes.
[0178] Model update and optimization: The edge gateway regularly retrains and updates the model using newly collected data to maintain model accuracy and reliability.
[0179] 2. Establish connections between IPC, edge gateway, and upper-level platform.
[0180] Establish connections between the IPC, edge gateway, and upper-level platform to enable secure push of video streams from the IPC to the upper-level platform.
[0181] (1) The logical address is generated by the AI learning model.
[0182] (2) IPC access:
[0183] 1) Automatic registration:
[0184] The IPC sends a data message (device ID, protocol type, etc.) to the edge gateway. The edge gateway automatically responds after receiving it, completing the IPC automatic registration.
[0185] The IPC only needs to send data to the logical address, and the gateway is responsible for mapping the logical address to the current dynamic IP.
[0186] 2) Dynamic IP adaptation:
[0187] When the edge gateway IP changes, the mapping table is automatically updated without manual intervention.
[0188] 3) Adaptation to the upper-level platform
[0189] A. Logical address replaces physical IP:
[0190] The edge gateway establishes a connection with the upper-level platform through the logical address and pushes the video stream through the logical address.
[0191] The edge gateway acts as a proxy and forwards the request to the current actual IP according to the mapping table.
[0192] B. Dynamic policy issuance:
[0193] The edge gateway provides an API interface or configuration interface, allowing the upper-level platform to dynamically update requirements (such as switching video formats and adjusting encryption).
[0194] Demand changes are synchronized to the AI module of the edge gateway in real time, triggering policy adjustments.
[0195] 4) Edge Gateway Connection:
[0196] Dynamic IP Adaptation Module Deployment: Logical Address Generation: The edge gateway periodically sends probe messages, monitors IPC and upper-level platform responses, and generates a unique logical address after receiving IPC registration information. Virtual Mapping Table Maintenance: Establishes a mapping relationship between dynamic IP and logical addresses and stores it in a high-performance in-memory database. When the edge gateway IP changes, the mapping table is updated to ensure that the logical address remains unchanged.
[0197] AI policy engine integration: Receives real-time requirements from the upper-level platform (such as H.265 transcoding and encryption) and dynamically generates push policies.
[0198] 3. Video stream adaptation strategy.
[0199] (1) Video stream push:
[0200] Video stream transmission: IPC pushes the video stream to the edge gateway through the logical address.
[0201] Temporary storage: After receiving the video stream, the edge gateway temporarily stores it locally.
[0202] (2) AI learning model processing:
[0203] Video stream analysis: The AI learning model analyzes the encoding format, protocol type, and other characteristics of the video stream in real time. Demand matching: Based on the requirements of the upper-level platform (such as H.264, H.265, SVAC (Smart Video / Audio Coding), encryption requirements, etc.), the AI learning model selects the appropriate conversion strategy.
[0204] (3) Format Conversion: Encoding conversion: The AI learning model converts the video stream into the format required by the upper-level platform, such as converting H.265 to H.264, or encrypting the video stream (H.264+). Encryption processing: The video stream is encrypted using the SM4 encryption algorithm.
[0205] 4: Push multi-channel heterogeneous video streams.
[0206] (1) Video stream splitting: Demand analysis: The edge gateway analyzes the specific requirements of each upper-level platform, including video format, protocol, encryption requirements, etc. Splitting strategy: Based on the demand analysis results, the video stream is split into multiple channels, and each channel is format converted for a different upper-level platform.
[0207] (2) Video stream push: Encapsulation: Encapsulate the video stream according to the needs of each upper-level platform to ensure that it complies with its protocol and format requirements. Format adjustment: Adjust the resolution, frame rate and other parameters of the video stream to adapt to the needs of different platforms.
[0208] (3) Multi-channel push: Push the packaged video stream to the corresponding upper-level platform through different communication channels. Dynamically adjust the video stream push strategy based on network bandwidth and the needs of the upper-level platform to ensure efficient transmission.
[0209] The video stream transmission method provided in the embodiment of the present application obtains the device metadata of the network camera device by parsing the registration request sent by the network camera device. The device metadata is processed to obtain the logical address corresponding to the network camera device. A mapping relationship between the local IP address and the logical address is established, and a connection is established between the network camera device and the upper platform based on the logical address. The video stream to be transmitted sent by the network camera device through the logical address is obtained. The video stream to be transmitted is transmitted to the upper platform through the logical address. The embodiment of the present application solves the limitations of traditional static IP configuration in a dynamic network environment through dynamic IP transparent access and virtual logical address mapping technology. When the edge side gateway IP address changes, the system will automatically update the mapping table without manual intervention. It can reduce the cost of manual maintenance of dynamic IP and the workload of manual configuration while significantly improving the reliability and maintenance efficiency of the system. It not only reduces the risk of errors caused by manual configuration, but also is suitable for large-scale deployment and dynamic network environments.
[0210] Reference Figure 3 , shows a schematic structural diagram of a video stream transmission device provided by an embodiment of the present application, which can be applied to an edge side gateway. Figure 3 As shown, the video stream transmission device 300 may include the following modules:
[0211] The metadata acquisition module 310 is used to parse the registration request sent by the network camera device and obtain the device metadata of the network camera device;
[0212] A logical address acquisition module 320 is configured to process the device metadata to obtain a logical address corresponding to the network camera device;
[0213] A connection establishing module 330 is used to establish a mapping relationship between a local IP address and the logical address, and to establish a connection between the network camera device and the upper platform based on the logical address;
[0214] The video stream acquisition module 340 is used to acquire the video stream to be transmitted sent by the network camera device through the logical address;
[0215] The video stream transmission module 350 is configured to transmit the video stream to be transmitted to the upper-level platform via the logical address.
[0216] Optionally, the device metadata includes: device brand information, protocol feature information and geographic location information of the network camera device,
[0217] The logical address acquisition module includes:
[0218] A topology determination unit, configured to determine network topology information of the network camera device based on the geographic location information and the communication relationship between the network camera device, the edge side gateway, and the upper-level platform;
[0219] A logical address generating unit is configured to generate the logical address according to the protocol feature information, the device brand information and the network topology information.
[0220] Optionally, the logical address generating unit includes:
[0221] An information input subunit, configured to input the protocol feature information, the device brand information, and the network topology information into a pre-trained logical address generation model, wherein the logical address generation model includes: a convolutional neural network and a graph neural network;
[0222] a vector acquisition subunit, configured to call the convolutional neural network to process the protocol feature information, the device brand information, and the network topology map information to obtain a protocol feature vector for the protocol feature information, a brand feature vector for the device brand information, and a topology feature vector for the network topology map information;
[0223] The logical address acquisition subunit is used to call the graph neural network to process the protocol feature vector, the brand feature vector and the topology feature vector to obtain a semantic address, and use the semantic address as the logical address.
[0224] Optionally, the device further comprises:
[0225] Network status acquisition module, used to obtain local network status data;
[0226] A data input module, configured to input the network status data and the device metadata into a pre-trained dynamic IP prediction model, wherein the dynamic IP prediction model includes: a long short-term memory network and a dynamic IP prediction network;
[0227] a feature vector acquisition module, configured to call the long short-term memory network to process the network status data and the device metadata to obtain a time series feature vector corresponding to the network status data and a device feature vector for the device metadata, wherein the time series feature vector is used to characterize the time series characteristics of the network status of the edge gateway;
[0228] a change probability acquisition module, configured to call the dynamic IP prediction network to process the time series feature vector and the device feature vector to obtain an IP change probability, where the IP change probability represents the probability that the IP address of the edge gateway will change within a future time window;
[0229] A target time prediction module, configured to predict a target time at which the IP address changes based on the IP change probability;
[0230] A target IP acquisition module is used to acquire the target IP address that has changed locally when the target time is reached;
[0231] A mapping relationship updating module is used to update the mapping relationship between the IP address and the logical address to the mapping relationship between the target IP address and the logical address.
[0232] Optionally, the video stream transmission module includes:
[0233] An information acquisition unit, configured to acquire the video stream demand information corresponding to the upper-level platform and the local current network load information;
[0234] A strategy generating unit, configured to generate a video stream transcoding strategy corresponding to the upper-level platform according to the video stream demand information and the network load information;
[0235] A video stream processing unit, configured to process the video stream to be transmitted based on the video stream transcoding strategy to obtain a processed video stream;
[0236] The video stream transmission unit is used to transmit the processed video stream to the upper-level platform.
[0237] Optionally, the policy generating unit includes:
[0238] A load information input subunit, configured to input the video stream demand information and the network load information into a strategy generation model, wherein the strategy generation model includes: a feature encoding network and a strategy generation network;
[0239] an information encoding subunit, configured to call the feature encoding network to perform feature encoding processing on the video stream demand information and the network load information respectively, to obtain a demand feature vector corresponding to the video stream demand information and a load feature vector corresponding to the network load information;
[0240] The strategy acquisition subunit is used to call the strategy generation network to process the demand feature vector and the load feature vector to obtain the video stream transcoding strategy.
[0241] Optionally, there are multiple upper-level platforms.
[0242] The video stream processing unit includes:
[0243] a channel establishing subunit, configured to establish a data transmission channel with each of the upper-level platforms;
[0244] The video stream acquisition sub-unit is used to process the video stream to be transmitted according to the video stream transcoding strategy corresponding to each of the upper-level platforms, and obtain the processed video stream corresponding to each of the upper-level platforms. The processed video stream contains the channel identifier of the corresponding data transmission channel, and the channel identifier is used to indicate that the processed video stream corresponding to each of the upper-level platforms is transmitted to the corresponding upper-level platform through the corresponding data transmission channel.
[0245] The video stream transmission device provided in the embodiment of the present application obtains the device metadata of the network camera device by parsing the registration request sent by the network camera device. The device metadata is processed to obtain the logical address corresponding to the network camera device. A mapping relationship between the local IP address and the logical address is established, and a connection is established between the network camera device and the upper platform based on the logical address. The video stream to be transmitted sent by the network camera device through the logical address is obtained. The video stream to be transmitted is transmitted to the upper platform through the logical address. The embodiment of the present application solves the limitations of traditional static IP configuration in a dynamic network environment through dynamic IP transparent access and virtual logical address mapping technology. When the edge side gateway IP address changes, the system will automatically update the mapping table without manual intervention. It can reduce the cost of manual maintenance of dynamic IP and the workload of manual configuration while significantly improving the reliability and maintenance efficiency of the system. It not only reduces the risk of errors caused by manual configuration, but also is suitable for large-scale deployment and dynamic network environments.
[0246] An embodiment of the present application further provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the above-mentioned video stream transmission method when executed by the processor.
[0247] Figure 4 FIG. 4 is a schematic diagram showing the structure of an electronic device 400 according to an embodiment of the present invention. Figure 4 As shown, electronic device 400 includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 402 or loaded from storage unit 408 into random access memory (RAM) 403. Various programs and data required for the operation of electronic device 400 can also be stored in RAM 403. CPU 401, ROM 402, and RAM 403 are connected to each other via bus 404. Input / output (I / O) interface 405 is also connected to bus 404.
[0248] Multiple components in the electronic device 400 are connected to the I / O interface 405, including an input unit 406, such as a keyboard, a mouse, a microphone, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disk, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0249] The various processes and processing described above may be performed by the processing unit 401. For example, the method of any of the above embodiments may be implemented as a computer software program, which is tangibly contained in a computer-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the CPU 401, one or more actions in the method described above may be performed.
[0250] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the above-mentioned video stream transmission method is implemented.
[0251] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0252] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0253] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminals (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal generate instructions for implementing the steps in the process. Figure 1a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0254] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing terminal to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0255] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal so that a series of operational steps are executed on the computer or other programmable terminal to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable terminal for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0256] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0257] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal comprising the element.
[0258] The above is a detailed introduction to the video stream transmission method, device, electronic device and computer-readable storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only intended to help understand the method and core ideas of the present application. At the same time, for those skilled in the art, based on the ideas of the present application, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present application.
Claims
1. A video stream transmission method, applied to an edge gateway, characterized in that: The method comprises: Parsing a registration request sent by a network camera device to obtain device metadata of the network camera device; Processing the device metadata to obtain a logical address corresponding to the network camera device; Establishing a mapping relationship between a local IP address and the logical address, and establishing a connection between the network camera device and the upper platform based on the logical address; Obtaining a video stream to be transmitted sent by the network camera device through the logical address; Transmitting the video stream to be transmitted to the upper-level platform through the logical address; The device metadata includes: device brand information, protocol feature information and geographic location information of the network camera device, The processing of the device metadata to obtain a logical address corresponding to the network camera device includes: determining network topology information of the network camera device based on the geographic location information and the communication relationship between the network camera device, the edge gateway, and the upper-level platform; and generating the logical address based on the protocol feature information, the device brand information, and the network topology information; The generating of the logical address according to the protocol feature information, the device brand information and the network topology map information includes: inputting the protocol feature information, the device brand information and the network topology map information into a pre-trained logical address generation model, the logical address generation model including: a convolutional neural network and a graph neural network; calling the convolutional neural network to process the protocol feature information, the device brand information and the network topology map information to obtain a protocol feature vector of the protocol feature information, a brand feature vector of the device brand information and a topology feature vector of the network topology map information; calling the graph neural network to process the protocol feature vector, the brand feature vector and the topology feature vector to obtain a semantic address, and using the semantic address as the logical address.
2. The method according to claim 1, characterized in that After parsing the registration request sent by the network camera device to obtain the device metadata of the network camera device, the method further includes: Get local network status data; Inputting the network status data and the device metadata into a pre-trained dynamic IP prediction model, wherein the dynamic IP prediction model includes: a long short-term memory network and a dynamic IP prediction network; Calling the long short-term memory network to process the network status data and the device metadata to obtain a time series feature vector corresponding to the network status data and a device feature vector of the device metadata, wherein the time series feature vector is used to characterize the time series characteristics of the network status of the edge gateway; Calling the dynamic IP prediction network to process the time series feature vector and the device feature vector to obtain an IP change probability, where the IP change probability represents a probability that the IP address of the edge gateway changes within a future time window; Predicting a target time when the IP address will change based on the IP change probability; When the target time is reached, obtaining the target IP address that has changed locally; The mapping relationship between the IP address and the logical address is updated to the mapping relationship between the target IP address and the logical address.
3. The method according to claim 1, characterized in that The transmitting the video stream to be transmitted to the upper-level platform through the logical address includes: Obtain the video stream demand information corresponding to the upper-level platform and the local current network load information; Generating a video stream transcoding strategy corresponding to the upper-level platform according to the video stream demand information and the network load information; Processing the video stream to be transmitted based on the video stream transcoding strategy to obtain a processed video stream; The processed video stream is transmitted to the upper-level platform.
4. The method according to claim 3, characterized in that Generating a video stream transcoding strategy corresponding to the upper-level platform according to the video stream demand information and the network load information includes: Inputting the video stream demand information and the network load information into a strategy generation model, wherein the strategy generation model includes: a feature encoding network and a strategy generation network; Calling the feature coding network to perform feature coding processing on the video stream demand information and the network load information respectively to obtain a demand feature vector corresponding to the video stream demand information and a load feature vector corresponding to the network load information; The strategy generation network is called to process the demand feature vector and the load feature vector to obtain the video stream transcoding strategy.
5. The method according to claim 3, characterized in that There are multiple upper-level platforms. The processing of the video stream to be transmitted based on the video stream transcoding strategy to obtain a processed video stream includes: Establishing a data transmission channel with each of the upper-level platforms; The video stream to be transmitted is processed separately according to the video stream transcoding strategy corresponding to each of the upper-level platforms to obtain the processed video stream corresponding to each of the upper-level platforms. The processed video stream contains the channel identifier of the corresponding data transmission channel, and the channel identifier is used to indicate that the processed video stream corresponding to each of the upper-level platforms is transmitted to the corresponding upper-level platform through the corresponding data transmission channel.
6. A video stream transmission device, applied to an edge gateway, characterized in that: The device comprises: A metadata acquisition module, configured to parse a registration request sent by a network camera device and obtain device metadata of the network camera device; A logical address acquisition module, configured to process the device metadata to obtain a logical address corresponding to the network camera device; A connection establishment module, configured to establish a mapping relationship between a local IP address and the logical address, and to establish a connection between the network camera device and a higher-level platform based on the logical address; A video stream acquisition module, configured to acquire the video stream to be transmitted sent by the network camera device via the logical address; A video stream transmission module, configured to transmit the video stream to be transmitted to the upper-level platform via the logical address; The device metadata includes: device brand information, protocol feature information, and geographic location information of the network camera device; the logical address acquisition module includes: a topology map determination unit, configured to determine network topology map information of the network camera device based on the geographic location information and the communication relationship between the network camera device, the edge gateway, and the upper-level platform; and a logical address generation unit, configured to generate the logical address based on the protocol feature information, the device brand information, and the network topology map information. The logical address generation unit includes: an information input subunit, used to input the protocol feature information, the device brand information and the network topology map information into a pre-trained logical address generation model, and the logical address generation model includes: a convolutional neural network and a graph neural network; a vector acquisition subunit, used to call the convolutional neural network to process the protocol feature information, the device brand information and the network topology map information to obtain the protocol feature vector of the protocol feature information, the brand feature vector of the device brand information and the topology feature vector of the network topology map information; a logical address acquisition subunit, used to call the graph neural network to process the protocol feature vector, the brand feature vector and the topology feature vector to obtain a semantic address, and use the semantic address as the logical address.
7. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the video stream transmission method according to any one of claims 1 to 5 when executing the program.
8. A computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the video stream transmission method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Remote video monitoring system and method based on public switched telephone network-Internet protocol (PSTN-IP) double-network cooperation
CN102307295A
Method and device for determining network topology and computer storage medium
CN112751714A