An internet of things video transmission method, system, device and medium
This IoT video transmission method, which uses dynamic frame extraction at the edge and frame interpolation at the server, solves the problems of bandwidth pressure and video integrity in IoT video transmission, and achieves efficient video data transmission and improved clarity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- E SURFING IOT CO LTD
- Filing Date
- 2026-01-22
- Publication Date
- 2026-06-02
AI Technical Summary
There is bandwidth pressure in IoT video transmission. Existing technologies cause image quality degradation through video compression encoding, and fixed frame extraction methods ignore the differences in the importance of video content, resulting in poor video integrity and viewing experience.
By performing inter-frame content analysis on IoT video stream data at the edge, key frames are dynamically extracted and then processed for frame interpolation on the server side to ensure the integrity and clarity of video transmission.
It effectively reduces the bandwidth required for IoT video transmission while improving the integrity and clarity of video transmission, thus enhancing the user's viewing experience.
Smart Images

Figure CN122137958A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet of Things (IoT) technology, and in particular to an IoT video transmission method, system, device, and medium. Background Technology
[0002] With the rapid development of IoT technology, video surveillance has been widely used in IoT scenarios such as smart cities, intelligent transportation, and industrial inspection. However, due to the high frame rate and high resolution video streams generated by front-end cameras requiring a large amount of uplink bandwidth, large-scale video surveillance faces enormous network bandwidth pressure, making IoT video transmission a key concern for relevant personnel.
[0003] Currently, related technologies typically reduce video transmission bandwidth in the following ways: 1. Video compression encoding: Video data is compressed using encoding standards such as H.264 / H.265, but the compression rate is limited and excessive compression will lead to a decrease in image quality and low video clarity, resulting in a poor video viewing experience for subsequent users; 2. Fixed frame dropping transmission: Discarding some video frames at fixed intervals, such as reducing from 30fps to 15fps. However, this method ignores the differences in the importance of video content and is prone to losing key action information of video content, resulting in low video integrity and a poor video viewing experience for subsequent users. Therefore, the problems existing in the current technology still need to be solved and optimized. Summary of the Invention
[0004] To address at least one of the aforementioned technical problems, this application provides an IoT video transmission method, system, device, and medium. This method can effectively reduce the bandwidth of IoT video transmission while improving the integrity and clarity of video transmission, thereby enhancing the user's video viewing experience.
[0005] According to a first aspect of this application, an Internet of Things (IoT) video transmission method is provided, applied at an edge device, the method comprising: Acquire first IoT video stream data; the first IoT video stream data includes several first video frames; Inter-frame content analysis is performed on the first IoT video stream data to obtain inter-frame difference data for each first video frame; Based on all the inter-frame difference data, the first IoT video stream data is dynamically extracted to obtain the second IoT video stream data; the second IoT video stream data includes several key frames, which are used to characterize the first video frame with the largest inter-frame difference data among all the first video frames within the corresponding frame interval window. The second IoT video stream data is sent to the server so that the server can perform frame interpolation on the second IoT video stream data. The server then obtains the third IoT video stream data and sends the third IoT video stream data to the user terminal.
[0006] In some embodiments, the step of performing inter-frame content analysis on the first IoT video stream data to obtain inter-frame difference data for each first video frame includes: Obtain edge operators; According to the edge operator, frame features are extracted from the second video frame and the third video frame to obtain the first frame feature of the second video frame and the second frame feature of the third video frame; the second video frame is any one of the first video frames in the first IoT video stream data; the third video frame is the first video frame adjacent to the second video frame in the first IoT video stream data; Based on the third video frame, perform inter-frame difference analysis on the second video frame to obtain inter-frame difference data of the second video frame.
[0007] In some embodiments, the step of dynamically extracting frames from the first IoT video stream data based on all the inter-frame difference data to obtain the second IoT video stream data includes: Obtain the original window frame group of the first IoT video stream data within the target interval window; the target interval window is any one of the frame interval windows. The original window frame group is compared with the inter-frame difference data to obtain the first target difference data. The first target difference data is used to characterize the largest inter-frame difference data among all the first video frames in the original window frame group. Based on the first target difference data, the original window frame group is filtered to obtain the key frame.
[0008] In some embodiments, the method further includes: Obtain the adaptive difference threshold; Based on the adaptive difference threshold, threshold analysis is performed on the original window frame group to obtain several second target difference data; the second target difference data is used to characterize the inter-frame difference data of all first video frames in the original window frame group that are greater than the adaptive difference threshold. Based on all the second target difference data, the original window frame group is frame filtered to obtain the value frame corresponding to each second target difference data. Based on the keyframes and all the value frames, the target window frame group is obtained.
[0009] In some embodiments, the method further includes: Obtain the target frame and several non-target frames within the target interval window, wherein the target frame is a value frame or the key frame; and the non-target frames are the first video frames within the target interval window other than the target frames. Based on all the non-target frames, frame motion features are extracted from the target frames to obtain motion feature information; Based on the motion feature information, the target frame is updated in terms of attributes to obtain the target frame with updated attributes.
[0010] According to a second aspect of this application, an Internet of Things (IoT) video transmission method is provided, applied to a server, the method comprising: Receive second IoT video stream data from the edge; The second IoT video stream data is subjected to frame interpolation processing to obtain the third IoT video stream data, and the third IoT video stream data is sent to the user terminal; The second IoT video stream data is obtained through the following steps: The edge device acquires first IoT video stream data; the first IoT video stream data includes several first video frames; The edge device performs inter-frame content analysis on the first IoT video stream data to obtain inter-frame difference data for each first video frame; The edge device dynamically extracts frames from the first IoT video stream data based on all the inter-frame difference data to obtain the second IoT video stream data; the second IoT video stream data includes several key frames, which are used to characterize the first video frame with the largest inter-frame difference data among all the first video frames within the corresponding frame interval window.
[0011] In some embodiments, performing frame interpolation on the second IoT video stream data to obtain the third IoT video stream data includes: Obtain a first real frame and a second real frame. The first real frame is any video frame in the second IoT video stream data, and the second real frame is the video frame that is adjacent to the first real frame on the time axis among all video frames in the second IoT video stream data. Bidirectional optical flow analysis is performed on the first real frame and the second real frame to obtain bidirectional optical flow data; Based on the bidirectional optical flow data, the first real frame and the second real frame are frame fused to obtain an intermediate frame, which is used to represent the video frame supplemented between the first real frame and the second real frame.
[0012] According to a third aspect of this application, an Internet of Things (IoT) video transmission system is provided, comprising: The edge device is used to acquire first IoT video stream data; the first IoT video stream data includes several first video frames; inter-frame content analysis is performed on the first IoT video stream data to obtain inter-frame difference data for each first video frame; based on all the inter-frame difference data, dynamic frame extraction is performed on the first IoT video stream data to obtain second IoT video stream data; the second IoT video stream data includes several key frames, the key frames being used to characterize the first video frame with the largest inter-frame difference data among all first video frames within the corresponding frame interval window. The server is configured to receive second IoT video stream data from the edge terminal; perform frame interpolation processing on the second IoT video stream data to obtain third IoT video stream data; and send the third IoT video stream data to the user terminal.
[0013] According to a fourth aspect of this application, an electronic device is provided, comprising: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor performs the method as described above.
[0014] According to a fifth aspect of this application, a computer-readable storage medium is provided, wherein a processor-executable program is stored, which, when executed by the processor, is used to implement the method as described above.
[0015] According to a sixth aspect of this application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium, wherein a processor of an electronic device reads the computer program from the computer-readable storage medium and executes the computer program, causing the electronic device to perform the method described above.
[0016] The beneficial effects of the technical solutions provided in this application are: This application provides an IoT video transmission method, system, device, and medium. The method acquires first IoT video stream data at an edge terminal. The first IoT video stream data includes several first video frames. Inter-frame content analysis is performed on the first IoT video stream data to obtain inter-frame difference data for each first video frame. Based on all the inter-frame difference data, dynamic frame extraction is performed on the first IoT video stream data to obtain second IoT video stream data. The second IoT video stream data includes several keyframes, which characterize the first video frame with the largest inter-frame difference data among all first video frames within a corresponding frame interval window. The second IoT video stream data is sent to a server, whereby the server performs frame interpolation processing on the second IoT video stream data, resulting in third IoT video stream data, which is then sent to a user terminal. This method, by performing inter-frame content analysis on the IoT video stream data at the edge terminal and dynamically extracting frames based on the inter-frame difference data of each video frame, combined with server-side frame interpolation of the extracted IoT video stream data, can effectively reduce the bandwidth of IoT video transmission while improving the integrity and clarity of video transmission, thereby enhancing the user's video viewing experience. Attached Figure Description
[0017] Figure 1 A flowchart illustrating an IoT video transmission method provided in an embodiment of this application; Figure 2 A detailed flowchart of step S120 provided for an embodiment of this application; Figure 3 A detailed flowchart of step S130 provided for an embodiment of this application; Figure 4 This is a schematic diagram of a first optional process for an Internet of Things (IoT) video transmission method provided in an embodiment of this application; Figure 5 This is a schematic diagram of a second optional process for an Internet of Things (IoT) video transmission method provided in an embodiment of this application; Figure 6 A flowchart illustrating another IoT video transmission method provided in an embodiment of this application; Figure 7 A detailed flowchart of step S620 provided for an embodiment of this application; Figure 8 A schematic diagram of the framework of an Internet of Things (IoT) video transmission system provided in an embodiment of this application; Figure 9 This is a structural block diagram of a computer device provided in an embodiment of this application. Detailed Implementation
[0018] The present application will be further described below with reference to the accompanying drawings and specific embodiments. The described embodiments should not be considered as limitations on the present application, and all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of the present application.
[0019] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0021] Currently, related technologies typically reduce video transmission bandwidth in the following ways: 1. Video compression encoding: Video data is compressed using encoding standards such as H.264 / H.265, but the compression rate is limited and excessive compression will lead to a decrease in image quality and low video clarity, resulting in a poor video viewing experience for subsequent users; 2. Fixed frame dropping transmission: Discarding some video frames at fixed intervals, such as reducing from 30fps to 15fps. However, this method ignores the differences in the importance of video content and is prone to losing key action information of video content, resulting in low video integrity and a poor video viewing experience for subsequent users. It should be noted that the aforementioned related technologies are only used to assist in understanding the technical solutions of this application and do not mean that they belong to the publicly disclosed prior art.
[0022] In view of this, embodiments of this application provide an IoT video transmission method, system, device, and medium. The method performs inter-frame analysis of IoT video stream data at the edge, which can extract frame feature information of each video frame in the IoT video stream data. This can avoid the loss of key action information in the video content to a certain extent, effectively improve video integrity, and thus improve the video viewing experience for subsequent users.
[0023] Furthermore, this method performs threshold analysis and frame filtering on the original window frame groups of IoT video stream data within each frame interval window at the edge. It can extract video frames with significant content changes based on the corresponding inter-frame difference data, fully consider the importance of video content, further avoid the loss of key action information in video content, effectively improve video integrity, and thus improve the video viewing experience for subsequent users.
[0024] Furthermore, this method dynamically extracts frames from IoT video stream data based on inter-frame difference data of each video frame at the edge. This allows for the extraction of key video frames from the original window frame group within each frame interval window of the IoT video stream data. By fully considering the importance of the video content and further avoiding the loss of key action information in the video content, the method effectively improves video integrity and thus enhances the video viewing experience for subsequent users. It also avoids the over-compression of video data, which helps improve video clarity.
[0025] The present application provides an IoT video transmission method, which can be specifically described through the following embodiments. First, an IoT video transmission method in the present application is described.
[0026] The IoT video transmission method provided in this application can be applied to IoT application scenarios, such as smart cities, intelligent transportation, and industrial inspection, which require video data transmission. In these IoT application scenarios, IoT service providers can use the method provided in this application to transmit IoT video data. This effectively reduces the bandwidth required for IoT video transmission while improving the integrity and clarity of the video transmission, thereby enhancing the user's video viewing experience.
[0027] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0028] Reference Figure 1 , Figure 1 This is a flowchart illustrating an IoT video transmission method provided in an embodiment of this application. The method is applied at the edge and includes, but is not limited to, steps S110 to S140: Step S110: Obtain first IoT video stream data; the first IoT video stream data includes several first video frames; In this embodiment of the application, a video stream data collected in real time by a front-end camera device in an IoT scenario can be transmitted to the edge device so that the edge device can obtain the first IoT video stream data.
[0029] Step S120: Perform inter-frame content analysis on the first IoT video stream data to obtain inter-frame difference data for each first video frame; In this embodiment of the application, inter-frame content analysis can be performed on each first video frame of the first IoT video stream data to obtain inter-frame difference data for each first video frame.
[0030] For example, in the first IoT video stream data, the second video frame is any one of the first video frames, and the third video frame is a first video frame adjacent to the first video frame. In a first implementation, the sum of the absolute differences or mean square error of all pixels between the second and third video frames can be calculated to extract global motion features between consecutive video frames, thereby taking into account the importance of the video content, and the calculation result can be determined as the inter-frame difference data of the second video frame.
[0031] Alternatively, in a second implementation, the difference in color histograms (such as RGB or HSV) between the second and third video frames can be calculated to capture the variation in overall tone between consecutive video frames, thereby taking into account the importance of the video content, and the calculation result can be determined as the inter-frame difference data of the second video frame.
[0032] Alternatively, in a third implementation, a lightweight CNN network (such as MobileNet, SqueezeNet) or a Vision Transformer network (such as MobileViT) can be used at the edge to extract the frame features of the second and third video frames, respectively. Then, the semantic changes between consecutive video frames are extracted by calculating the cosine distance or Euclidean distance between the frame features of the second and third video frames, and the calculation result is determined as the inter-frame difference data of the second video frame.
[0033] Reference Figure 2 In some embodiments, step S120, performing inter-frame content analysis on the first IoT video stream data to obtain inter-frame difference data for each first video frame, includes: Step S210: Obtain the edge operator; Step S220: Based on the edge operator, extract frame features from the second video frame and the third video frame to obtain the first frame feature of the second video frame and the second frame feature of the third video frame; the second video frame is any one of the first video frames in the first IoT video stream data; the third video frame is the first video frame in the first IoT video stream data that is adjacent to the second video frame. Step S230: Based on the third video frame, perform inter-frame difference analysis on the second video frame to obtain inter-frame difference data of the second video frame.
[0034] In the fourth implementation of inter-frame content analysis, the edge operator can be the Sobel operator, the Canny operator, etc. Frame feature extraction can be performed by extracting the gradient information of the second video frame and the gradient information of the third video frame respectively through the edge operator to obtain the feature gradient map of the second video frame (i.e., the first frame feature) and the feature gradient map of the third video frame (i.e., the second frame feature).
[0035] It is understandable that inter-frame difference analysis can be used to calculate the difference between the features of the first frame and the features of the second frame. The specific calculation method can be to calculate the mean square error between the features of the first frame and the features of the second frame, etc., to obtain the structural changes between consecutive video frames, thereby taking into account the importance of the video content, and determining the calculation result as the inter-frame difference data of the second video frame.
[0036] Step S130: Based on all the inter-frame difference data, dynamically extract frames from the first IoT video stream data to obtain the second IoT video stream data; the second IoT video stream data includes several key frames, which are used to characterize the first video frame with the largest inter-frame difference data among all the first video frames within the corresponding frame interval window. In this embodiment of the application, dynamic frame extraction can be performed on the first IoT video stream data based on the inter-frame difference data of each first video frame, and the extracted video frames can be determined as the second IoT video stream data. The extracted video frames can be key frames or value frames of the first IoT video stream data.
[0037] Reference Figure 3 In some embodiments, step S130, which involves dynamically extracting frames from the first IoT video stream data based on all the inter-frame difference data to obtain the second IoT video stream data, includes: Step S310: Obtain the original window frame group of the first IoT video stream data within the target interval window; the target interval window is any one of the frame interval windows. Step S320: Compare the inter-frame difference data of the original window frame group to obtain the first target difference data. The first target difference data is used to characterize the largest inter-frame difference data among all the inter-frame difference data of the first video frames in the original window frame group. Step S330: Based on the first target difference data, perform frame filtering on the original window frame group to obtain the key frame.
[0038] In this embodiment, the first IoT video stream data can be divided into frames based on a frame interval window to obtain several first video frames under each frame interval window. Specifically, for any frame interval window, denoted as the target interval window, several first video frames within the target interval window can be denoted as the original window frame group; the comparison of inter-frame difference data can be a comparison of the size relationship between the inter-frame difference data of each first video frame in the original window frame group to determine the largest inter-frame difference data among all the inter-frame difference data of the first video frames in the target interval window, which is denoted as the first target difference data.
[0039] It is understood that frame filtering can be done by identifying the first video frame corresponding to the first target difference data in the original window frame group as the key frame to obtain the key frame of the target interval window; then, in the first embodiment, the second IoT video stream data can be determined based on the key frame of each frame interval window.
[0040] Reference Figure 4 In some embodiments, the method further includes: Step S410: Obtain the adaptive difference threshold; Step S420: Perform threshold analysis on the original window frame group according to the adaptive difference threshold to obtain several second target difference data; the second target difference data is used to characterize the inter-frame difference data of all first video frames in the original window frame group that are greater than the adaptive difference threshold. Step S430: Based on all the second target difference data, perform frame filtering on the original window frame group to obtain the value frame corresponding to each second target difference data. Step S440: Obtain the target window frame group based on the key frame and all the value frames.
[0041] In this embodiment, the adaptive difference threshold can be a statistical difference value of the first IoT video stream data. Specifically, this statistical difference value can be the mean, median, or other values of the inter-frame difference data of each first video frame in the first IoT video stream data; alternatively, it can be the mean, median, or other values of the inter-frame difference data of each first video frame in several original window frame groups. For any original window frame group, threshold analysis can involve comparing the adaptive difference threshold with the inter-frame difference data of each first video frame in that original window frame group, and determining the inter-frame difference data that is greater than the adaptive difference threshold as the second target difference data.
[0042] It is understood that frame filtering can be done by identifying the first video frame corresponding to the second target difference data in the original window frame group as the value frame to obtain the value frame of the corresponding frame interval window; then, in the second embodiment, the key frame and all value frames of the corresponding frame interval window can be integrated into the target window frame group of the frame interval window; the target window frame groups of the remaining frame interval windows are similarly integrated, and the second IoT video stream data is determined based on the target window frame groups of all frame interval windows.
[0043] Reference Figure 5 In some embodiments, the method further includes: Step S510: Obtain the target frame and several non-target frames within the target interval window. The target frame is a value frame or the key frame. The non-target frames are the first video frames within the target interval window other than the target frames. Step S520: Based on all the non-target frames, extract frame motion features from the target frames to obtain motion feature information; Step S530: Update the attributes of the target frame according to the motion feature information to obtain the target frame with updated attributes.
[0044] In this embodiment, a non-target frame within the target interval window can be the first video frame within that window, excluding the value frame and / or keyframe. Frame motion feature extraction can be based on all non-target frames, extracting motion vectors and other information of the target frame, denoted as motion feature information. Specifically, this can be obtained using Motion Vector Prediction (MVP) technology.
[0045] It is understandable that attribute update can be achieved by adding motion feature information to the target frame so that the motion feature information is determined as the attribute information of the target frame, thereby obtaining the target frame after attribute update.
[0046] Step S140: Send the second IoT video stream data to the server so that the server performs frame interpolation on the second IoT video stream data, and the server obtains the third IoT video stream data and sends the third IoT video stream data to the user terminal.
[0047] In this embodiment, the edge device can send the second IoT video stream data to the server through the uplink bandwidth channel, so that the server can perform frame interpolation on the received IoT video stream data, thereby enabling the server to send the third IoT video stream data obtained by frame interpolation to the user terminal through the downlink bandwidth channel, so as to realize the playback of IoT video stream data on the user terminal.
[0048] Reference Figure 6 This application provides an IoT video transmission method applied to a server, the method comprising: Step S610: Receive second IoT video stream data from the edge device; Step S620: Perform frame interpolation processing on the second IoT video stream data to obtain the third IoT video stream data, and send the third IoT video stream data to the user terminal; The second IoT video stream data is obtained through the following steps: The edge device acquires first IoT video stream data; the first IoT video stream data includes several first video frames; The edge device performs inter-frame content analysis on the first IoT video stream data to obtain inter-frame difference data for each first video frame; The edge device dynamically extracts frames from the first IoT video stream data based on all the inter-frame difference data to obtain the second IoT video stream data; the second IoT video stream data includes several key frames, which are used to characterize the first video frame with the largest inter-frame difference data among all the first video frames within the corresponding frame interval window.
[0049] In the embodiments of this application, the contents of steps S610 to S620 are similar to those of the aforementioned step S140 and can be easily deduced by analogy.
[0050] Reference Figure 7 In some embodiments, step S620, performing frame interpolation on the second IoT video stream data to obtain the third IoT video stream data, includes: Step S710: Obtain the first real frame and the second real frame. The first real frame is any video frame in the second IoT video stream data, and the second real frame is the video frame that is adjacent to the first real frame on the time axis among all video frames in the second IoT video stream data. Step S720: Perform bidirectional optical flow analysis on the first real frame and the second real frame to obtain bidirectional optical flow data; Step S730: Based on the bidirectional optical flow data, perform frame fusion on the first real frame and the second real frame to obtain an intermediate frame. The intermediate frame is used to represent the video frame supplemented between the first real frame and the second real frame.
[0051] In this embodiment of the application, for any video frame (i.e., the first real frame) of the second IoT video stream data, a video frame (i.e., the second real frame) adjacent to the first real frame can be obtained. Specifically, the video frame can be a value frame or a key frame in the second IoT video stream data.
[0052] In the first embodiment, bidirectional optical flow analysis can be based on bidirectional optical flow technology to extract the motion vector between the first real frame and the second real frame, which is denoted as bidirectional optical flow data. Frame fusion can be based on the motion vector indicated by the bidirectional optical flow data, and generate corresponding video frames based on the first real frame and the second real frame through linear and equal deformation, which are denoted as intermediate frames. There are already many ways to generate supplementary video frames based on the motion vector and the first real frame and the second real frame, which will not be described in detail here.
[0053] In the second embodiment, bidirectional optical flow analysis can be based on bidirectional optical flow technology to extract the motion vector between the first real frame and the second real frame, and combine it with the motion feature information determined by the aforementioned corresponding target frame to determine the bidirectional optical flow data. Specifically, the combination method can be to perform operations such as summing and averaging or weighted summing on the motion vector between the first real frame and the aforementioned motion feature information. The subsequent frame fusion content is similar to the frame fusion content in the first embodiment and can be simply deduced by analogy.
[0054] It is understandable that after obtaining the intermediate frames of each video frame and adjacent video frames of the second IoT video stream data, the third IoT video stream data can be obtained by simply integrating all the determined intermediate frames and the original video frames of the second IoT video stream data.
[0055] Figure 8 A system block diagram of an Internet of Things (IoT) video transmission system provided in this application embodiment includes an edge terminal and a server terminal; The edge terminal 801 is used to acquire first IoT video stream data; the first IoT video stream data includes a plurality of first video frames; inter-frame content analysis is performed on the first IoT video stream data to obtain inter-frame difference data for each first video frame; based on all the inter-frame difference data, dynamic frame extraction is performed on the first IoT video stream data to obtain second IoT video stream data; the second IoT video stream data includes a plurality of key frames, the key frames being used to characterize the first video frame with the largest inter-frame difference data among all first video frames within the corresponding frame interval window. The server 802 is used to receive second IoT video stream data from the edge terminal; perform frame interpolation processing on the second IoT video stream data to obtain third IoT video stream data and send the third IoT video stream data to the user terminal.
[0056] It is worth mentioning that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0057] Figure 9 A schematic diagram of the structure of a computer device provided in this application embodiment includes: At least one processor 980; At least one memory 920 is used to store at least one program; When the at least one program is executed by the at least one processor 980, the at least one processor 980 performs the method as described in the foregoing embodiments.
[0058] This application also provides a computer-readable storage medium storing a processor-executable program, which, when executed by the processor 980, is used to implement the methods described in the foregoing embodiments.
[0059] Specifically, computer equipment can be either a user terminal or a server.
[0060] This application uses a computer device as a user terminal as an example, as detailed below: like Figure 9 As shown, the computer device 900 may include an RF (Radio Frequency) circuit 910, a memory 920 including one or more computer-readable storage media, an input unit 930, a display unit 940, a sensor 950, an audio circuit 960, a WiFi module 970, a processor 980 including one or more processing cores, and a power supply 990, among other components. Those skilled in the art will understand that... Figure 9The device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0061] The RF circuit 910 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and hands it over to one or more processors 980 for processing; additionally, it transmits uplink data to the base station. Typically, the RF circuit 910 includes, but is not limited to, an antenna, at least one amplifier, a tuner, one or more oscillators, a Subscriber Identity Module (SIM) card, a transceiver, a coupler, an LNA (Low Noise Amplifier), a duplexer, etc. Furthermore, the RF circuit 910 can also communicate wirelessly with networks and other devices. Wireless communication can use any communication standard or protocol, including but not limited to GSM (Global System for Mobile communication), GPRS (General Packet Radio Service), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), LTE (Long Term Evolution), email, SMS (Short Messaging Service), etc.
[0062] The memory 920 can be used to store software programs and modules. The processor 980 executes various functional applications and data processing by running the software programs and modules stored in the memory 920. The memory 920 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device 900 (such as audio data, telephone directory, etc.). In addition, the memory 920 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 920 may also include a memory controller to provide access to the memory 920 by the processor 980 and the input unit 930. Although Figure 9 The RF circuit 910 is shown, but it is understood that it is not a necessary component of the computer device 900 and can be omitted as needed without changing the nature of the invention.
[0063] The input unit 930 can be used to receive input digital or character information, and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control. Specifically, the input unit 930 may include a touch-sensitive surface 932 and other input devices 931. The touch-sensitive surface 932, also known as a touch display screen or touchpad, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch-sensitive surface 932), and drive the corresponding connection device according to a pre-set program. Optionally, the touch-sensitive surface 932 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to the processor 980, and can receive and execute commands from the processor 980. In addition, the touch-sensitive surface 932 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch-sensitive surface 932, the input unit 930 may also include other input devices 931. Specifically, other input devices 931 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.
[0064] Display unit 940 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of computer device 900. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. Display unit 940 may include display panel 941, optionally configured as LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), etc. Further, touch-sensitive surface 932 may cover display panel 941. When touch-sensitive surface 932 detects a touch operation on or near it, it transmits the information to processor 980 to determine the type of touch event. Subsequently, processor 980 provides corresponding visual output on display panel 941 according to the type of touch event. Although in Figure 9 In this embodiment, the touch-sensitive surface 932 and the display panel 941 are implemented as two separate components to realize input and output functions. However, in some embodiments, the touch-sensitive surface 932 and the display panel 941 can be integrated to realize input and output functions.
[0065] The computer device 900 may also include at least one sensor 950, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 941 according to the ambient light level, and the proximity sensor can turn off the display panel 941 and / or backlight when the computer device 900 is moved to the ear. As a type of motion sensor, a gravity acceleration sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometers, taps), etc. Other sensors that the computer device 900 may also be equipped with, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.
[0066] Audio circuitry 960, speaker 961, and microphone 962 provide an audio interface between the user and computer device 900. Audio circuitry 960 converts received audio data into electrical signals, which are then transmitted to speaker 961, where they are converted into sound signals for output. Conversely, microphone 962 converts collected sound signals into electrical signals, which are received by audio circuitry 960, converted back into audio data, and then processed by processor 980 before being transmitted via RF circuitry 910 to another control device, or output to memory 920 for further processing. Audio circuitry 960 may also include an earphone jack to facilitate communication between peripheral headphones and computer device 900.
[0067] Computer device 900 can transmit information with the wireless transmission module set up on the battle equipment via WiFi module 970.
[0068] The processor 980 is the control center of the computer device 900. It connects various parts of the control device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 920, and by calling data stored in the memory 920, it performs various functions of the computer device 900 and processes data, thereby providing overall monitoring of the control device. Optionally, the processor 980 may include one or more processing cores; optionally, the processor 980 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the aforementioned modem processor may also not be integrated into the processor 980.
[0069] The computer device 900 also includes a power supply 990 (such as a battery) that supplies power to various components. Preferably, the power supply can be logically connected to the processor 980 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 990 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0070] Although not shown, the computer device 900 may also include a camera, Bluetooth module, etc., which will not be described in detail here.
[0071] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the methods described in the foregoing embodiments.
[0072] This application also discloses a computer program product or computer program, which includes computer instructions stored in the aforementioned computer-readable storage medium; the processor of the aforementioned electronic device can read the computer instructions from the aforementioned computer-readable storage medium, and the processor executes the computer instructions, causing the electronic device to perform the aforementioned method embodiment.
[0073] It is understood that the content of the above method embodiments is applicable to this computer program product or computer program embodiment. The specific functions implemented by this computer program product or computer program embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0074] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatuses.
[0075] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0076] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0077] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0078] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0079] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0080] The step numbers in the above method embodiments are set only for ease of explanation and do not limit the order of the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0081] The above is a detailed description of the preferred embodiments of this application, but this application is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application, and these equivalent modifications or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A video transmission method for the Internet of Things, characterized in that, Applied to the edge, the method includes: Acquire first IoT video stream data; the first IoT video stream data includes several first video frames; Inter-frame content analysis is performed on the first IoT video stream data to obtain inter-frame difference data for each first video frame; Based on all the inter-frame difference data, the first IoT video stream data is dynamically extracted to obtain the second IoT video stream data; the second IoT video stream data includes several key frames, which are used to characterize the first video frame with the largest inter-frame difference data among all the first video frames within the corresponding frame interval window. The second IoT video stream data is sent to the server so that the server can perform frame interpolation on the second IoT video stream data. The server then obtains the third IoT video stream data and sends the third IoT video stream data to the user terminal.
2. The method according to claim 1, characterized in that, The step of performing inter-frame content analysis on the first IoT video stream data to obtain inter-frame difference data for each first video frame includes: Obtain edge operators; According to the edge operator, frame features are extracted from the second video frame and the third video frame to obtain the first frame feature of the second video frame and the second frame feature of the third video frame; the second video frame is any one of the first video frames in the first IoT video stream data; the third video frame is the first video frame adjacent to the second video frame in the first IoT video stream data; Based on the third video frame, perform inter-frame difference analysis on the second video frame to obtain inter-frame difference data of the second video frame.
3. The method according to claim 1, characterized in that, The step of dynamically extracting frames from the first IoT video stream data based on all the inter-frame difference data to obtain the second IoT video stream data includes: Obtain the original window frame group of the first IoT video stream data within the target interval window; the target interval window is any one of the frame interval windows. The original window frame group is compared with the inter-frame difference data to obtain the first target difference data. The first target difference data is used to characterize the largest inter-frame difference data among all the first video frames in the original window frame group. Based on the first target difference data, the original window frame group is filtered to obtain the key frame.
4. The method according to claim 3, characterized in that, The method further includes: Obtain the adaptive difference threshold; Based on the adaptive difference threshold, threshold analysis is performed on the original window frame group to obtain several second target difference data; the second target difference data is used to characterize the inter-frame difference data of all first video frames in the original window frame group that are greater than the adaptive difference threshold. Based on all the second target difference data, the original window frame group is frame filtered to obtain the value frame corresponding to each second target difference data. Based on the keyframes and all the value frames, the target window frame group is obtained.
5. The method according to claim 3 or 4, characterized in that, The method further includes: Obtain the target frame and several non-target frames within the target interval window, wherein the target frame is a value frame or the key frame; and the non-target frames are the first video frames within the target interval window other than the target frames. Based on all the non-target frames, frame motion features are extracted from the target frames to obtain motion feature information; Based on the motion feature information, the target frame is updated in terms of attributes to obtain the target frame with updated attributes.
6. A video transmission method for the Internet of Things, characterized in that, Applied to the server side, the method includes: Receive second IoT video stream data from the edge; The second IoT video stream data is subjected to frame interpolation processing to obtain the third IoT video stream data, and the third IoT video stream data is sent to the user terminal; The second IoT video stream data is obtained through the following steps: The edge device acquires first IoT video stream data; the first IoT video stream data includes several first video frames; The edge device performs inter-frame content analysis on the first IoT video stream data to obtain inter-frame difference data for each first video frame; The edge device dynamically extracts frames from the first IoT video stream data based on all the inter-frame difference data to obtain the second IoT video stream data; the second IoT video stream data includes several key frames, which are used to characterize the first video frame with the largest inter-frame difference data among all the first video frames within the corresponding frame interval window.
7. The method according to claim 6, characterized in that, The step of performing frame interpolation on the second IoT video stream data to obtain the third IoT video stream data includes: Obtain a first real frame and a second real frame. The first real frame is any video frame in the second IoT video stream data, and the second real frame is the video frame that is adjacent to the first real frame on the time axis among all video frames in the second IoT video stream data. Bidirectional optical flow analysis is performed on the first real frame and the second real frame to obtain bidirectional optical flow data; Based on the bidirectional optical flow data, the first real frame and the second real frame are frame fused to obtain an intermediate frame, which is used to represent the video frame supplemented between the first real frame and the second real frame.
8. An Internet of Things (IoT) video transmission system, characterized in that, Including edge computing and server computing; The edge device is used to acquire first IoT video stream data; the first IoT video stream data includes several first video frames; inter-frame content analysis is performed on the first IoT video stream data to obtain inter-frame difference data for each first video frame; based on all the inter-frame difference data, dynamic frame extraction is performed on the first IoT video stream data to obtain second IoT video stream data; the second IoT video stream data includes several key frames, the key frames being used to characterize the first video frame with the largest inter-frame difference data among all first video frames within the corresponding frame interval window. The server is used to receive second IoT video stream data from the edge terminal; The second IoT video stream data is subjected to frame interpolation processing to obtain the third IoT video stream data, and the third IoT video stream data is sent to the user terminal.
9. An electronic device, characterized in that, include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor performs the method as described in any one of claims 1-5 or the method as described in any one of claims 6-7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-5 or the method as described in any one of claims 6-7.