Adaptive uploading methods, systems, devices, and storage media for video image data

By extracting and combining keyframes from video image data through the communication gateway on the edge computing side, determining data attributes based on visual activity levels, and generating an adaptive upload strategy, the problem of redundant transmission in IoT video image data processing is solved, and data transmission with low power consumption and low bandwidth usage is achieved.

CN122137798APending Publication Date: 2026-06-02SHENZHEN QIANHAI BEE BEAT DIGITAL TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN QIANHAI BEE BEAT DIGITAL TECHNOLOGY CO LTD
Filing Date
2026-03-16
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies in IoT video image data processing fail to distinguish the differences in data value of video content, resulting in a large amount of redundant transmission in scenarios where the image is static for a long time or changes little. This increases the energy consumption and bandwidth resource usage of edge devices, making it difficult to meet the needs of battery-powered or low-power application scenarios.

Method used

The communication gateway on the edge computing side extracts and combines key frames of video image data to form multiple key segments. It also determines data attributes based on visual activity levels, generates an adaptive upload strategy, distinguishes between hot and cold data, optimizes upload methods and timing, and avoids redundant transmission.

Benefits of technology

It reduces the proportion of redundant data, alleviates bandwidth consumption and device power consumption, and is suitable for battery-powered or low-power application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122137798A_ABST
    Figure CN122137798A_ABST
Patent Text Reader

Abstract

This application relates to an adaptive uploading method, system, device, and storage medium for video image data. The method includes: acquiring video image data to be uploaded; extracting and combining keyframes from the video image data to obtain multiple key segments, each key segment consisting of multiple keyframes; determining the data attributes of each key segment based on its visual activity over time, including hot and cold data attributes; and generating an uploading strategy to control the uploading method and timing of the video image data based on the data attributes corresponding to each key segment. This method enables data segmentation, attribute classification, and scheduling control at the edge via a communication gateway, avoiding repeated transmission of static or minimally changing images, reducing redundant data, and is more suitable for battery-powered or low-power applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video image data processing technology, and in particular to an adaptive uploading method, system, computer device, and storage medium for video image data. Background Technology

[0002] In the field of video image data processing technology for the Internet of Things (IoT), the processing of acquired video image data at the edge computing side and its uploading to the backend system are involved to achieve remote monitoring and status analysis. Related video data uploading methods typically achieve video data transmission by continuously encoding the complete video stream and uploading it at fixed intervals or frequencies. However, this method fails to distinguish the differences in the data value of video content and is prone to generating a large amount of redundant transmission in scenarios where the image is static for a long time or with little change. This leads to increased power consumption and bandwidth resource usage of edge devices, making it difficult to meet the needs of battery-powered or low-power application scenarios. Summary of the Invention

[0003] Therefore, it is necessary to provide an adaptive uploading method, system, computer device, and computer-readable storage medium for video image data to address the aforementioned technical problems.

[0004] In a first aspect, this application provides an adaptive uploading method for video image data, applied to a communication gateway on the edge computing side, the method comprising: The video image data to be uploaded is obtained, and key frames are extracted and combined from the video image data to obtain multiple key segments, each of which is composed of multiple key frames. Based on the visual activity of each key segment in the time dimension, the data attributes of each key segment are determined to identify the corresponding data attributes, which include hot data attributes and cold data attributes. Based on the data attributes corresponding to each key segment, an upload strategy is generated to control the upload method and upload timing of the video image data.

[0005] Secondly, this application also provides an adaptive uploading system for video image data, applied to a communication gateway on the edge computing side, the system comprising: The acquisition module is used to acquire video image data to be uploaded, extract and combine key frames from the video image data to obtain multiple key segments, each of which is composed of multiple key frames. The attribute determination module is used to determine the data attributes of each key segment based on the visual activity of each key segment in the time dimension, so as to determine the data attributes corresponding to each key segment. The data attributes include hot data attributes and cold data attributes. The strategy generation module is used to generate an upload strategy for controlling the upload method and upload timing of the video image data based on the data attributes corresponding to each key segment.

[0006] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the above steps.

[0007] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the above steps.

[0008] The aforementioned adaptive uploading method, system, computer equipment, and computer-readable storage medium for video image data firstly involves local reception and key frame extraction of video image data via a communication gateway on the edge computing side, combining it into multiple key segments. This allows the data to undergo structural reorganization before entering the external network, thus forming controllable data units at the edge. Secondly, data attributes are determined based on a quantitative analysis of the visual activity level of each key segment, distinguishing between hot data attributes with significant image changes and cold data attributes with fewer changes. Thirdly, the uploading method and timing are arranged according to the differences in data attributes, enabling the adaptive uploading of key segments corresponding to hot and cold data attributes. Based on this, the entire technical solution completes data segmentation, attribute classification, and scheduling control at the edge through the communication gateway, avoiding repeated transmission of long-term static or less-changing images, reducing the proportion of redundant data, thereby reducing bandwidth consumption and controlling device power consumption, making it more suitable for battery-powered or low-power application scenarios. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a flowchart illustrating an adaptive uploading method for video image data in one embodiment; Figure 2 This is a block diagram of an adaptive video image data uploading system in one embodiment; Figure 3 This is an internal structural diagram of a computer device used to implement an adaptive uploading method for video image data in one embodiment; Figure 4This is an internal structural diagram of a computer-readable storage medium for an adaptive uploading method for video image data in one embodiment. Detailed Implementation

[0011] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0012] In one embodiment, such as Figure 1 As shown, an adaptive uploading method for video image data is provided. This embodiment illustrates the application of this method to a communication gateway on the edge computing side. In this embodiment, the method includes the following steps S101 to S103.

[0013] Step S101: Obtain the video image data to be uploaded, extract and combine key frames from the video image data to obtain multiple key segments, each of which consists of multiple key frames.

[0014] For example, a communication gateway located on the edge computing side between the video capture terminal and the remote platform receives the video image data to be uploaded. The communication gateway is an intermediate device deployed on the field side, possessing data access, caching, and external communication capabilities. It is used both to interface with the continuous video frame data output from the video capture terminal and to preprocess the data locally before sending it to the remote platform. After the communication gateway acquires the video image data, it parses each video frame in chronological order and compares them based on the degree of difference between adjacent frames. Essentially, it identifies video frames that represent phased changes in the scene by detecting changes in image content, structural changes, or brightness trends, and determines these video frames as keyframes.

[0015] Based on this, according to the distribution of keyframes on the timeline, multiple keyframes that are temporally adjacent and reflect the same continuous change process are classified and integrated to form a data unit with a complete time range, which is a key segment. Each key segment corresponds to a relatively independent change interval in the video, thereby transforming the video image data that was originally arranged frame by frame into a segmented data structure composed of multiple key segments.

[0016] Step S102: Based on the visual activity of each key segment in the time dimension, determine the data attributes of each key segment to identify the corresponding data attributes, which include hot data attributes and cold data attributes.

[0017] For example, after constructing multiple key segments, the changes in the image presented by each key segment within its coverage time range are analyzed and processed. That is, it is equivalent to comparing the key frames that make up a key segment in chronological order, so as to obtain a quantitative result describing the visual activity of the key segment in the time dimension by statistically analyzing the visual differences between adjacent key frames. Here, the visual activity level represents the comprehensive reflection of the frequency and magnitude of changes in the image content within a unit time range. The higher the value, the more frequent or obvious the changes in the image within that time range.

[0018] Furthermore, the visual activity level of each key segment is compared with a preset judgment standard. That is, the data attributes of each key segment are judged according to its position in different ranges. When the visual activity level of a key segment is higher than the set threshold, it is judged as a hot data attribute, while when the visual activity level is in a lower range and remains stable for a certain period of time, it is judged as a cold data attribute. Among them, hot data attributes are used to represent data types with obvious change characteristics that need to be given priority attention, while cold data attributes are used to represent data types with small changes and relatively stable images.

[0019] For example, in a corridor monitoring scenario, key segments of hot data attributes can represent video segments where people continuously enter and exit or walk quickly within a certain time period, and the main body position shifts significantly between key frames; key segments of cold data attributes can represent video segments where the background remains stable during periods when no one passes by, with only slight fluctuations in lighting or minor noise changes.

[0020] Step S103: Based on the data attributes corresponding to each key segment, generate an upload strategy to control the upload method and upload timing of video image data.

[0021] For example, the processing order and sending method of each key segment are differentiated based on the difference between hot data attributes and cold data attributes. For instance, for key segments determined to be hot data attributes, since the screen changes frequently and the content updates quickly within their corresponding time range, they are marked as priority processing objects when generating the upload strategy and directly included in the queue to be sent in chronological order so that they can quickly enter the upload process. For key segments determined to be cold data attributes, based on their slow screen changes and stable characteristics within a certain time range, their sending time is postponed, and multiple key segments with adjacent cold data attributes are sequentially integrated to form a continuous data set before being sent.

[0022] After completing the above differentiation and integration, the sending methods and sending order of each key segment are uniformly arranged to generate an upload strategy covering all key segments. This upload strategy clearly specifies the order in which each key segment is sent and whether it is sent in combination with other key segments, thereby achieving control over the overall video image data upload process at the level of method and timing.

[0023] In the aforementioned adaptive uploading method for video image data, in step S101, the video image data is locally received and key frames are extracted by the communication gateway on the edge computing side, and combined into multiple key segments. This allows the data to undergo structural reorganization before entering the external network, thereby forming controllable data units based on segments at the edge. In step S102, the data attributes of each key segment are determined based on a quantitative analysis of the visual activity level, thus distinguishing between hot data attributes with significant image changes and cold data attributes with fewer image changes. In step S103, the uploading method and uploading sequence are arranged according to the differences in data attributes, so that the key segments corresponding to hot data attributes and cold data attributes are uploaded adaptively. Based on this, in the entire technical solution, data segmentation, attribute classification, and scheduling control are completed at the edge through the communication gateway, avoiding repeated transmission of long-term static or less changing images, reducing the proportion of redundant data, thereby reducing bandwidth consumption and controlling device energy consumption, making it more suitable for battery-powered or low-power application scenarios.

[0024] In an exemplary embodiment, keyframes are extracted and combined from video image data to obtain multiple key segments, including steps S201 to S202.

[0025] Step S201: Based on the inter-frame correlation characteristics of each video frame in the video image data, key frames are extracted from the video image data and each key frame is marked in the original video bitstream to obtain a video bitstream containing key frame marks.

[0026] For example, each video frame in the video image data is analyzed frame by frame in chronological order. While keeping the original bitstream structure and timestamp information unchanged, the inter-frame correlation characteristics between adjacent video frames are calculated and analyzed. The inter-frame correlation characteristics are reflected in the magnitude of changes in image content, structural changes, and brightness change trends. In the specific processing, the degree of difference in pixel distribution between adjacent video frames is statistically analyzed to obtain a value reflecting the magnitude of changes in the image content. At the same time, the changes in the distribution of the main outlines or regions in the image are compared to determine whether the structure has changed significantly. Combined with the fluctuation trend of the overall brightness value in the time series, it is possible to identify whether there are phased changes in lighting conditions or shooting scene.

[0027] After obtaining the above comparison results, a comprehensive judgment is made. When adjacent video frames meet any of the preset judgment conditions in terms of the magnitude of image content change, structural change, and brightness change trend, such as when the difference in pixel distribution between two adjacent frames exceeds a preset difference threshold, or when the change in the main outline or area distribution of two adjacent frames in the picture shows a structural change exceeding a preset change threshold, or when the difference in the overall average brightness of two adjacent frames exceeds a preset brightness fluctuation threshold, then the two adjacent frames are considered representative at the current time position and are identified as keyframes. Subsequently, keyframe identifiers are embedded in the original video bitstream structure to enable accurate identification in subsequent processing, thereby obtaining a video bitstream containing keyframe identifiers.

[0028] Step S202: Score the consecutive keyframes in the video stream, and combine the consecutive keyframes whose scores meet the preset threshold conditions into a key segment.

[0029] For example, after obtaining the video stream containing keyframe identifiers, the frames identified as keyframes are extracted in chronological order, and adjacent keyframes are analyzed based on their arrangement on the timeline, i.e., multiple keyframes that appear consecutively in time are considered as a candidate keyframe set. Based on this, each candidate keyframe set is scored, with scoring criteria including the number of keyframes in the set, the correlation between the first and last keyframes, and the stability of changes between adjacent keyframes, to reflect the degree of concentration of changes and overall continuity of the candidate keyframe set within the corresponding time range.

[0030] Next, the scoring results of each candidate keyframe set at each level are compared with preset threshold conditions. If the scoring results at each level all reach the preset threshold conditions, it is equivalent to: the candidate keyframe set meeting or exceeding the preset minimum frame number requirement in terms of the number of keyframes to avoid insufficient information due to too few frames; meeting the preset time continuity or content connection requirements in terms of the correlation between the first and last keyframes to ensure that the set does not have obvious breaks on the timeline and maintains an internal connection between the first and last frames; and being within the preset fluctuation range in terms of the stability of changes between adjacent keyframes to ensure that the intensity of changes between keyframes remains relatively balanced and to prevent the internal consistency of the segment from being affected by drastic fluctuations in the degree of change. Therefore, the candidate keyframe set is determined to have the integrity of an independent time segment, and the keyframes in the set are integrated according to the original time order to form a key segment.

[0031] In this embodiment, in step S201, key frames are extracted by analyzing the inter-frame correlation characteristics between video frames and key frame identifiers are embedded in the original video bitstream, thereby achieving explicit marking of representative frames; in step S202, consecutive key frames in the video bitstream are scored and filtered in combination with preset threshold conditions to obtain each key segment; based on this, in the entire technical solution, the transformation from single-frame level recognition to segment level structure is realized, so that the video image data has clear segmentation boundaries at the structural level, providing a stable data organization foundation for subsequent further processing based on segments.

[0032] In an exemplary embodiment, data attributes of each key segment are determined based on the visual activity of each key segment in the time dimension, in order to determine the data attributes corresponding to each key segment, including steps S301 to S303.

[0033] Step S301: Analyze the temporal change characteristics of each key segment to dynamically determine the temporal analysis segment used for visual activity analysis of the corresponding key segment.

[0034] For example, the keyframes arranged chronologically within each key segment are analyzed as a whole, and their temporal change characteristics are analyzed based on their evolution along the timeline. These temporal change characteristics are specifically reflected in three aspects: change rhythm, distribution of change intensity, and change duration interval. During the analysis, the change rhythm is determined based on the order and frequency of appearance of the keyframes to determine whether the temporal change occurs densely in a short period of time or is evenly distributed over a longer period of time. Then, based on the magnitude of change corresponding to each keyframe, the distribution of change intensity within the entire time range of the key segment is identified to distinguish between segments with more prominent temporal changes and segments with relatively gentle changes. Furthermore, the change duration interval, i.e., the duration for which a certain change intensity is continuously maintained on the timeline, is used to determine whether the temporal change remains stable within a certain time range.

[0035] After considering the above factors, a continuous time interval is selected within the overall time range covered by the key segment. This interval exhibits a relatively concentrated rhythm of change, a high intensity of change within the numerical distribution of change intensity within the key segment, and consistently meets the preset duration requirement. This interval serves as the temporal analysis segment corresponding to the key segment. The temporal analysis segment is always confined within the time range of the corresponding key segment, and its duration is less than or equal to the overall time length of the key segment, without exceeding the start and end time boundaries of the key segment. Thus, by analyzing the temporal change characteristics within the key segment, a temporal analysis segment corresponding to the main change process is dynamically determined within its own time range for subsequent assessment of visual activity.

[0036] Step S302: Within the time-series analysis segment corresponding to each key segment, the visual activity level of the corresponding key segment is evaluated based on the visual difference relationship between key frames of the same key segment, so as to obtain the activity level representation results corresponding to the key segments.

[0037] For example, after the temporal analysis segment corresponding to each key segment is determined, the evaluation scope of each key segment is limited to the time range covered by the corresponding temporal analysis segment. The visual differences between key frames in the corresponding temporal analysis segment within each key segment are extracted and organized in chronological order. These visual differences originate from the differences between key frames at the image content level (e.g., the proportion of image area change, the degree of positional shift of the main subject, or the concentration of local changes). Subsequently, the difference values ​​between all adjacent key frames in the temporal analysis segment are statistically summarized in chronological order. That is, by accumulating, averaging, or normalizing the difference values ​​according to the time length, a quantitative result that can comprehensively characterize the frequency and magnitude of change in the temporal analysis segment is formed, namely, the activity level characterization result, which is used to describe the visual activity level of the key segment in the time dimension. Its value reflects the density and overall intensity level of changes between key frames in the temporal analysis segment.

[0038] For example, if there are 10 sets of difference values ​​between adjacent keyframes in a certain time series analysis segment, and the difference values ​​are 5, 8, 6, 7, 9, 10, 6, 8, 7, and 9 respectively, then the above difference values ​​are first accumulated to obtain 75, and then divided by the total duration of the keyframe interval or the number of keyframes to perform averaging, and the corresponding average difference value is obtained. This average difference value is used as the result of the activity level of the key segment in the time series analysis segment.

[0039] Step S303: Based on the activity level representation results corresponding to each key segment, determine the data attributes of each key segment to identify the corresponding data attributes.

[0040] For example, the activity level representation results corresponding to each key segment are compared against a set judgment interval. The judgment interval is used to divide different activity level values ​​into several levels. For instance, the interval with activity level values ​​above the upper limit benchmark value is defined as a high-activity interval, the interval below the lower limit benchmark value is defined as a low-activity interval, and the interval in the middle range is defined as a medium-activity interval. The upper and lower limit benchmark values ​​can be pre-calibrated and determined by combining conventional industry experience values, historical operational statistics, and production operation patterns. Based on this, the activity level representation results corresponding to each key segment are matched one by one with the aforementioned judgment interval. When the activity level value of a key segment is in the high-activity interval, the key segment is judged as a hot data attribute, indicating that its image changes frequently and with large fluctuations within the corresponding time range. When the activity level value is in the low-activity interval, it is judged as a cold data attribute, indicating that its image changes less and tends to be stable within the corresponding time range.

[0041] For cases within the moderately active range, the data is further merged into adjacent ranges based on the proximity of the values. For example, if the activity level is within the moderately active range but its value is higher than the median of that range and the difference between it and the lower limit of the high-activity range is less than a preset proximity threshold, it is merged into the high-activity range and identified as a hot data attribute. Conversely, if its value is lower than the median of that range and the difference between it and the upper limit of the low-activity range is less than a preset proximity threshold, it is merged into the low-activity range and identified as a cold data attribute.

[0042] In this embodiment, in step S301, the temporal change characteristics of each key segment are analyzed, and the temporal analysis segment is dynamically determined within the time range of the key segment, thereby limiting the subsequent evaluation to an effective time range that can represent the main change process; in step S302, the activity level characterization result is formed by statistically summarizing the visual differences between adjacent key frames within the temporal analysis segment, thereby representing the activity level of the key segment in the time dimension in numerical form; in step S303, the data attributes of each key segment are determined based on the activity level characterization result of each key segment, thereby realizing the attribute division of key segments with different change levels; based on this, in the entire technical solution, the refined identification and attribute division of the change characteristics of key segments are realized, providing a clear basis for subsequent differentiated processing based on data attributes.

[0043] In an exemplary embodiment, within the time-series analysis segment corresponding to each key segment, the visual activity level of the corresponding key segment is evaluated based on the visual difference relationship between each key frame of the same key segment, so as to obtain the activity level characterization result corresponding to the corresponding key segment, including steps S401 to S403.

[0044] Step S401: Within the temporal analysis segment corresponding to the current key segment, the visual differences in the temporal order of each key frame of the current key segment are compared frame by frame to obtain the difference values ​​between adjacent key frames.

[0045] For example, when assessing the visual activity of the current key segment, the assessment scope is limited to the temporal analysis segment corresponding to the current key segment. All keyframes within this temporal analysis segment are read sequentially according to time, and frame-by-frame comparison is performed using adjacent keyframes as units. During the comparison, difference values ​​are calculated based on changes in image content. Specifically, by statistically analyzing the differences between adjacent frames in aspects such as the proportion of change in the image area, the degree of subject position shift, or the size of the concentrated area of ​​local change, a difference value reflecting the intensity of change between the two frames is obtained. Based on this, a corresponding difference value is generated for each group of adjacent keyframes and arranged in chronological order, thereby constructing a difference value sequence covering the entire temporal analysis segment. This difference value sequence, indexed by time, records the continuous distribution of the intensity of change between keyframes, reflecting the changes of the current key segment at various time positions within this temporal analysis segment.

[0046] Step S402: Based on the distribution of each difference value in time sequence, the continuity constraint is applied to the changes between key frames of the current key segment to limit the scope of each difference value.

[0047] For example, by analyzing the relative magnitudes of the difference values ​​and the changing trends of adjacent difference values, a continuity constraint is imposed on the change process between keyframes to limit the effective range of each difference value in the statistical process. Specifically, when several adjacent difference values ​​remain relatively close in numerical range and appear consecutively in time sequence, for example, when the difference values ​​of three consecutive sets of adjacent keyframes are 12, 14, and 13, the corresponding consecutive positions are regarded as the same change stage and treated as a whole range in the subsequent statistical process. When the difference value shows a significant jump on the time axis, for example, suddenly rising from 13 to 35 or falling to 5, the position is regarded as the dividing point of the change stage, so that the difference values ​​before and after are respectively limited to their corresponding time intervals in the subsequent statistical process.

[0048] Step S403: Based on the difference values ​​after limiting the scope of action, comprehensively characterize the changes of the current key segment within the time-series analysis segment to generate the activity level characterization result corresponding to the current key segment.

[0049] For example, after defining the scope of each difference value, based on the boundaries of each change stage formed in the aforementioned steps, statistical processing is performed on each of the divided continuous change stages. For instance, in the example in the aforementioned steps, when the difference values ​​of three consecutive sets of adjacent keyframes are 12, 14, and 13, these three sets of values ​​have been defined as the same change stage. Therefore, the difference values ​​within this change stage are first accumulated to obtain 39, and then divided by the number of the three sets of adjacent keyframes included in this change stage to obtain an average value of 13, which is used as a characterization of the change intensity of this change stage. When the difference value suddenly rises from 13 to 35 or drops to 5, this position has been identified as the dividing point of the change stage. In the statistical process, the value intervals before and after the dividing point are processed independently. For example, the difference values ​​within the change stage before 13 are accumulated and calculated separately, and the difference values ​​within the change stage after 35 or 5 are accumulated and calculated separately, thereby avoiding the mixing of change stages with significant differences in statistics.

[0050] After obtaining the average value corresponding to each continuous change stage, the average value is weighted and summed based on the time proportion of each change stage within the time series analysis segment. For example, if a change stage covers 60% of the time series analysis segment and another change stage covers 40%, the average value is multiplied by the corresponding time proportion and then summed to obtain the final comprehensive value. This comprehensive value is used as the result representing the activity level of the current key segment within the time series analysis segment.

[0051] In this embodiment, in step S401, the difference values ​​between adjacent keyframes are obtained by comparing them frame by frame within the time-series analysis segment, thereby reflecting the data structure of the distribution of the intensity of change within the key segment over time; in step S402, the distribution of each difference value in the time dimension is constrained to limit the scope of statistical analysis of each difference value; in step S403, the changes of the key segment within the time-series analysis segment are comprehensively characterized based on the difference values ​​within the limited scope to generate a unified activity level characterization result; based on this, the structured and quantitative characterization of the change process of the key segment is achieved in the entire technical solution, so that the activity level value corresponds to the actual change stage.

[0052] In an exemplary embodiment, based on the activity level characterization results corresponding to each key segment, data attribute determination is performed on each key segment to determine the data attributes corresponding to each key segment, including steps S501 to S502.

[0053] Step S501: If the activity level representation result corresponding to the current key segment continues to reflect continuous changes between key frames within the corresponding time series analysis segment, then the data attribute corresponding to the current key segment is determined to be a hot data attribute.

[0054] Step S502: If the activity level representation result corresponding to the current key segment continues to reflect the situation of limited change between key frames in the corresponding time series analysis segment, then the data attribute corresponding to the current key segment is determined to be a cold data attribute.

[0055] For example, after obtaining the activity level representation result corresponding to the current key segment, the judgment is not based solely on the value falling into a certain judgment interval on a single occasion, but rather on a comprehensive judgment based on its stage distribution within the corresponding time series analysis segment. Specifically, the activity level representation result is derived from the statistical summary of multiple continuous change stages. Therefore, while generating a comprehensive value reflecting the overall activity level, the stage intensity and time proportion information corresponding to each continuous change stage are retained. On this basis, the comprehensive value of the current key segment is compared with the high-activity interval and the low-activity interval: on the one hand, when the average difference of each change stage constituting the comprehensive value... If the outlier values ​​all fall within the high-activity range and occupy the majority of the time within the time-series analysis segment, it indicates that the key segment continuously exhibits frequent changes between keyframes within the corresponding time range, and the intensity remains stable at a high level. For example, if three consecutive change stages are divided within a certain time-series analysis segment, and the average difference values ​​of each change stage are 14, 15, and 13, respectively, and all are within the high-activity range, and the time coverage of this segment exceeds 70%, then the overall value formed by the synthesis is determined to fall within the high-activity range, and its high-intensity change state exists continuously in the time dimension. Thus, the data attribute of the current key segment is determined to be a hot data attribute.

[0056] On the other hand, when the average difference values ​​of each change stage constituting the comprehensive value fall into the low-activity range and occupy the main time proportion within the time analysis segment, it indicates that the comprehensive value is not driven up by short-term high-intensity changes, but rather reflects the state of limited and small-amplitude changes between key frames in the time dimension. For example, if three consecutive change stages are divided within a certain time analysis segment, and the average difference values ​​of each change stage are 3, 5, and 4, respectively, and all are in the low-activity range, and the time proportion covering this segment exceeds 70%, then the comprehensive value formed by the synthesis is determined to fall into the low-activity range, and its stable change state exists continuously in the time dimension. Thus, the data attribute of the current key segment is determined to be a cold data attribute.

[0057] In this embodiment, in steps S501 and S502, based on the changes continuously reflected in the corresponding time-series analysis segments by the activity level characterization results corresponding to the key segments, key segments of hot data attributes that continuously exhibit high-intensity changes and key segments of cold data attributes that continuously exhibit stable changes are identified. Based on this, in the entire technical solution, the data attribute division reflects both the numerical magnitude and the continuous change state in the time dimension.

[0058] In an exemplary embodiment, an upload strategy for controlling the upload method and upload timing of video image data is generated based on the data attributes corresponding to each key segment, including steps S601 to S602.

[0059] Step S601: The upload method of the original data segment of each key segment with hot data attribute in the video image data is determined as the real-time upload method, and the upload method of the original data segment of each key segment with cold data attribute in the video image data is determined as the condition-triggered upload method.

[0060] For example, after the data attributes of each key segment are determined, the key segment is used as an index to locate its start and end time range in the original video image data, thereby determining the original data segment of each key segment in the original video image data. Based on this, according to the data attributes corresponding to the key segment, the upload method of its original data segment is set: on the one hand, when a key segment is determined to be a hot data attribute, it is considered that its corresponding original data segment contains continuously changing content in the time dimension. Therefore, the upload method of the original data segment is determined to be an instant upload method, so that the original data segment directly enters the sending process without delay after generation or caching. On the other hand, when a key segment is determined to be a cold data attribute, it is considered that its corresponding original data segment has limited changes and is generally stable. Therefore, the upload method of the original data segment is determined to be a condition-triggered upload method, that is, the sending process is performed only when preset conditions are met, such as after reaching a set time window or accumulating a certain number of original data segments with cold data attributes before unified sending.

[0061] Step S602: Under the constraint of the communication gateway's battery power, prioritize each raw data segment according to the upload method corresponding to each raw data segment to determine the upload sequence of each raw data segment and generate an upload strategy for video image data.

[0062] For example, after setting the upload method for the original data segments corresponding to each key segment, the original data segments are further sorted under the battery power constraint of the communication gateway to determine their upload sequence. Specifically, the remaining battery power value of the communication gateway is read, and the current state of high or low battery is determined according to the preset battery level range. For example, when the remaining battery power is higher than 50%, it is considered a high battery state, and when it is lower than 50%, it is considered a low battery state. Based on this, the original data segments that have been set to the instant upload method are given a higher sorting weight, so that they enter the sending queue first. For the original data segments that have been set to the conditional upload method, their arrangement position is adjusted according to the battery status. For example, when the battery is high, some original data segments with cold data attributes are allowed to be inserted after the original data segments that are uploaded instantly and sent sequentially. When the battery is low, all original data segments with cold data attributes are arranged uniformly at the end of the sending queue and sent uniformly only after a preset time interval is reached. Finally, an information structure containing the order of transmission of each original data segment is generated based on the sorted transmission queue. This information structure constitutes the final upload strategy, which guides the original video image data to be uploaded sequentially according to the predetermined order under the corresponding power conditions.

[0063] In this embodiment, in step S601, the data attributes of key segments are mapped to their corresponding original data segments, and the instant upload method and condition-triggered upload method are determined respectively, thereby achieving differentiated control of different original data segments at the transmission method level; in step S602, each original data segment is prioritized and an upload sequence is formed according to the current battery level of the communication gateway, so that the data sending order matches the battery level constraint; based on this, the entire technical solution achieves joint scheduling of video image data in two dimensions: upload method and upload sequence, so that the transmission process forms a coordinated relationship between content characteristics and device battery level.

[0064] In an exemplary embodiment, under the constraint of the communication gateway's battery power, the original data segments are prioritized according to their respective upload methods to determine the upload sequence of each original data segment and generate an upload strategy for video image data, including steps S701 to S702.

[0065] Step S701: Based on the battery power and consumption trend of the communication gateway, determine the priority of the original data segments corresponding to all key segments with hot data attributes in the real-time upload method, and the trigger constraint of the original data segments corresponding to all key segments with cold data attributes in the condition-triggered upload method.

[0066] For example, after determining whether the communication gateway is in a high-battery or low-battery state, the battery consumption trend is further analyzed by combining the changes in battery power over a continuous period of time. For example, when the battery power continuously decreases in a short period of time and the decrease exceeds a preset change benchmark, it is identified as a rapid consumption trend. When the battery power decreases less or remains relatively stable over a certain period of time, it is identified as a slow consumption trend. Based on the clear battery consumption trend, the priority of the original data segments corresponding to key segments with hot data attributes in the real-time upload mode is determined. For example, when it is in a high-battery state and in a slow consumption trend, it is given the highest priority so that it can enter the transmission queue without waiting. When it is in a low-battery state or in a rapid consumption trend, it still maintains a high priority, but the instantaneous power consumption is reduced by limiting the number of simultaneous transmissions or adjusting the transmission interval.

[0067] Meanwhile, for the original data segments corresponding to key segments with cold data attributes, the degree of trigger constraint for the original data segments corresponding to key segments with cold data attributes in the condition-triggered upload method is determined. For example, when the battery is high and in a slow consumption trend, it is allowed to be sent after the basic time interval condition is met (e.g., the preset minimum time interval has been reached since the last transmission). When the battery is low or in a rapid consumption trend, it can only be sent after a longer time interval is reached or more data is accumulated.

[0068] Step S702: Based on the determined combination of priority protection level and trigger constraint level, sort the priority of each raw data segment to determine the upload sequence of each raw data segment and generate the upload strategy for video image data.

[0069] For example, the priority level can be set as a binary parameter, where a value of 1 indicates a high priority level, meaning that the original data segment must be ranked first in the sorting process and is not subject to the sending interval restriction, and a value of 0 indicates a low priority level, meaning that although it can be sent first, it is subject to the concurrency limit or sending rhythm restriction. At the same time, the trigger constraint level can also be set as a binary parameter, where a value of 1 indicates a high trigger constraint level, meaning that it can only be sent after the extended time interval or cumulative quantity condition is met, and a value of 0 indicates a low trigger constraint level, meaning that it can be sent after the basic time interval condition is met.

[0070] For example, when the communication gateway is in a high-battery state and slowly consuming power, the original data segments corresponding to hot data attributes have a higher priority, while the original data segments corresponding to cold data attributes have a lower trigger constraint. Under this combination of priority and trigger constraint, all original data segments corresponding to hot data attributes are set to enter the transmission queue first without restriction on the transmission interval, while all original data segments corresponding to cold data attributes are set to enter the transmission queue after the basic time interval condition is met. When prioritizing under this combination, the original data segments of all hot data attributes are first arranged in chronological order, and then the original data segments of cold data attributes that meet the basic time interval condition are inserted sequentially after them.

[0071] For example, when the communication gateway is in a low-power state or in a rapid power consumption trend, the original data segments corresponding to hot data attributes have a lower priority, while the original data segments corresponding to cold data attributes have a higher trigger constraint. Under this combination of priority and trigger constraint, all original data segments corresponding to hot data attributes are set to be sent first and the number of concurrent transmissions is limited to control instantaneous power consumption. At the same time, all original data segments corresponding to cold data attributes are set to only enter the transmission queue after the extended time interval is reached or a larger amount of data is accumulated. When prioritizing under this combination, the original data segments of all hot data attributes are fixed at the front of the transmission queue, while the original data segments of all cold data attributes are uniformly delayed and added to the transmission queue in turn after the enhanced trigger conditions are met (such as the extended time interval has been reached since the last transmission or the number of accumulated original data segments of cold data attributes has reached a preset limit).

[0072] In this embodiment, in steps S701 and S702, the priority of the original data segments with hot data attributes and the trigger constraint of the original data segments with cold data attributes are determined according to the battery power and consumption trend of the communication gateway. Then, based on the combination of the determined priority and trigger constraint, all original data segments are prioritized, so that the upload strategy can be adaptively adjusted according to the battery power. Based on this, the entire technical solution realizes the hierarchical scheduling of video image data in an energy-constrained environment, so that the transmission order reflects the content change characteristics and meets the battery power constraints.

[0073] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0074] Based on the same inventive concept, this application also provides an adaptive uploading system for video image data to implement the adaptive uploading method for video image data described above. The solution provided by this system is similar to the implementation described in the above method; therefore, the specific limitations in one or more embodiments of the adaptive uploading system for video image data provided below can be found in the limitations of the adaptive uploading method for video image data described above, and will not be repeated here.

[0075] In one exemplary embodiment, such as Figure 2 As shown, an adaptive video image data uploading system is provided, applied to a communication gateway on the edge computing side. It includes: an acquisition module 201, an attribute determination module 202, and a policy generation module 203, wherein: The acquisition module 201 is used to acquire video image data to be uploaded, extract and combine key frames from the video image data to obtain multiple key segments, each of which is composed of multiple key frames. The attribute determination module 202 is used to determine the data attributes of each key segment based on the visual activity of each key segment in the time dimension, so as to determine the data attributes corresponding to each key segment. The data attributes include hot data attributes and cold data attributes. The strategy generation module 203 is used to generate an upload strategy for controlling the upload method and upload timing of video image data based on the data attributes corresponding to each key segment.

[0076] The modules in the aforementioned adaptive video image data uploading system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0077] In one exemplary embodiment, a computer device is provided, the computer device including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps in the above-described method embodiments.

[0078] This computer device serves as a communication gateway on the edge computing side, and its internal structure diagram can be seen as follows: Figure 3 As shown, the computer device includes a processor, memory, input / output interfaces, and a communication interface. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements the steps in the adaptive video image data uploading method described above.

[0079] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0080] In one exemplary embodiment, such as Figure 4 The diagram shows the internal structure of a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the above-described method embodiments.

[0081] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods.

[0082] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0083] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An adaptive uploading method for video image data, characterized in that, A communication gateway applied to the edge computing side, the method comprising: The video image data to be uploaded is obtained, and key frames are extracted and combined from the video image data to obtain multiple key segments, each of which is composed of multiple key frames. Based on the visual activity of each key segment in the time dimension, the data attributes of each key segment are determined to identify the corresponding data attributes, which include hot data attributes and cold data attributes. Based on the data attributes corresponding to each key segment, an upload strategy is generated to control the upload method and upload timing of the video image data.

2. The method according to claim 1, characterized in that, The process of extracting and combining keyframes from the video image data yields multiple key segments, including: Based on the inter-frame correlation characteristics of each video frame in the video image data, keyframes are extracted from the video image data and each keyframe is marked in the original video bitstream to obtain a video bitstream containing keyframe marks. The continuous keyframes in the video stream are scored, and the continuous keyframes whose scores meet the preset threshold conditions are combined into a key segment.

3. The method according to claim 1, characterized in that, The process of determining the data attributes of each key segment based on its visual activity over time, in order to identify the corresponding data attributes, includes: The temporal variation characteristics of each key segment are analyzed to dynamically determine the temporal analysis segments used for visual activity analysis of the corresponding key segments. Within the time-series analysis segments corresponding to each key segment, the visual activity level of the corresponding key segment is evaluated based on the visual difference relationship between each key frame of the same key segment, so as to obtain the activity level representation results corresponding to each key segment. Based on the activity level representation results corresponding to each key segment, the data attributes of each key segment are determined to identify the corresponding data attributes.

4. The method according to claim 3, characterized in that, Within the temporal analysis segments corresponding to each key segment, the visual activity level of the corresponding key segment is evaluated based on the visual difference relationship between key frames of the same key segment, so as to obtain the activity level representation results corresponding to the respective key segments, including: Within the time-series analysis segment corresponding to the current key segment, the visual differences in the temporal order of each key frame of the current key segment are compared frame by frame to obtain the difference values ​​between each adjacent key frame. Based on the distribution of each difference value in time sequence, the changes between each key frame of the current key segment are constrained to limit the scope of effect of each difference value. Based on the difference values ​​after limiting the scope of application, the changes of the current key segment within the time series analysis segment are comprehensively characterized to generate the activity level characterization result corresponding to the current key segment.

5. The method according to claim 3, characterized in that, The step of determining the data attributes of each key segment based on the activity level representation results of each key segment, and thus determining the corresponding data attributes of each key segment, includes: If the activity level representation result corresponding to the current key segment continues to reflect continuous changes between key frames within the corresponding time-series analysis segment, then the data attribute corresponding to the current key segment is determined to be a hot data attribute. If the activity level representation result corresponding to the current key segment continues to reflect a situation where changes between key frames are limited within the corresponding time-series analysis segment, then the data attribute corresponding to the current key segment is determined to be a cold data attribute.

6. The method according to claim 1, characterized in that, The step of generating an upload strategy based on the data attributes corresponding to each key segment to control the upload method and timing of the video image data includes: The upload method for each key segment with hot data attributes in the original data segment of the video image data is determined to be an instant upload method, and the upload method for each key segment with cold data attributes in the original data segment of the video image data is determined to be a condition-triggered upload method. Under the constraint of the communication gateway's battery power, the original data segments are prioritized according to their respective upload methods to determine the upload sequence of each original data segment and generate the upload strategy for the video image data.

7. The method according to claim 6, characterized in that, Under the constraint of the communication gateway's battery power, the process of prioritizing each raw data segment according to its corresponding upload method to determine the upload sequence of each raw data segment and generate the video image data upload strategy includes: Based on the battery power and consumption trend of the communication gateway, determine the priority of the original data segments corresponding to all key segments with hot data attributes in the real-time upload method, and the trigger constraint of the original data segments corresponding to all key segments with cold data attributes in the condition-triggered upload method. Based on the determined combination of priority protection and trigger constraint, the priorities of each raw data segment are sorted to determine the upload sequence of each raw data segment and generate the upload strategy for the video image data.

8. An adaptive uploading system for video image data, characterized in that, The system includes a communication gateway applied to the edge computing side. The acquisition module is used to acquire video image data to be uploaded, extract and combine key frames from the video image data to obtain multiple key segments, each of which is composed of multiple key frames. The attribute determination module is used to determine the data attributes of each key segment based on the visual activity of each key segment in the time dimension, so as to determine the data attributes corresponding to each key segment. The data attributes include hot data attributes and cold data attributes. The strategy generation module is used to generate an upload strategy for controlling the upload method and upload timing of the video image data based on the data attributes corresponding to each key segment.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.