Data transmission method, device and equipment
By evaluating the content and transmission dimensions of data units on the terminal device side, high-value data is identified and differentiated, solving the problems of resource waste and critical data delay in existing technologies, and achieving efficient data transmission and reliable backend AI analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
- Filing Date
- 2026-01-22
- Publication Date
- 2026-04-21
AI Technical Summary
Existing data transmission mechanisms cannot effectively distinguish between high-value and low-value data in network environments with limited or fluctuating bandwidth, leading to resource waste and delays or loss of critical data, which affects the continuity and reliability of backend AI analysis.
By jointly evaluating each data unit in the perception-related data stream from both the content and transmission dimensions at the terminal device side, data units with high business semantic value and high acquisition quality are identified, and priority scores are determined based on the evaluation results to achieve differentiated transmission scheduling.
With limited bandwidth, the system avoids the ineffective consumption of low-value data, significantly improves the success rate and timeliness of receiving key data, ensures the accuracy and real-time performance of backend AI analysis, and enhances the overall efficiency of the system.
Smart Images

Figure CN121907783A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of communication and data processing technology, and particularly to data transmission methods. This specification also relates to data transmission apparatus and a computing device. Background Technology
[0002] With the deep integration of artificial intelligence and multimodal perception technologies, real-time data transmission between front-end acquisition devices (such as cameras, microphones, and sensors) and back-end artificial intelligence (AI) analysis systems has become a critical factor affecting overall system performance and user experience. Currently, in scenarios such as real-time video analysis, remote monitoring, and interactive AI applications, how to efficiently and reliably transmit the acquired front-end perception data to the back-end server is a pressing technical problem that needs to be solved.
[0003] However, existing data transmission mechanisms generally employ static strategies. These schemes often make transmission decisions based solely on the arrival order of data packets or coarse type tags (such as audio and video stream tags), treating them equally. This "indiscriminate" transmission mode leads to significant resource misallocation: on the one hand, a large amount of low-quality or irrelevant data, which has no practical value for backend AI analysis, continuously occupies valuable network bandwidth, resulting in resource waste; on the other hand, in weak network environments with limited or fluctuating bandwidth, truly critical high-value data may be delayed or lost, causing backend AI analysis to be interrupted or fail, ultimately damaging the continuity and reliability of business functions. Summary of the Invention
[0004] In view of this, one or more embodiments of this specification provide a data transmission method, apparatus, and device to solve the network resource mismatch problem caused by the existing indiscriminate data transmission mode, thereby improving the continuity and reliability of the overall system's business functions.
[0005] According to a first aspect of one or more embodiments of this specification, a data transmission method is provided, comprising: Acquire the perception-related data stream to be transmitted; the perception-related data stream includes multiple data units; For each of the data units, an evaluation is performed based on at least one of the content dimension and the transmission dimension to obtain an evaluation result; the content dimension is related to at least one of the content importance and acquisition quality of the data unit; the transmission dimension is related to at least one of the dependency weight of the data unit and the network transmission cost under the current network condition. Based on the evaluation results, the priority score of each data unit is determined; According to the priority score, the transmission scheduling of each data unit is differentiated.
[0006] According to a second aspect of one or more embodiments of this specification, a data transmission apparatus is provided, comprising: The data acquisition module is used to acquire the perception-related data stream to be transmitted; the perception-related data stream includes multiple data units. The data evaluation module is used to evaluate each data unit based on at least one of the content dimension and the transmission dimension to obtain an evaluation result; the content dimension is related to at least one of the content importance and acquisition quality of the data unit; the transmission dimension is related to at least one of the dependency weight of the data unit and the network transmission cost under the current network condition. The priority score calculation module is used to determine the priority score of each data unit based on the evaluation results. The transmission scheduling module is used to perform differentiated transmission scheduling on each of the data units according to the priority score.
[0007] According to a third aspect of one or more embodiments of this specification, a computing device is provided, including a memory, a processor, and computer instructions stored in the memory and executable on the processor, wherein the processor, when executing the computer instructions, implements the steps of the data transmission method.
[0008] One embodiment of this specification can achieve at least the following beneficial effects: By having the terminal device quantitatively evaluate each data unit in the data stream to be transmitted based on at least one dimension, namely the content dimension (evaluating its content importance and acquisition quality) and the transmission dimension (evaluating its dependency weight and transmission cost under the current network conditions), the content importance, acquisition quality, and dependency weight of the data unit are identified, and the transmission cost of each data unit under the current actual network conditions is dynamically perceived. Then, based on the above multi-dimensional evaluation results, the priority score of each data unit can be comprehensively determined, and differentiated transmission scheduling is performed on each data unit according to the priority score. Thus, the on-demand and optimal allocation of network bandwidth resources can be realized, enabling the system to automatically and in real time prioritize the successful transmission of data units that are more critical to business implementation based on the priority score when network resources are limited. This effectively solves the problems of network resource waste and the inability to guarantee the arrival of critical data under weak network conditions, and significantly improves the analysis success rate, business reliability, and overall system efficiency of backend AI services. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is a schematic diagram illustrating an application scenario of a data transmission method provided in the embodiments of this specification. Figure 2 A schematic flowchart illustrating a data transmission method provided in an embodiment of this specification; Figure 3 This specification provides a schematic diagram of a data transmission process from a terminal device to a server in a practical application scenario, as illustrated in the embodiments of this specification. Figure 4 The embodiments provided in this specification correspond to Figure 2 A schematic diagram of the structure of a data transmission device; Figure 5 A structural block diagram of a computing device provided according to an embodiment of this specification is shown. Detailed Implementation
[0011] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0012] This specification uses specific terms to describe embodiments thereof. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of those different embodiments or examples, without contradiction.
[0013] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “an,” “an,” “the,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification includes any or all possible combinations of one or more associated listed items.
[0014] The terms “comprising,” “including,” or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitation, the presence of additional identical or equivalent elements in the process, method, product, or apparatus that includes said elements is not excluded.
[0015] Although the terms "first," "second," etc., may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, "first" may also be referred to as "second," and similarly, "second" may also be referred to as "first," without departing from the scope of one or more embodiments of this specification. Ordinal numbers such as "first," "second," etc., do not necessarily indicate order; often they are used to facilitate the distinction of objects. For example, "first server" and "second server" usually refer to two servers. To distinguish these two servers, they are described as "first server" and "second server." Of course, sometimes these two servers may be the same server.
[0016] Depending on the context, the word "if" as used here can be interpreted as "when," "when," or "in response to determination."
[0017] In this specification, unless explicitly stated otherwise, "receiving and sending data" does not necessarily mean direct receiving and sending; it can also mean indirect receiving and sending. For example, A receiving data sent by B can be understood as A directly receiving the data sent by B, or it can be understood as A indirectly receiving the data sent by B through other entities such as C. Similarly, B sending data to A can be understood as B sending the data directly to A, or it can be understood as B indirectly sending the data to A through other entities such as C. Here, C can be one entity, or it can be two or more entities.
[0018] In this specification, unless explicitly stated otherwise, the relationships between structures can be direct or indirect. For example, when describing "A is connected to B," unless it is explicitly stated that A and B are directly connected, it should be understood that A can be directly connected to B or indirectly connected to B. Similarly, when describing "A is on top of B," unless it is explicitly stated that A is directly above B (AB is adjacent and A is above B), it should be understood that A can be directly above B or indirectly above B (AB is separated by other elements, and A is above B). And so on.
[0019] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data shall comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation entry points shall be provided for users to choose to authorize or refuse.
[0020] With the deep integration of artificial intelligence and multimodal perception technologies, terminal devices (such as smartphones, smart locks, and industrial cameras) need to upload sensor information, including audio, video, and sensor data, to cloud-based AI systems for analysis in real time. However, existing data transmission mechanisms generally employ a static strategy—regardless of whether the data content contains valid business information (such as faces or abnormal events), it is compressed and transmitted at a fixed bitrate or frame rate. This "one-size-fits-all" approach exposes serious flaws in bandwidth-constrained weak network environments: on the one hand, a large amount of low-value data (such as background images in unmanned scenes and invalid audio clips) continuously consumes uplink bandwidth; on the other hand, key high-value data that truly affects AI analysis results (such as clear frontal face frames and specific sound events) is easily discarded or delayed due to a lack of priority protection, leading to backend recognition failures or delayed responses.
[0021] In light of this, the inventors of this application, after analysis, believe that the core deficiency of traditional solutions lies in the transmission system's lack of perception and quantitative assessment capabilities regarding the business semantic value carried by data units (e.g., whether a video frame contains key target events), acquisition quality (e.g., whether the picture is clear and the audio is pure), the dependency weight of data units, and the actual transmission cost of the current network. Therefore, under limited bandwidth, how to avoid wasting resources on low-value data while ensuring the priority and reliable delivery of high-value, high-recovery-cost data has become a key bottleneck in improving the overall efficiency of edge-cloud collaborative AI systems.
[0022] To address the shortcomings of related technologies, the embodiments in this specification achieve refined priority score decision-making by jointly evaluating each data unit in the perception-related data stream on the terminal device side using both content and transmission dimensions. On one hand, the content dimension evaluation identifies data units with high business semantic value (e.g., video frames containing faces) or high acquisition quality (e.g., clear focus and high signal-to-noise ratio). On the other hand, the transmission dimension evaluation quantifies the dependency weight of each data unit (e.g., the impact range on subsequent decoding as an I-frame, P-frame, or LTR frame) and its actual transmission cost under the current network conditions (e.g., including FEC redundancy and potential retransmission overhead). Based on the priority score determined by this dual evaluation, differentiated scheduling can prioritize high-priority data units to occupy the limited bandwidth within the sliding time window, while simultaneously applying stronger FEC and retransmission protection to high-priority data. Therefore, under weak network conditions, it avoids the ineffective consumption of bandwidth by low-value data and significantly improves the successful reception rate and timeliness of critical data, thereby ensuring the accuracy and real-time performance of backend AI analysis.
[0023] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0024] Figure 1 This is a schematic diagram illustrating an application scenario of a data transmission method provided in an embodiment of this specification.
[0025] like Figure 1 As shown, the terminal device 100 may be equipped with data acquisition devices. For example, the terminal device 100 may be equipped with auditory data acquisition devices (such as microphones) to acquire audio data, visual data acquisition devices (such as cameras) to acquire image or video data, and sensors (such as speedometers, gyroscopes, temperature sensors, etc.) to acquire sensor data, and is not limited to these.
[0026] The terminal device 100 may also include a data processing unit. For example, the data processing unit can perform preliminary analysis on the raw sensing data to obtain derived sensing semantic data (such as posture data, motion trajectory data, point cloud data, etc.).
[0027] After acquiring the perception-related data stream (including original perception data and / or derived perception semantic data) to be transmitted, the terminal device 100 can evaluate each data unit based on at least one dimension, namely the content dimension and the transmission dimension, to obtain an evaluation result. Then, based on the evaluation result, a priority score is determined for each data unit. Thus, differentiated transmission scheduling is performed on each data unit according to the priority score. Based on this differentiated transmission scheduling, the perception-related data stream can be transmitted from the terminal device 100 to the server 200.
[0028] In such Figure 1 In the application scenario shown, server 200 can connect to one or more terminal devices 100 via LAN connection, WAN connection, Internet connection or other types of data network. Figure 1 The server 200 in the text can include, but is not limited to, any device, equipment, platform, or equipment cluster with computing and processing capabilities. Figure 1 The terminal device 100 can be any physical device, including but not limited to: smartphones, computers, network devices, smart home devices, wearable devices, and smart medical devices. Computers can include, but are not limited to, PCs, laptops, and tablets. Network devices can include, but are not limited to, routers, switches, network cards, and hubs. Smart home devices can include, but are not limited to, smart TVs, smart air conditioners, smart humidifiers, smart water heaters, smart kitchen appliances, smart doors and windows, and smart air purifiers. Wearable devices can include, but are not limited to, smart bracelets, smartwatches, and smart glasses. Smart medical devices can include, but are not limited to, smart blood pressure monitors, smart scales, smart blood glucose meters, and smart massage chairs. The authentication device can be a server-side device or a cluster of devices, such as a server or a server cluster.
[0029] In practical applications, Figure 1 The terminal device 100 can also be referred to as the terminal side, front end, front-end device, edge device / end / side, data transmission device / end / side, data acquisition device / end / side, etc. The server 200 can also be referred to as the back end, back end system, cloud, cloud system, data receiving device / end / side, data analysis device / end / side, etc., specifically such as cloud business system, back end AI analysis system, etc.
[0030] like Figure 1 As shown in the embodiments of this application, rapid business relevance analysis is performed at the edge before the data reaches the backend data analysis system. The data transmission system thus realized has business front-end capability, which can reduce the transmission of useless / low-value data, save network resources, and improve network resource utilization efficiency on the one hand, and ensure the transmission of useful / high-value data on the other hand, improve the parsing success rate of the backend analysis system, thereby improving business reliability.
[0031] This specification provides a technical solution for intelligent and differentiated transmission scheduling in bandwidth-constrained or unstable network environments. This solution involves content understanding and dynamic assessment of transmission costs for perceived data (such as video, audio, and sensor data). This solution can be widely applied to scenarios with high requirements for data transmission quality and timeliness, such as real-time video recording, video conferencing, remote control, and edge AI analysis.
[0032] The embodiments in this specification provide a data transmission method. This application also relates to a data transmission device and a computing device, which will be described in detail in the following embodiments.
[0033] Figure 2 This is a flowchart illustrating a data transmission method provided in an embodiment of this specification.
[0034] From a programming perspective, the entity executing the process can be a program mounted on a terminal device. It can be understood that this method can be executed by any device, equipment, platform, or cluster of devices with computing and processing capabilities.
[0035] like Figure 2 As shown, the process may include the following steps: Step 202: Obtain the perception-related data stream to be transmitted; the perception-related data stream includes multiple data units.
[0036] This specification provides a data transmission method for multimodal services in its embodiments. The perception-related data stream to be transmitted may include multimodal data. Specifically, it may include various data collected by different types of sensing devices, including but not limited to video, audio, images, sensor information, etc., and this multimodal data typically needs to be combined and analyzed in a backend AI system.
[0037] The data transmission method can be applied to terminal devices. The perception-related data stream may include raw perception data collected by the terminal device and perception semantic data generated by an edge AI model deployed on the terminal device. Optionally, the perception-related data stream may include raw perception data of the physical world (such as camera visual perception data, microphone auditory perception data, sensor environmental perception data, etc.). Optionally, the perception-related data stream may include derived perception semantic data (such as posture data, motion trajectory data, point cloud data, etc.) obtained after the terminal device performs preliminary analysis of the raw perception data.
[0038] Furthermore, the perception-related data stream can represent a structured data sequence with time attributes generated from the acquisition unit or processing unit deployed on the terminal device. In practical applications, the perception-related data stream may include, but is not limited to: visual perception data, such as video frame sequences and image sequences; auditory perception data, such as audio frame sequences and acoustic feature sequences; sensor time-series data, such as data sampled according to sampling periods generated by accelerometers, gyroscopes, temperature sensors, etc.; semantic parsing data, such as structured data sequences obtained by parsing the original perception data through AI models, including target detection box sequences, speech-to-text streams, behavior recognition label sequences, etc.; metadata and control data, such as device status information and acquisition parameters.
[0039] The data units in the perception-related data stream include at least one of the following: video frames, audio frames, accelerometer data, gyroscope data, infrared distance sensor readings, speech recognition text, human posture trajectory data, abnormal event tags, etc.
[0040] Step 204: For each of the data units, evaluate them based on at least one of the content dimension and the transmission dimension to obtain the evaluation result; the content dimension is related to at least one of the content importance and acquisition quality of the data unit; the transmission dimension is related to at least one of the dependency weight of the data unit and the network transmission cost under the current network state.
[0041] Optionally, for each data unit, evaluation can be performed based on both the content dimension and the transmission dimension to obtain a content dimension score and a transmission dimension score, respectively; and at least one of the content evaluation parameters and the transmission evaluation parameters can be used to determine the priority score.
[0042] The content dimension can be an analytical dimension used to evaluate the value or quality of the data to be transmitted. Specifically, the content dimension can be used to evaluate two aspects: the importance of the business content carried by the data unit and the quality of data collection. Optionally, only one aspect can be evaluated.
[0043] The transmission dimension can be an analytical dimension used to evaluate the dependencies of the data to be transmitted, network bandwidth usage, and recovery costs. Specifically, the transmission dimension can be used to evaluate two aspects: the dependency weight of data units and the network transmission cost (e.g., transmission resource consumption) under the current network conditions. Optionally, only one aspect can be evaluated.
[0044] The Network Transmission Cost can represent the overall overhead of data transmission over a network, and can be related to multiple aspects such as data size, encoding complexity, bandwidth usage, and retransmission delay.
[0045] In one or more embodiments of this specification, in order to achieve business-driven data transmission optimization, the system (terminal device) first performs a business-pre-analysis of the data units in the perception-related data stream before data transmission, evaluating the actual value of each data unit to the current business. This analysis process is completed in real time on the terminal side, ensuring that data transmission decisions can reflect business needs in a timely manner.
[0046] The evaluation of the data unit based on the content dimension in step 204 may specifically include evaluating the content importance of the data unit.
[0047] Evaluating the data unit based on the content dimension may specifically include: determining the relevance of the data unit to the current business scenario; and generating a content importance score based on the relevance.
[0048] The acquired data units can first undergo lightweight business semantic analysis, and then the results of the business semantic analysis can be mapped to a content importance score. The content importance score reflects the content importance of the data unit.
[0049] Furthermore, determining the relevance of the data unit to the current business scenario may specifically include: using a local artificial intelligence model to detect whether the data unit involves a predefined target object or event type. In the embodiments of this specification, data units involving predefined target objects or event types can obtain a relatively high content importance score.
[0050] In practical applications, AI models deployed on terminal devices or specific scenarios, events, or objects of data to be transmitted can be identified and analyzed according to business rules to determine their impact on business processes.
[0051] More specifically, a business semantic awareness model deployed on the terminal side can be used to detect whether the data unit contains predefined key target objects or event types related to the current business scenario. Furthermore, the business semantic awareness model can be configured according to different business scenarios to identify specific semantic content related to that business scenario.
[0052] In practical applications, the system (terminal device) can adopt a modular business semantic analysis framework, which may include: a business scenario configuration module, configured to load corresponding predefined business information (such as predefined target objects or event types) according to the current application scenario (such as identity verification, security inspection, interactive dialogue, industrial inspection, etc.); a lightweight business semantic perception model, which can be configured as a neural network model optimized for different data types, capable of running efficiently on the terminal device to achieve business semantic analysis; and a scoring mapping module, configured to convert the analysis results into a unified content importance score.
[0053] Furthermore, the system (terminal device) can employ differentiated analysis methods for different types of data units.
[0054] Optionally, for video data units, lightweight visual AI models deployed at the edge can be used for real-time object detection or event recognition. Object detection includes identifying specific objects, such as faces, vehicles, and specific landmarks. Event recognition can include detecting predefined event patterns, such as traffic accidents, crowd gatherings, and abnormal behavior. For example, in identity verification scenarios, video frames detecting clear faces can obtain high content importance scores. Similarly, in security patrols, video frames detecting unauthorized intrusions or left-behind items can obtain high content importance scores. Furthermore, in live interactive scenarios, video frames detecting specific interactive gestures or facial expressions can obtain high content importance scores.
[0055] Optionally, for audio data units, specific sound events, such as gunshots, explosions, cries for help, and the sound of breaking glass, can be identified using audio event detection models deployed at the edge. Alternatively, specific voice keywords, such as specific command words, wake-up words, or sensitive words, can also be identified. These models can be deep learning-based acoustic event detection models or traditional signal processing and pattern recognition methods. For example, in security scenarios, audio segments detecting cries for help or abnormal sounds can receive high content importance scores. Similarly, in voice interaction scenarios, audio frames containing clear voice commands can receive high content importance scores.
[0056] Optionally, for sensor data units, such as infrared sensors, accelerometers, and gyroscopes, business semantic analysis can be performed based on preset rules or simple models. Specifically, whether a critical event has occurred can be determined based on whether its data value exceeds a preset threshold (safety or anomaly threshold) or conforms to a specific pattern (such as a specific motion trajectory, equipment vibration pattern, etc.). For example, in industrial inspection, an abnormally high temperature sensor reading may indicate equipment failure, and the importance score of the data unit generated at that time will increase. Similarly, in health monitoring, abnormal heart rate sensor data may indicate a user health problem, and the importance score of the data unit generated at that time will increase.
[0057] Optionally, business model parsing data units typically contain structured information that has undergone preliminary analysis, such as human pose sequences extracted from videos or text converted from audio. For these business model parsing data units, a score can be awarded based on whether the information they contain meets preset quality assessment criteria. These criteria may include, but are not limited to, confidence levels (i.e., the confidence level of the model output), key fields (checking whether the data contains predefined key fields or markers), and completeness. For example, in a digital human scenario, the completeness of the pose data (whether it contains full-body key points) and confidence level can affect the score. Similarly, in a voice interaction scenario, whether the recognized text contains key instructions or sensitive words can affect the score.
[0058] It's important to note that the business semantic awareness model used to detect predefined target objects or event types can be customized according to business scenarios, supporting different domains (such as security, live streaming, and industrial inspection). Different business scenarios identify different target objects or event types, therefore the logic for generating content importance scores also differs. For example, in a security scenario, identifying an intruder will result in a high score; while in a live streaming scenario, identifying viewer donations or interactions may lead to a higher score.
[0059] Furthermore, the content importance score can employ a unified quantification system. In an optional embodiment, the content importance score can be quantified as a value from 0 to 100, with higher scores indicating greater importance. The score can be binary (0 or 100), multi-level (e.g., 0, 50, 100), or even a continuous value. The specific mapping rules for the score can be determined by business requirements; for example, a score of 100 might be assigned if a key event is detected, and 0 if it is not detected; or it can be calculated comprehensively based on indicators such as the confidence level, quantity, and size of detected targets.
[0060] Based on the embodiments in this specification, by using content importance scores as input for subsequent priority ranking, the priority ranking results used for data transmission scheduling can refer to the business importance carried by the data unit, rather than simply based on the data unit's acquisition / acquisition time or size. This enables the data transmission system to understand the business scenario and ensures the priority transmission of critical business data during the data transmission phase. For example, when bandwidth is limited, video frames containing critical events (such as accidents) will be transmitted first, while video frames not containing critical events may be discarded or delayed, thereby ensuring that the backend AI analysis system can process critical events in a timely manner and improve the overall system's business response efficiency.
[0061] For example, in identity verification scenarios, when a user performs facial recognition, the solution based on this application embodiment can prioritize transmitting frames containing clear faces, while blurry or faceless frames may be discarded or downgraded for transmission, thereby reducing a large amount of invalid data transmission and improving the success rate and speed of backend recognition. As another example, in security patrol scenarios, the solution based on this application embodiment can transmit at a low frame rate under normal conditions, and immediately switch to high-priority, high-quality transmission once intrusion is detected, thereby significantly reducing daily bandwidth consumption and critical event transmission latency. Furthermore, in interactive dialogue scenarios, in voice assistant interactions, audio segments containing clear voice commands can be prioritized for transmission, filtering out irrelevant sounds, thereby significantly reducing audio data transmission, lowering cloud processing load, and improving response speed.
[0062] Furthermore, since the content importance score is calculated in real time on the terminal, this reduces reliance on cloud analytics services, lowers transmission latency, and can adapt to unstable network conditions. Moreover, the terminal-side calculation model can be updated based on models distributed from the cloud to adapt to new business scenarios or optimize recognition performance.
[0063] Based on the embodiments described in this specification, business-driven content importance assessment is implemented, providing key inputs for subsequent priority score decisions, thereby improving the intelligence and business value of data transmission as a whole.
[0064] Based on the embodiments described in this specification, a business semantic pre-analysis mechanism based on a lightweight AI model is introduced on the terminal side, achieving deep coupling between data transmission strategy and business value. Specifically, for various data units such as video, audio, and sensors, customized business semantic perception models (such as target detection, event recognition, keyword wake-up, etc.) are run in real time to accurately identify predefined targets or events related to the current scene and generate a quantified content importance score accordingly. In this way, the system can dynamically distinguish the "business value" of data units—for example, marking frames containing intrusion behavior in security scenarios, filtering frames containing clear faces in identity verification, and extracting audio segments containing clear commands in voice interaction. This process is completed entirely on the terminal side, which not only reduces reliance on cloud analysis and transmission latency but also influences the calculation of priority scores, thereby enabling differentiated transmission scheduling for each data unit. Thus, when bandwidth is limited, the transmission of high-value data is proactively prioritized, while low-value data is suppressed or discarded. Therefore, based on the embodiments described in this specification, a shift from "indiscriminate transmission" to "business value-driven transmission" is achieved. While reducing cloud load, it ensures that critical business events can be reported and processed with low latency and high reliability, which is conducive to improving the overall intelligence, response efficiency and resource utilization of the system.
[0065] In one or more embodiments of this specification, the system (terminal device) evaluates the acquisition quality of the data unit by combining the real-time status information of the acquisition device, thereby avoiding the transmission of low-quality, unidentifiable data to the cloud and saving bandwidth at the source.
[0066] The evaluation of the data unit based on the content dimension in step 204 may specifically include evaluating the acquisition quality of the data unit.
[0067] Specifically, evaluating the data unit based on the content dimension may include: acquiring real-time status information of the acquisition device that generated the data unit; the real-time status information reflecting the acquisition quality of the data unit by the acquisition device; and generating a content quality score based on the real-time status information. The content quality score can also reflect the acquisition quality of the data unit. For example, higher acquisition quality can result in a higher content quality score.
[0068] The real-time status information reflects the real-time status of the device acquiring the data unit when acquiring it. In practical applications, the real-time status information can be used as metadata for the data unit, acquired, stored, and used by the terminal.
[0069] The real-time status information may include at least one of the following: camera status information, microphone status information, and device self-diagnostic information.
[0070] Furthermore, the camera status information may include at least one of the following: focus distance, aperture value, exposure parameters, image sharpness detection result, and camera anomaly marker.
[0071] Focusing distance, or focusing distance, refers to the object distance at which the camera focuses. In practical applications, the camera's focusing state directly affects image clarity; focusing too close or too far can lead to blurry images. Typically, different business scenarios have corresponding optimal focusing distance ranges. For example, in identity verification scenarios, the optimal focusing distance is the distance between the face and the camera (e.g., 30cm-100cm). If the user is too close (<30cm) or too far (>100cm) from the camera, the system (terminal device) can lower the quality score of that frame and prompt the user to adjust the distance. In industrial inspection scenarios, the optimal distance between the camera and the product is a fixed value; a deviation exceeding ±5% is considered low quality. Furthermore, in telemedicine scenarios, wound or skin detection has strict focal length requirements; images deviating from the standard focal length have significantly lower quality scores.
[0072] Aperture value affects the amount of light entering the camera and depth of field. An inappropriate aperture value can lead to images that are too dark or too bright, or images where the subject is not sharp. For example, if the aperture is too large (F-number too small), the amount of light entering the camera is large and suitable for low-light environments, but in some business scenarios, the depth of field may be too shallow, and the subject may be out of focus. For example, if the aperture value is lower than F2.0, the score for inspection scenarios that require sharpness across the entire scene can be lowered. On the other hand, if the aperture is too small (F-number too large), the depth of field is deep and the entire scene is sharp, but in some business scenarios, the amount of light entering the camera may be insufficient, potentially leading to increased noise. For example, if the aperture value is higher than F11, the score for low-light environments can be lowered.
[0073] Exposure parameters can include exposure time and ISO. Improper exposure can lead to overexposure or underexposure of the image, affecting the brightness and detail of the image.
[0074] For image sharpness information, image processing algorithms can be used to quickly detect the sharpness of video frames (including edge sharpness analysis and / or noise detection) to obtain the image sharpness detection results.
[0075] For camera anomaly marking, image data can be analyzed to detect whether the camera screen is distorted, whether the frame rate is much lower than the set value, etc., and marking will be carried out when anomalies are detected.
[0076] Therefore, for video frame-type data units, the acquisition quality of the video frame can be evaluated based on at least one of the following: camera focus distance, aperture value, exposure parameters, image sharpness information, and camera anomaly markers. Specifically, this can include: determining whether the focus distance is within the acceptable focus distance range for the current business scenario; if not, lowering the content quality score; determining whether the aperture value is within the reasonable working range; if not, lowering the content quality score; determining whether the exposure parameters indicate an exposure anomaly (including overexposure and / or underexposure); if an exposure anomaly exists, lowering the content quality score; determining whether the image sharpness information meets a preset image sharpness threshold; if not, lowering the content quality score; and if a camera anomaly marker is detected, the content quality score can be set to the lowest value.
[0077] Furthermore, the microphone status information may include at least one of the following: background noise level, signal-to-noise ratio, and microphone anomaly marker.
[0078] Background noise levels can be measured in decibels; excessively high background noise can interfere with speech signals. In practical applications, terminal devices can have built-in voiceprint feature libraries for common noise types to facilitate the identification of steady-state noise (such as air conditioner noise, fan noise, and traffic background noise), transient noise (such as keyboard typing, door closing, and coughing), and impulse noise (such as clapping and impact sounds). Signal-to-noise ratio (SNR) represents the ratio of signal to noise; a higher SNR indicates better audio quality. Microphone anomaly markers can detect microphone malfunctions, such as fixed output values, zero output, or excessive noise.
[0079] Therefore, for audio frame-type data units, the acquisition quality of the audio frame can be evaluated based on at least one of the microphone background noise level, signal-to-noise ratio (SNR), and microphone anomaly marker. Specifically, this may include: reducing the content quality score if the detected background noise level exceeds a preset noise level threshold; reducing the content quality score if the calculated SNR is lower than a preset SNR threshold; and setting the content quality score to the lowest value if a microphone anomaly marker is detected.
[0080] Furthermore, the device self-diagnostic information may include at least one of the following: device temperature, vibration amplitude, and processor frequency reduction flag.
[0081] Device temperature can include battery temperature and CPU temperature, etc. Excessive temperature may cause the device to throttle, affecting data acquisition performance. Vibration amplitude can be obtained using accelerometers and gyroscopes; severe vibration may cause unstable data acquisition. Regarding the processor throttling indicator, when the acquisition device overheats or has insufficient power, the processor may operate at a reduced frequency, which may degrade the quality of the acquired data.
[0082] Therefore, the quality of the acquired data units can also be evaluated based on at least one of the acquisition device's temperature and vibration amplitude. Specifically, this can include: if the acquisition device temperature exceeds a preset temperature threshold, it is determined that the acquisition device may be in a frequency reduction state, and the content quality score is lowered; if abnormal vibration amplitude is detected, it is determined that the acquisition device is in an unstable state, and the content quality score is lowered.
[0083] In the embodiments of this specification, the operating status of the acquisition device (such as camera focus, exposure information, microphone signal-to-noise ratio, etc.) can be used to evaluate the quality of the data content.
[0084] In practical applications, different quality thresholds can be set for different business scenarios, including image sharpness thresholds, noise level thresholds, signal-to-noise ratio thresholds, and temperature thresholds. Optionally, different quality thresholds can be set for different models of acquisition devices.
[0085] Furthermore, after obtaining real-time status information, the system (terminal device) can map this status information into a content quality score, ranging from 0 to 100, with higher scores indicating better data collection quality. The content quality score can be generated using rule-based methods or machine learning model-based methods.
[0086] Optionally, the method for generating content quality scores based on rules can include: setting a scoring function for each state parameter, and then weighted and fused the scores of multiple parameters. Taking a camera as an example, a focus distance score can be calculated for the focus distance, an aperture score for the aperture value, an exposure score for the exposure parameters, a sharpness score for the image sharpness detection results, and an anomaly score for the camera anomaly markers. For example, a focus distance score can be calculated based on the deviation between the current focus distance and the optimal focus distance; the larger the deviation, the lower the focus distance score. An aperture score can be calculated based on whether the aperture value is within a reasonable range. An exposure score can be calculated based on the overall brightness or histogram distribution of the image. Based on an edge detection algorithm, the gradient magnitude of the image can be calculated; the higher the gradient magnitude, the higher the sharpness score. Anomaly marker score: if marked as an anomaly, the score is directly 0. Then, the above scores can be weighted and summed to obtain the overall score of the camera state. Similarly, similar rules can be used for microphone state and device self-diagnosis state.
[0087] Alternatively, the system can train a machine learning model and deploy it on the terminal, taking the various state information as input and having the model directly output a content quality score. For example, a neural network model can be used, trained on a large amount of labeled data, to learn the relationship between state information and data quality.
[0088] Based on the embodiments described in this specification, by introducing content quality scoring, low-quality data can be identified and discarded or have its priority reduced before transmission, filtering out low-quality data and reducing invalid transmissions. Furthermore, bandwidth is saved by avoiding wasting it on junk data unusable for business analysis. Additionally, it ensures high-quality data transmitted to the backend, which is beneficial for improving the recognition accuracy of AI models. For example, in an identity verification scenario, if the camera is out of focus, causing a blurry face, the system (terminal device) will assign a low-quality score, thereby reducing the priority score of that frame or even discarding it. Simultaneously, the system (terminal device) can prompt the user to adjust their position or refocus. As another example, in an audio interaction scenario, if the ambient noise is too high, the system (terminal device) will reduce the priority score of the current audio frame to prevent transmission noise from interfering with backend speech recognition. In this way, the transmission system can guarantee data quality from the source and optimize overall transmission efficiency.
[0089] Optionally, the system (terminal device) can provide real-time feedback based on the quality score to prompt the user to adjust their acquisition behavior to obtain higher-quality content. For example, when a focus anomaly is detected, the user can be prompted to "move the device closer / away." Similarly, when an exposure anomaly occurs, the camera's exposure parameters can be automatically adjusted. Furthermore, when sensor anomalies are continuously detected, the user can be reminded to check the device. Thus, real-time feedback can also improve data acquisition quality, avoid generating and transmitting junk data, and ultimately optimize overall data transmission efficiency.
[0090] In the embodiments of this specification, for non-collected data, such as semantic parsing data, the collection quality dimension can be set to a default value. For example, when the data unit is non-collected data, the content quality score or corresponding weight can be set to a low value (e.g., set to zero).
[0091] Based on the embodiments described in this specification, by introducing a content quality assessment mechanism based on the real-time status of the acquisition device, data quality can be effectively guaranteed and transmission resources optimized from the data source. Specifically, the terminal side acquires real-time status information of the camera (focus distance, aperture value, exposure parameters, image sharpness, etc.), microphone (background noise, signal-to-noise ratio, etc.), and the device itself (temperature, vibration amplitude, etc.), and maps these multi-dimensional statuses to a unified content quality score according to preset rules or machine learning models. In this way, the transmission system can accurately identify low-quality or invalid data caused by factors such as device defocusing, abnormal exposure, excessive noise, or device malfunction before transmission, thereby proactively discarding such data or significantly reducing its transmission priority, avoiding the occupation of valuable network bandwidth by junk data. Simultaneously, this mechanism can provide real-time acquisition quality feedback for high-value business scenarios (such as identity verification and industrial inspection), guiding users or terminal devices to adjust their behavior to improve data quality. Ultimately, based on the scheme exemplified in this specification, by using the content quality score to influence the priority score, and thus affecting the differentiated transmission scheduling of each data unit, it is possible to achieve quality filtering and control from the source of collection, save uplink bandwidth, and ensure that the data transmitted to the cloud or backend analysis module has high availability. This is beneficial to improving the overall data transmission efficiency and the accuracy and reliability of backend AI model processing.
[0092] In one or more embodiments of this specification, for a data unit of video frame type, its frame type and inter-frame decoding dependency relationship are parsed, a decoding dependency graph is constructed, and a dependency weight score is output according to the number of subsequent data units affected by the data unit in the decoding dependency graph.
[0093] The evaluation of the data unit based on the content dimension in step 204 may specifically include evaluating the dependency weight of the data unit.
[0094] Specifically, the data unit includes video frames; evaluating the data unit based on the transmission dimension may include: parsing the frame type of the video frame; the frame type includes keyframes, prediction frames, and long-term reference frames; constructing a decoding dependency graph according to the frame type; the decoding dependency graph describes the decoding dependencies between the video frames; and calculating the dependency weight score of the video frame based on the decoding dependency graph. The dependency weight score can be used to characterize the importance of the corresponding video frame in the decoding link.
[0095] The frame type is obtained by parsing the syntax elements in the video encoded data. Specifically, parsing the frame type of the video frame can be understood as parsing the frame type of the video data unit and whether it is a long-term reference frame.
[0096] Mainstream video coding standards (such as H.264, H.265 / HEVC, and AV1) define the syntax and encoding methods for keyframes and prediction frames. A keyframe (I-frame) is an independent coded frame containing complete image information. It can be decoded independently without referencing other frames and is used as a reference for decoding other P-frames. Losing a keyframe usually leads to decoding failures for several subsequent frames. A prediction frame (P-frame) only stores the difference information from the previous keyframe and is used as a reference for decoding previously encoded frames (I-frames or P-frames). Prediction frames (P-frames) are relatively small, approximately 10%-30% the size of an I-frame.
[0097] Before the encoder starts working, some key parameters can be configured—the GOP structure (i.e., "group of pictures" structure), which specifically defines the period at which keyframes (I-frames) appear. For example, setting the GOP length to 120 frames (equivalent to 4 seconds in a 30fps video) means that at least one I-frame can be encoded every 4 seconds to ensure random access and fault tolerance.
[0098] In practical applications, content-driven encoding optimization can also be implemented. The appearance of I-frames is determined not only by the period but also by sudden changes in content. Specifically, within a fixed GOP framework, the encoder can analyze the video content in real time using rate-distortion optimization algorithms to determine whether each video frame is encoded as an I-frame or a P-frame, in order to obtain good image quality with a small amount of data. Suppose that in the middle of the GOP period, the shot suddenly cuts from indoors to outdoors. The encoder can identify that this frame has a low similarity to the previous frame. At this time, the encoder can break the preset GOP structure and insert an I-frame. For non-I-frames, the encoder can encode them as P-frames and determine the decoding dependency frames of the P-frame, such as the most recent previous frame (I or P) or certain LTR frames.
[0099] A Long-Term Reference Frame (LTR frame) is a special reference frame that is explicitly marked by the encoder as a long-term reference for multiple subsequent frames. In practical applications, some predicted frames (P frames) can often be designated as LTR frames.
[0100] In the embodiments described in this specification, the Long-Term Reference Frame (LTR) marking function can be dynamically enabled. Specifically, during normal encoding, all frames use only Short-Term Reference Frames (LTRs) and LTRs are not enabled. During video encoding, the terminal device can enable or disable the LTR function as needed based on the real-time scenario, content importance, or network conditions. Furthermore, based on the analysis results of edge service content, it can be determined in real time whether the current frame meets preset importance conditions (e.g., contains key semantic information, detects key events (e.g., face appearance, abnormal behavior)). If the preset importance conditions are met, the encoder determines that the current frame is a high-quality frame (P-frame or I-frame), and then marks the current high-quality frame as an LTR frame. For example, the current frame can be marked as an LTR frame through video encoding standard signaling (e.g., the MMCO instruction in H.264 / AVC, or the RPS mechanism in H.265 / HEVC), that is, the frame is added to the "Long-Term Reference Frame List" in the long-term reference buffer so that subsequent frames can selectively reference the LTR frame.
[0101] Specifically, P frames following an LTR frame can be directly designated to reference the LTR frame, without depending on (or not only depending on) the nearest I / P frame. For example, frame 0 (I0) is marked as an LTR frame; frames 1-4 (P1-P4) are normal prediction frames, each referencing the previous frame during decoding; frame 5 (P5) is explicitly designated to reference frame 0 (I0), not frame 4 (P4); therefore, the decoding of frame 5 (P5) does not depend on frames 1-4 (P1-P4), but on frame 0 (I0).
[0102] The content-driven LTR tagging based on the embodiments of this specification is triggered only when edge AI detects "high-value content," thus the tagged LTR frames naturally have high business semantic importance. Furthermore, by allocating LTR resources on demand, "wasting reference slots for unimportant frames" is avoided.
[0103] Based on the embodiments of this specification, by implementing the dynamic LTR marking function, it is possible to dynamically evaluate the decoding impact range and dependency weight score of a video frame. For example, after a normal P frame is set to LTR, its "decoding dependency impact range" jumps from 0 to N (e.g., N = the number of subsequent frames referencing it), and consequently, the corresponding dependency weight score (RecoveryCost Score) increases sharply. This can improve the priority score and protection strength of the video frame during network transmission.
[0104] Furthermore, constructing the decoding dependency graph based on the frame type may specifically include: analyzing the decoding dependencies between the video frames based on the frame type, and constructing a decoding dependency graph. More specifically, an inter-frame decoding dependency graph may be constructed based on the frame type and reference frame management signaling in the video stream. The decoding dependency graph can represent the frame dependency structure, specifically representing the decoding dependencies between each video frame. The frame type and decoding dependencies can be determined from the encoded bitstream data.
[0105] More specifically, constructing the decoding dependency graph may include: analyzing the reference frame list for each video frame, establishing directed edges (directed dependency chains) from the referenced frame to the reference frames, thereby forming a directed acyclic graph. Here, the I-frame serves as the decoding starting point, P-frames reference previous I-frames or P-frames, and LTR frames are referenced by multiple subsequent video frames over a long period.
[0106] Furthermore, calculating the dependency weight score of the video frame based on the decoding dependency graph may specifically include: determining the importance of the video frame in the decoding link based on the decoding dependency graph, and generating a dependency weight score.
[0107] The dependency weight score reflects the overall impact of the lost video frame on the decoding success rate and decoding quality of subsequent video frames. The impact on the decoding success rate of subsequent video frames can represent the impact on the decoding chain; for example, it can reflect how many subsequent video frames cannot be decoded after the frame is lost. The impact on the decoding quality of subsequent video frames reflects the degree to which the lost frame reduces video quality because subsequent frames cannot reference it (even if decoding does not fail). Specifically, the dependency weight score is used to characterize the number of decoding failures and / or the degree of degradation in reconstruction quality caused by the loss of the lost video frame.
[0108] In one or more embodiments of this specification, calculating the dependency weight score of the video frame based on the decoding dependency graph specifically includes: evaluating the decoding interruption impact score caused by the loss of the video frame based on the decoding dependency graph; evaluating the video quality degradation impact score caused by the loss of the video frame; and calculating the dependency weight score based on the decoding interruption impact score and the video quality degradation impact score.
[0109] Furthermore, the decoding interruption impact score is determined based on the number of subsequent frames that are affected and cannot be decoded. The decoding interruption impact score can be positively correlated with the number of frames that are completely undecoded due to the loss of a necessary reference frame. For example, if an I-frame is lost and all subsequent P-frames become invalid, the decoding interruption impact score can be positively correlated with the number of subsequent P-frames.
[0110] Furthermore, the impact of the video quality degradation is determined based on at least one of the following factors: whether the video frame is a long-term reference frame (LTR frame); the content complexity of the video frame (the greater the information content, the higher the reference value); and the reference interval distance from subsequent frames to the video frame (the greater the distance, the worse the alternative reference).
[0111] Furthermore, the calculation of the video quality degradation impact score specifically includes: when the video frame is an LTR frame, the video quality degradation impact score of the LTR frame is proportional to the time interval between subsequent reference frames and the LTR frame. In practical applications, for example, the longer the time interval, the greater the loss of compression efficiency due to the lack of high-quality references, and the higher the video quality degradation impact score.
[0112] Furthermore, the evaluation of the content complexity of the long-term reference frame includes at least one of the following: analyzing the texture complexity of the long-term reference frame; analyzing the motion complexity of the long-term reference frame; and analyzing the ratio of the coding bit rate to the average coding bit rate of the long-term reference frame.
[0113] Furthermore, calculating the dependency weight score based on the decoding interruption impact score and the video quality degradation impact score specifically includes: performing a weighted sum of the decoding interruption impact score and the video quality degradation impact score to obtain the dependency weight score. Optionally, the weight of the decoding interruption impact score can be higher than the weight of the video quality degradation impact score.
[0114] In one or more embodiments of this specification, the dependency weight score setting for data units that are not video frames can be a default value. For example, when the data unit is not a video frame, the dependency weight score or its weight can be set to a low value (e.g., zero).
[0115] Based on the embodiments described in this specification, by constructing a decoding dependency graph and designing a two-dimensional dynamic evaluation mechanism, the criticality of each video frame in the decoding stream can be intelligently and accurately quantified. Specifically, a decoding dependency graph is first constructed based on frame type and reference relationship. Then, dependency weight scores are calculated from two core dimensions: the impact of decoding interruption and the impact of video quality degradation. The former quantifies the number of subsequent frames that cannot be decoded due to frame loss, directly ensuring decoding continuity. The latter comprehensively evaluates factors such as whether a frame is a dynamically marked long-term reference frame (LTR), its content complexity, and reference distance, reflecting the potential damage of frame loss to reconstruction quality. In particular, through a content-driven dynamic LTR marking mechanism, terminal devices can upgrade high-quality frames to LTRs as needed based on business semantic importance (such as critical events), thereby intelligently expanding their decoding impact range and dynamically increasing their weight scores, realizing the linkage between encoding optimization and transmission priority. Finally, by weighted fusion of the two scores, a comprehensive score representing the importance of each data unit (video frame or non-video frame) can be output. This score, by influencing the priority score, enables the transmission system to accurately identify and prioritize the protection of key frames that have a significant impact on decoding continuity and visual quality when the network is limited. This improves the overall robustness of video transmission, reconstruction quality, and bandwidth resource utilization efficiency.
[0116] In one or more embodiments of this specification, the system employs a dynamic, quantitative network transmission cost evaluation model. This model comprehensively considers the attributes of the data unit itself, real-time network status, and service timeliness requirements to accurately calculate the expected network resource cost required to transmit each data unit.
[0117] The network transmission cost may include the initial transmission cost and the packet loss recovery cost (also known as recovery cost).
[0118] In step 204, the data unit is evaluated based on the transmission dimension. Specifically, this may include: calculating the amount of redundant data required to ensure the success rate of the first transmission based on the original data size of the data unit, the current network status parameters, and the preset data transmission success rate target, thus obtaining the first transmission cost; and determining the transmission cost score based on the first transmission cost.
[0119] The initial transmission cost can refer to the total amount of data that needs to be transmitted to achieve the required transmission success rate when a data unit is first sent.
[0120] The current network state parameters include at least one of packet loss rate, round-trip time, and available bandwidth. Optionally, the current network state parameters used to calculate the initial transmission cost specifically include the packet loss rate. The amount of redundant data required in the initial transmission cost can be calculated using a forward error correction coding model based on the packet loss rate.
[0121] In an optional embodiment, calculating the initial transmission cost may specifically include calculating the amount of redundant data required to ensure successful initial transmission using a forward error correction coding model. The amount of redundant data may have a non-linear relationship with the original data size due to factors such as scale effects, packet optimization, and diminishing marginal returns. In practical applications, the forward error correction coding model can establish a relationship between packet loss rate, data size, and required redundancy based on information theory and coding theory. For example, for 10KB of data with a 20% packet loss rate, 5KB of redundancy may be required (totaling 15KB); while for 20KB of data with a 20% packet loss rate, 8KB of redundancy may be required (totaling 28KB, which is not a simple doubling of 15KB).
[0122] Furthermore, evaluating the data unit based on the transmission dimension may further include: estimating the amount of retransmitted data required to recover from potential packet loss within the remaining effective transmission time of the data unit and the current network state parameters, thereby obtaining the packet loss recovery cost. Correspondingly, determining the transmission cost score based on the initial transmission cost specifically includes: determining the transmission cost score based on the initial transmission cost and the packet loss recovery cost. In practical applications, the transmission cost score can be determined using a preset function based on the initial transmission cost and the packet loss recovery cost. For example, the transmission cost score can be determined based on the sum of the initial transmission cost and the packet loss recovery cost.
[0123] Among them, the packet loss recovery cost can refer to the amount of retransmitted data required to recover lost data within the time constraint when packet loss is detected after the initial transmission.
[0124] The remaining valid transmission time reflects the timeliness / expiration time of data in the current business scenario, indicating that data needs to arrive at the receiving end within a certain time limit to be valid. Specifically, the remaining valid transmission time can represent the remaining time before the expiration time defined by the business for the data unit. In practical applications, the data timeliness is usually defined by the transmission latency requirements of the upper-layer business for the data unit, and is usually related to the business scenario. For example, face recognition requires ≤200ms, and industrial quality inspection requires ≤500ms.
[0125] The current network state parameters include at least one of packet loss rate, round-trip time, and available bandwidth. Optionally, the current network state parameters used to calculate the initial transmission cost specifically include packet loss rate, round-trip time, and available bandwidth. The additional amount of retransmitted data required for the packet loss recovery cost can be calculated using a retransmission probability model based on round-trip time and packet loss rate.
[0126] In an optional embodiment, calculating the packet loss recovery cost may specifically include calculating the amount of additional retransmitted data required to complete effective recovery within the remaining time using a retransmission probability model, thereby obtaining the packet loss recovery cost. More specifically, the retransmission probability model estimates the maximum number of retransmissions that can be performed within the remaining effective transmission time based on the current network's round-trip latency and packet loss rate, and calculates the amount of retransmitted data based on the maximum number of retransmissions and the packet loss rate. In practical applications, when the expected recovery time exceeds the service timeliness, the packet loss recovery cost can be set to a maximum value, significantly reducing the priority score of the data unit.
[0127] In an optional embodiment, determining the transmission cost score specifically includes adding the initial transmission cost to the packet loss recovery cost to obtain the total transmission cost, and mapping it to a normalized transmission cost score.
[0128] In an optional embodiment, the transmission cost score is negatively correlated with the sum of the initial transmission cost and the packet loss recovery cost; that is, the higher the total cost, the lower the score. As an example, the transmission cost score can be normalized as follows: Transmission_Cost_Score = 100 × (1 min(Total_Cost / Cost_Max, 1)), where Cost_Max is the preset maximum acceptable transmission cost threshold, and Total_Cost is the sum of the initial transmission cost and the packet loss recovery cost.
[0129] In an optional embodiment, when the expected total transmission cost exceeds the maximum amount of data that the remaining effective transmission time can support, the transmission cost score is set to a minimum value, so that the data unit is preferentially discarded.
[0130] In practical applications, both the initial transmission cost and the packet loss recovery cost can be calculated per frame of data or per segment of data (containing one or more frames of data). The data unit can be a single frame of video, a single frame of audio, or a logical data segment containing multiple frames.
[0131] To facilitate understanding of the calculation of transmission cost score, an example is provided below.
[0132] Consider the following scenario: Network conditions: bandwidth 1000kbps, packet loss rate 20%, RTT 50ms. Service requirements: target transmission success rate 99%, target transmission time 200ms. Data units include: Data segment 1: 10KB, containing 10 data packets (1KB each); Data segment 2: 20KB, containing 20 data packets (1KB each).
[0133] The cost of the first transmission is calculated as follows.
[0134] For data segment 1 (10KB, 10 packets), the required number of redundant packets is calculated to be 5 (redundancy rate 50%); the total number of packets for the first transmission is 15; the cost of the first transmission is 15KB.
[0135] For data segment 2 (20KB, 20 packets), the required number of redundant packets is calculated to be 8 (redundancy rate 40%); the total number of packets in the first transmission is 28; and the cost of the first transmission is 28KB.
[0136] The cost of packet loss recovery is calculated as follows.
[0137] Assume packet loss is detected in data segment 1 at 50ms. Number of lost packets: 2 (estimated based on a 20% packet loss rate). Remaining time: 200ms - 50ms = 150ms. RTT: 50ms. Maximum number of retransmissions: floor(150ms / 50ms) = 3 times. Calculate the amount of data to be retransmitted: Success rate of a single retransmission: 1 - 0.2 = 0.8; Expected number of transmissions to recover 2 packets within 3 retransmissions: 2.5 packets; Packet loss recovery cost: 2.5KB.
[0138] Assume packet loss is detected in data segment 2 at 150ms. Number of lost packets: 4 (estimated based on a 20% packet loss rate). Remaining time: 200ms - 150ms = 50ms. RTT: 50ms. Maximum retransmission count: floor(50ms / 50ms) = 1 time. Calculate the amount of data to be retransmitted: Only one retransmission opportunity, requiring a high success rate; multiple redundant packets need to be retransmitted simultaneously; expected number of transmissions: 10 packets; packet loss recovery cost: 10KB.
[0139] The total transmission cost is calculated as follows.
[0140] Total cost of data segment 1: 15 + 2.5 = 17.5KB. Total cost of data segment 2: 28 + 10 = 38KB. The total transmission cost is mapped to a transmission cost score as follows.
[0141] Let's assume that transmission costs are mapped to a rating range of 0-10 (the lower the cost, the higher the rating). For example, data segment 1: 17.5KB, corresponding to a rating of 8.0; data segment 2: 38KB, corresponding to a rating of 6.5.
[0142] Based on the embodiments described in this specification, the system (terminal device) can intelligently quantify the transmission cost of each data unit through the aforementioned transmission cost evaluation mechanism, providing a precise basis for priority decision-making. This mechanism is particularly suitable for edge computing scenarios with dynamically changing network conditions and diverse service requirements, achieving optimized allocation of network resources and effective assurance of service quality.
[0143] More specifically, a joint quantification model of initial transmission cost and packet loss recovery cost is introduced to achieve a refined assessment of network transmission costs. Unlike traditional static strategies that only consider the size of the original data, this scheme dynamically calculates the FEC redundancy required to ensure transmission reliability based on factors such as the current packet loss rate, RTT, and service timeliness before transmission; and assesses possible packet loss scenarios, combining the remaining effective time constraint to accurately estimate retransmission overhead. This two-stage cost modeling mechanism enables the system (terminal device) to identify the difference between "seemingly small data volume but difficult to recover" and "large data volume but easy to recover" (such as early high-redundancy frames), thereby making better scheduling decisions under limited bandwidth. This mechanism significantly improves the effective delivery rate of high-value data in weak network environments, while avoiding the waste of bandwidth resources on data that is "theoretically important but cannot be recovered in time," providing a stable and reliable data input guarantee for real-time cloud analysis.
[0144] Based on the embodiments described in this specification, a refined and forward-looking calculation of the resource consumption of data unit transmission is achieved by constructing a dynamic and quantitative evaluation method for network transmission. Specifically, this method comprehensively considers data unit attributes, real-time network status (packet loss rate, latency, bandwidth), and service timeliness requirements, decomposing the total transmission cost into two parts for joint evaluation: the initial transmission cost and the packet loss recovery cost. The former, based on a forward error correction coding model, calculates the amount of redundant data required to achieve the target transmission success rate; the latter, combining the remaining effective transmission time and network round-trip latency, estimates the potential retransmission overhead required to recover packet loss within the time constraint using a retransmission probability model. By mapping the sum of the two to a normalized transmission cost score, the system (terminal device) can accurately identify data units that are "small in size but difficult to recover near the deadline" or "large in size but with sufficient redundancy for easy assurance," thereby making better decisions during resource scheduling. Compared to traditional static methods, this mechanism transforms transmission cost assessment from simply looking at the size to dynamically calculating resource consumption. This allows the transmission system to prioritize high-value data transmissions that need to be delivered promptly in weak network or congested environments, significantly improving the utilization efficiency of network resources and the reliability and real-time performance of cloud service data streams.
[0145] Step 206: Determine the priority score of each data unit based on the evaluation results.
[0146] Specifically, a priority score can be calculated for each data unit in the perception-related data stream, and then the priority score of each data unit can be determined based on the priority score. For example, the higher the priority score, the higher the priority rating.
[0147] In optional embodiments of this specification, the calculation of the priority score can be performed independently within each data stream. More specifically, it can be performed independently within data streams of the same data type. For example, video frames in a video frame sequence are ordered relative to each other. Similarly, audio frames in an audio frame sequence are ordered relative to each other. Furthermore, data units in a sensor data sequence are ordered relative to each other. Single-modal scoring is possible, but parallel intelligent transmission of multi-channel sensing-related data can also be supported.
[0148] In an optional embodiment, the priority score can be calculated between data streams. More specifically, it can be performed between data streams of different data types. Therefore, cross-modal unified scoring can be performed, and multi-channel sensing-related data can be comprehensively scheduled and transmitted based on the same score.
[0149] In one or more embodiments of this specification, a unified and quantifiable multi-dimensional scoring system is established to integrate heterogeneous evaluation information into a single decision-making basis.
[0150] Specifically, in step 204, each data unit can be evaluated based on its content dimension to obtain a content score; and evaluated based on its transmission dimension to obtain a transmission score. Furthermore, for each data unit, the content score and transmission score can be weighted and calculated to obtain a comprehensive priority score.
[0151] In practical applications, when using multi-dimensional comprehensive evaluation, if the reference value of a certain dimension's score is low or not applicable in some scenarios, the weight coefficient corresponding to that evaluation dimension can be set to a lower value.
[0152] In step 206, determining the priority score of each data unit based on the evaluation results may specifically include: for each data unit, performing a weighted calculation of the content importance score and content quality score of the content dimension, and the dependency weight score and transmission cost score of the transmission dimension to obtain a priority score.
[0153] Furthermore, the priority score can be positively correlated with the content importance score, content quality score, and dependency weight score, and negatively correlated with the transmission cost score.
[0154] In practical applications, to ensure the comparability and synergy of evaluation results across all dimensions, the evaluation of all dimensions is based on a single data unit (such as a video frame, an audio frame, or a sensor sampling packet) as the basic granularity. This ensures that scores from different dimensions (content importance, content quality, dependency weight, and transmission cost) can be calculated on the same scale.
[0155] In practical applications, to ensure the comparability and synergy of evaluation results across all dimensions, the original evaluation values can be standardized—normalization: the original scores for each dimension (e.g., 0-100 points, or values based on different units such as confidence level, distance, milliseconds, bytes, etc.) are mapped to a common numerical range, such as [0, 1] or [0, 100], using a linear or non-linear function. For example, the transmission cost score might initially be the estimated amount of additional data (e.g., 150KB), which can be converted into a standardized score positively correlated with priority by dividing by a preset maximum cost threshold (e.g., 500KB) and inverting the result.
[0156] Furthermore, by performing multi-dimensional fusion of the scoring results of the content dimension and the transmission dimension through a weighting formula, a comprehensive priority score can be obtained.
[0157] Optionally, the weighted fusion formula used for weighted calculation can be: Priority Score = α * Content Importance Score + β * Content Quality Score - γ * Transmission Cost Score + δ * Dependency Weight Score, where α, β, γ, and δ are configurable weight coefficients.
[0158] For example, the weighted fusion formula is: Priority=α*Content_Importance+β*Content_Quality-γ*Transmission_Cost+δ*Dependency_Weight.
[0159] Among them, Priority is the priority score, Content_Importance is the content importance score, Content_Quality is the content quality score, Transmission_Cost is the transmission cost score, and Dependency_Weight is the dependency weight score, where α, β, γ, and δ are weight coefficients.
[0160] In practical applications, Content_Importance_Score, Content_Quality_Score, Transmission_Cost_Score, and Decoding_Importance_Score are respectively the standardized (normalized) content importance score, content quality score, transmission cost score, and dependency weight score.
[0161] In one or more embodiments of this specification, the weighting coefficients used for weighted calculations are configurable coefficients.
[0162] Furthermore, the weighting coefficients used for weighted calculation can correspond to the current business scenario. For example, the weighting coefficients α, β, γ, and δ in the above weighted fusion formula can be configured according to different business scenarios.
[0163] In practical applications, the weighting coefficients can be dynamically adjusted according to the type of business scenario. The weighting coefficients are used to quantify the relative importance attached to different objectives such as "content value," "transmission efficiency," and "decoding continuity" under different business scenarios.
[0164] The business scenarios mentioned can include, but are not limited to: identity verification, industrial inspection, live streaming interaction, security alarms, human-computer interaction dialogue, and digital human generation. For example, in the identity verification scenario, the weight of content importance (α) and decoding dependency (δ) can be increased, while the weight of transmission cost (γ) can be decreased. Similarly, in the security inspection scenario, the weight of content importance (α) can be dynamically adjusted based on event detection results, and the weight of decoding dependency (δ) can be increased when an event occurs. Furthermore, in the high-concurrency live streaming scenario, the weight of transmission cost (γ) can be increased to optimize bandwidth efficiency. Finally, in the industrial quality inspection scenario, for video frames with detected defects, both the weight of content importance (α) and the weight of content quality (β) can be increased simultaneously.
[0165] In one or more embodiments of this specification, the weight coefficients can be automatically optimized using an online learning model, which adjusts the weights of each dimension inversely based on the accuracy of the cloud-based AI analysis results. In practical applications, the system can use the business results feedback of historical transmission data (such as the accuracy, success rate, and average processing latency of cloud-based AI recognition) as optimization targets. Through a lightweight online learning model (such as a simple optimizer based on gradient descent), the values of the weight coefficients (α, β, γ, δ) are periodically fine-tuned, ensuring that the scheduling strategy continuously approaches the optimized business effect.
[0166] Furthermore, the end server can aggregate operational data from massive numbers of terminals, perform centralized analysis and model training, generate optimized weight configuration templates, and uniformly distribute them to each terminal device to achieve rapid strategy iteration and global optimization. Correspondingly, the method also includes: receiving a weight coefficient update instruction from the business server; adjusting the weight coefficients (α, β, γ, δ) in the weighted fusion formula according to the update instruction; wherein the update instruction is generated based on the business server's statistical analysis of historical transmission data in terms of business identification success rate and response time.
[0167] In the embodiments of this specification, each type of data in the perception-related data stream can be independently evaluated and scheduled based on priority scores; wherein, video data units are compared in priority with each other, and audio data units are compared in priority with each other.
[0168] Based on the embodiments described in this specification, a decision-making core combining business intelligence and network awareness is constructed through the aforementioned multi-dimensional fusion scoring mechanism. This transforms the data transmission system from a passive pipeline into an intelligent edge capable of proactively optimizing business outcomes. The complex transmission scheduling problem is converted into a computable and analyzable multi-dimensional optimization problem, making the data transmission system behavior more predictable and debuggable. Furthermore, by adjusting weights, the same core algorithm can be quickly adapted to vastly different business domains such as security, live streaming, IoT, and remote operation, demonstrating high flexibility and reusability.
[0169] In the embodiments of this specification, the terminal device's actions, such as parsing the frame type (I-frame / P-frame / LTR frame), analyzing decoding dependencies, constructing a decoding dependency graph, and calculating a comprehensive priority score, can be performed before the actual encoding of the video frames. Therefore, the comprehensive priority score can be used to determine which frames will be encoded and transmitted.
[0170] Based on the embodiments described in this specification, a unified multi-dimensional fusion scoring mechanism is established. Heterogeneous evaluation results, such as content importance, content quality, decoding dependency weight, and network transmission cost, are dynamically weighted and fused using configurable weight coefficients to generate a comprehensive priority score for each data unit. This constructs an intelligent scheduling decision-making core that combines business semantic understanding and network state awareness. This mechanism supports unified scoring and ranking within a single-modal data stream or across cross-modal data streams. It also allows for dynamic adjustment of the weight ratio of each dimension based on the needs of different business scenarios such as identity verification, security patrol, and live streaming interaction. Furthermore, it can continuously optimize the weight configuration based on business feedback through an online learning model. Therefore, the data transmission system can transform the complex transmission scheduling problem into a computable and iterative optimization process. Before encoding and transmission, it proactively identifies data units with high business value, high dependency, and controllable transmission costs, assigning them higher transmission priority. This significantly improves the delivery rate and timeliness of high-value data in weak network environments.
[0171] Step 208: Perform differentiated transmission scheduling for each data unit according to the priority score.
[0172] The transmission scheduling may include a series of actions such as determining the sending order based on priority, selecting the sending content, and allocating network resources.
[0173] Furthermore, the differentiated transmission scheduling for each data unit may specifically include: prioritizing the transmission of data with high priority scores (such as data with scores higher than a preset threshold) and applying enhanced protection measures (such as FEC redundancy, retransmission, etc.); delaying the transmission of data with low priority scores (such as data with scores lower than a preset threshold) and discarding it when bandwidth is limited.
[0174] In one or more embodiments of this specification, different scheduling strategies can be adopted based on the current network conditions. When network bandwidth is sufficient, all data units can be transmitted. When network bandwidth is insufficient, a subset of data units can be selected for priority transmission.
[0175] In step 208, the differentiated transmission scheduling of each data unit according to the priority score may specifically include: calculating the maximum amount of data that can be transmitted within a predefined time window based on the current network bandwidth; selecting at least a portion of the data units whose total data volume does not exceed the maximum data volume from the data units to be transmitted, and adding them to the transmission queue based on the priority score of each data unit; the priority score of the data units added to the transmission queue is greater than or equal to the priority score of the data units not added to the transmission queue. In practical applications, the selection is performed in descending order of the priority scores.
[0176] Specifically, the data units to be transmitted within the time window can be sorted from highest to lowest priority score; then, data units are selected sequentially and added to the transmission queue until the total amount of data in the transmission queue reaches the maximum data volume. Further, if the total amount of data to be transmitted within the time window exceeds the maximum data volume, data units not yet added to the transmission queue are discarded.
[0177] In practical applications, the predefined time window is a sliding time window. Within each sliding time window, data units can preempt the transmission resources within the window according to their priority scores; data units with higher scores get priority in obtaining transmission opportunities until the transmission resources within the window are fully occupied.
[0178] Based on the embodiments of this specification, different transmission priorities can be assigned to different data units according to the comprehensive priority score evaluation result, thereby realizing priority-based Scheduling, so that high-value or high-recovery-cost data units are sent first.
[0179] Furthermore, the differentiated transmission scheduling also includes applying enhanced protection measures to high-priority data units. Specifically, during the network transmission phase, differentiated protection measures can be further applied to each data unit in the transmission queue based on its priority score.
[0180] Furthermore, the differentiated transmission scheduling of each data unit according to the priority score may further include: performing differentiated protection on the data units added to the transmission queue; the differentiated protection includes at least one of the following: A forward error correction coding redundancy ratio is allocated to data units with a priority score higher than a first threshold; the redundancy ratio is positively correlated with the priority score; that is, the redundancy ratio can increase as the priority score increases.
[0181] For data units with a priority score higher than the second threshold, retransmission is performed after packet loss is detected; the number of retransmissions is positively correlated with the priority score; that is, the number of retransmissions is dynamically adjusted according to the priority score, and the higher the priority, the more retransmissions are performed.
[0182] Data units with priority scores higher than the third threshold are placed into a high-priority transmission queue, which has priority scheduling rights or is guaranteed minimum bandwidth when the network is congested.
[0183] The first threshold, the second threshold, and the third threshold may be the same or different. The first threshold, the second threshold, and the third threshold may be determined statistically based on historical data transmission records or based on expert experience, and are not limited thereto.
[0184] Forward Error Correction (FEC) is a packet loss mitigation technique that does not rely on retransmissions. Specifically, the sender sends additional redundant packets (checksum packets) along with the original data. Therefore, even if the receiver loses some original packets, it can recover the lost data itself as long as it receives enough original and redundant packets without requesting retransmissions. This avoids RTT latency and is suitable for low-latency scenarios (such as video calls and AI inspections). In practical applications, adding FEC redundancy to a data frame means generating redundant packets for that frame and sending them together, improving its packet loss mitigation capability, but at the cost of increased bandwidth overhead. In the embodiments of this specification, for frames with high priority scores, the FEC coding redundancy ratio is actively increased; for frames with low priority scores, the FEC coding redundancy ratio is reduced or eliminated to save bandwidth.
[0185] In this context, data units with higher priority scores are allocated a higher redundancy ratio and are allowed a greater number of retransmissions. In practical applications, the number of retransmissions can be dynamically determined based on the priority score, the current network round-trip time, and the remaining effective transmission time after packet loss.
[0186] To make it easier to understand, an example is provided here.
[0187] Under weak network conditions (e.g., a measured available uplink bandwidth of 500 Kbps), with a sliding time window of 200 ms, the maximum amount of data that can be transmitted within one sliding time window is 100 KB. Assume there are 4 video frames to be transmitted, and the priority scores for each frame are as follows: first frame F1-95 points, second frame F2-20 points, third frame F3-85 points, and fourth frame F4-10 points.
[0188] First, based on the scheduling decision of comprehensive priority, the frames to be transmitted are selected. The four frames are arranged in descending order of priority score: F1 (95) > F3 (85) > F2 (20) > F4 (10). Within a 200 ms sliding window, the frames are added to the transmission queue in this order, with the total amount of original data not exceeding 100 KB: F1 (60 KB) is added, total = 60 KB; F3 (35 KB) is added, total = 95 KB ≤ 100 KB, accepted; F2 (10 KB) is added, total = 105 KB > 100 KB, rejected. The final transmission queue is: {F1, F3}. Frames discarded: F2, F4.
[0189] Subsequently, differentiated transmission protection based on comprehensive priority is implemented. After F1 and F3 enter the network transmission stage, further hierarchical protection is applied according to their priority scores. For example, a 30% FEC redundancy ratio is allocated to the first frame, and a 15% redundancy ratio is allocated to the third frame. It is also assumed that if packet loss occurs, the first frame can be retransmitted twice, and the third frame can be retransmitted once.
[0190] As can be seen, by prioritizing and selecting data based on scores, transmission resources (100KB budget) are ensured to be allocated to the data combination with the highest overall value (F1+F3). Subsequently, tiered protection is applied based on the overall score, allocating more resources (higher redundancy, more retransmission opportunities) to the higher-value F1 to ensure its reliable delivery. In this way, under bandwidth constraints, both "what to send" (maximizing business value) and "how to ensure delivery" (differentiated reliability) are optimized.
[0191] The solution based on the embodiments of this specification not only avoids low-value data occupying bandwidth during the selection phase, but also provides high-value data with reliability guarantees matching its importance during the transmission phase, thereby improving the overall service efficiency of data transmission in weak network environments.
[0192] Based on the embodiments described in this specification, a two-stage differentiated transmission scheduling mechanism based on priority scores is constructed to achieve joint optimization of data unit transmission decisions and transmission assurance under weak network or bandwidth-constrained conditions. This mechanism first sorts and filters data units to be transmitted according to priority scores when network resources are limited, adding data to the transmission queue in descending order until the bandwidth budget of the current time window is fully utilized, thus ensuring that transmission resources are always preferentially allocated to the data combination with the highest overall value. Subsequently, during the data transmission phase, a hierarchical protection strategy is further applied based on priority scores—higher forward error correction redundancy ratios are dynamically allocated to high-priority data units, allowing more retransmissions and placing them in the high-priority queue, thereby providing them with reliability guarantees commensurate with their business importance. Through this two-layer scheduling of prioritizing transmission followed by hierarchical protection, the data transmission system can not only proactively avoid reducing bandwidth consumption by low-value data but also provide enhanced transmission reliability for high-value data, ultimately achieving synergistic optimization of maximizing business value, efficiently utilizing network resources, and improving transmission robustness.
[0193] While one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is merely one possible execution order among many steps and does not represent the only possible execution order. The order of some steps may be adjusted according to actual needs, or some steps may be omitted. When the claims involve method steps, changes in the order of such steps, or parallel execution between steps, are also within the scope of protection of the claims.
[0194] Figure 2 In this process, terminal devices quantitatively evaluate each data unit in the data stream to be transmitted based on at least one dimension: content (assessing its content importance and acquisition quality) and transmission (assessing its dependency weight and transmission cost under the current network conditions). This identifies data units with high business value and high acquisition quality and dynamically senses the transmission cost of each data unit under the current actual network conditions. Then, based on the above multi-dimensional evaluation results, a priority score for each data unit can be comprehensively determined, and differentiated transmission scheduling can be performed on each data unit according to the priority score. This enables on-demand and optimal allocation of network bandwidth resources, allowing the data transmission system to automatically and in real time prioritize the successful transmission of data units that are more critical to business implementation based on priority scores, even when network resources are limited. This effectively solves the problems of network resource waste and the inability to guarantee the arrival of critical data under weak network conditions, significantly improving the analysis success rate, business reliability, and overall efficiency of the data transmission system for backend AI services.
[0195] The various technical features in the above embodiments can be combined arbitrarily, as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they have not been described one by one. Therefore, the arbitrary combination of various technical features in the above embodiments is also within the scope of this specification.
[0196] Figure 3 This is a schematic diagram illustrating the process of transmitting data from a terminal device to a server in a practical application scenario provided by an embodiment of this specification.
[0197] like Figure 3 As shown, in step 302, the data to be transmitted is obtained.
[0198] In practical applications, data to be transmitted can be acquired through cameras, microphones, and other sensors. The specific data acquisition channels and data types correspond to the application scenario.
[0199] Step 304: Evaluate each data unit in the data to be transmitted in the content dimension to obtain the content importance score and content quality score of each data unit.
[0200] In step 306, the execution order of the steps for determining the content importance score and the steps for determining the content quality score is not specifically limited; they can be executed simultaneously or sequentially.
[0201] Step 306: Evaluate each data unit in the data to be transmitted in the transmission dimension to obtain the dependency weight score and transmission cost score of each data unit.
[0202] In step 308, the execution order of the steps for determining the weighted score and the transmission cost score is not specifically limited; they can be executed simultaneously or sequentially.
[0203] In practical applications, the execution order of steps 304 and 306 is not specifically limited; they can be executed simultaneously or sequentially.
[0204] Step 308: Weighted fusion of content importance score, content quality score, dependency weight score, and transmission cost score to determine the priority score of each data unit.
[0205] Step 310: Perform differentiated transmission scheduling for each data unit according to priority scores.
[0206] Specifically, in the process of differentiated scheduling, data units with higher priority scores can be transmitted first, and transmission protection measures can be implemented.
[0207] Based on embodiments of the present invention, a systematic multi-dimensional evaluation and intelligent scheduling process is constructed. At the terminal side, collaborative quantitative analysis of the data to be transmitted is performed on both content and transmission dimensions (including content importance, content quality, decoding dependency weight, and network transmission cost). Based on configurable weights, the four scores are merged into a unified priority score, and then differentiated transmission scheduling is implemented for data units according to this score. This achieves multi-objective optimization of data value, acquisition quality, decoding criticality, and transmission cost. Ultimately, when network resources are limited, it can intelligently prioritize the transmission of data with high business value, high reliability requirements, and controllable transmission costs. This significantly improves bandwidth utilization efficiency, ensures the real-time and continuity of critical services, and enhances the robustness of data transmission and overall system performance in weak network environments.
[0208] Based on the same idea, embodiments of this specification also provide apparatus corresponding to the above methods.
[0209] Figure 4 The embodiments provided in this specification correspond to Figure 2 A schematic diagram of the structure of a data transmission device.
[0210] like Figure 4 As shown, the device may include: Data acquisition module 402 is used to acquire the perception-related data stream to be transmitted; the perception-related data stream includes multiple data units; The data evaluation module 404 is used to evaluate each of the data units based on at least one of the content dimension and the transmission dimension to obtain an evaluation result; the content dimension is related to at least one of the content importance and the acquisition quality of the data unit; the transmission dimension is related to at least one of the dependency weight of the data unit and the network transmission cost under the current network state. Priority score calculation module 406 is used to determine the priority score of each data unit based on the evaluation results; The transmission scheduling module 408 is used to perform differentiated transmission scheduling on each of the data units according to the priority score.
[0211] based on Figure 4 The embodiments of this specification also provide some specific implementation schemes of the method, which are described below.
[0212] Optionally, the data evaluation module 404 is specifically used to: determine the degree of relevance between the data unit and the current business scenario; and generate a content importance score based on the degree of relevance.
[0213] Optionally, determining the relevance of the data unit to the current business scenario specifically includes: using a local artificial intelligence model to detect whether the data unit involves a predefined target object or event type.
[0214] Optionally, the data evaluation module 404 is specifically used to: acquire real-time status information of the acquisition device that generates the data unit; the real-time status information reflects the acquisition quality of the data unit by the acquisition device; and generate a content quality score based on the real-time status information. The content quality score reflects the acquisition quality of the data unit. For example, the higher the acquisition quality, the higher the content quality score can be.
[0215] Optionally, the data unit includes video frames; the data evaluation module 404 is specifically used for: parsing the frame type of the video frame; the frame type includes keyframes, prediction frames, and long-term reference frames; constructing a decoding dependency graph according to the frame type; the decoding dependency graph describes the decoding dependency relationship between each video frame; and calculating the dependency weight score of the video frame based on the decoding dependency graph.
[0216] Optionally, calculating the dependency weight score of the video frame based on the decoding dependency graph specifically includes: evaluating the decoding interruption impact score caused by the loss of the video frame based on the decoding dependency graph; evaluating the video quality degradation impact score caused by the loss of the video frame; and calculating the dependency weight score based on the decoding interruption impact score and the video quality degradation impact score.
[0217] Optionally, the data evaluation module 404 is specifically used to: calculate the amount of redundant data required to ensure the success rate of the first transmission based on the original data size of the data unit, the current network status parameters, and the preset data transmission success rate target, and obtain the first transmission cost; and determine the transmission cost score based on the first transmission cost.
[0218] Optionally, evaluating the data unit based on the transmission dimension further includes: estimating the amount of retransmitted data required to recover from possible packet loss within the remaining effective transmission time based on the remaining effective transmission time of the data unit and the current network state parameters, thereby obtaining the packet loss recovery cost; the step of determining the transmission cost score based on the initial transmission cost specifically includes: determining the transmission cost score based on the initial transmission cost and the packet loss recovery cost.
[0219] Optionally, the priority score calculation module 406 is specifically used to: for each of the data units, perform a weighted calculation of the content importance score and content quality score of the content dimension, as well as the dependency weight score and transmission cost score of the transmission dimension, to obtain a priority score.
[0220] Optionally, the weighting coefficients used for weighted calculation can correspond to the current business scenario.
[0221] Optionally, the transmission scheduling module is specifically configured to: calculate the maximum amount of data that can be transmitted within a predefined time window based on the current network bandwidth; select at least a portion of the data units whose total data amount does not exceed the maximum data amount from the data units to be transmitted and add them to the transmission queue based on the priority score of each data unit; the priority score of the data units added to the transmission queue is greater than or equal to the priority score of the data units not added to the transmission queue.
[0222] Optionally, the transmission scheduling module is further configured to: perform differentiated protection on data units added to the transmission queue; the differentiated protection includes at least one of the following: allocating a forward error correction coding redundancy ratio to data units with a priority score higher than a first threshold; the redundancy ratio is positively correlated with the priority score; performing retransmission on data units with a priority score higher than a second threshold after detecting packet loss; the number of retransmissions is positively correlated with the priority score; and placing data units with a priority score higher than a third threshold into a high-priority transmission queue, the high-priority transmission queue having priority scheduling rights or being guaranteed minimum bandwidth when the network is congested.
[0223] It is understood that the modules mentioned above refer to computer programs or program segments used to perform one or more specific functions. Furthermore, the distinction between these modules does not imply that the actual program code must also be separate.
[0224] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0225] The above is an illustrative scheme of a data transmission device according to this embodiment. It should be noted that the technical solution of this data transmission device and the technical solution of the data transmission method described above belong to the same concept. For details not described in detail in the technical solution of the data transmission device, please refer to the description of the technical solution of the data transmission method described above.
[0226] Based on the same idea, this specification also provides devices corresponding to the above methods in its embodiments.
[0227] Figure 5 A structural block diagram of a computing device provided according to an embodiment of this specification is shown.
[0228] The computing device 500 includes: Memory 510 and processor 520; The memory 510 is used to store computer programs / instructions, and the processor 520 is used to execute the computer programs / instructions, which, when executed by the processor 520, implement the steps of the data transmission method.
[0229] Specifically, the components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and the database 550 is used to store data.
[0230] The computing device 500 also includes an access device 540, which enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 540 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0231] In one embodiment of this specification, the aforementioned components of the computing device 500 and Figure 5 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 5 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.
[0232] The computing device 500 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 500 can also be a mobile or stationary server.
[0233] The processor 520 implements the data transmission method when executing the computer instructions.
[0234] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above-described data transmission method belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the above-described data transmission method.
[0235] An embodiment of this specification also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the steps of the data transmission method described above.
[0236] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above-described data transmission method belong to the same concept, and all details not described in detail in the technical solution of the storage medium can be found in the description of the technical solution of the above-described data transmission method.
[0237] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described data transmission method.
[0238] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above data transmission method belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the above data transmission method.
[0239] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the apparatus and device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The apparatus, device and method provided in the embodiments of this specification are corresponding to each other, and therefore the apparatus and device also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the corresponding apparatus and device will not be repeated here.
[0240] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0241] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program a digital system themselves to "integrate" it onto a PLD, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0242] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0243] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0244] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.
[0245] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, embodiments of this specification can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0246] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, produce a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0247] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0248] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0249] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0250] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0251] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital character versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0252] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0253] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.
Claims
1. A data transmission method, comprising: Acquire the perception-related data stream to be transmitted; The perception-related data stream includes multiple data units; For each of the data units, an evaluation is performed based on at least one of the content dimension and the transmission dimension to obtain an evaluation result; the content dimension is related to at least one of the content importance and the acquisition quality of the data unit. The transmission dimension is related to at least one of the dependency weight of the data unit and the network transmission cost in the current network state; Based on the evaluation results, the priority score of each data unit is determined; According to the priority score, the transmission scheduling of each data unit is differentiated.
2. The method as described in claim 1, wherein evaluating the data unit based on the content dimension specifically includes: Determine the degree of relevance between the data unit and the current business scenario; Based on the degree of relevance, a content importance score is generated.
3. The method as described in claim 2, wherein determining the relevance of the data unit to the current business scenario specifically includes: A local artificial intelligence model is used to detect whether the data unit involves a predefined target object or event type.
4. The method as described in claim 1, wherein evaluating the data unit based on the content dimension specifically includes: Obtain the real-time status information of the acquisition device that generates the data unit; The real-time status information reflects the acquisition quality of the data unit by the acquisition device; Based on the real-time status information, a content quality score is generated.
5. The method of claim 1, wherein the data unit comprises a video frame; evaluating the data unit based on the transmission dimension specifically includes: The frame types of the video frames are analyzed; the frame types include keyframes, prediction frames, and long-term reference frames. Construct a decoding dependency graph according to the frame type; The decoding dependency graph describes the decoding dependencies between the video frames. Based on the decoding dependency graph, the dependency weight score of the video frame is calculated.
6. The method as described in claim 5, wherein calculating the dependency weight score of the video frame based on the decoding dependency graph specifically includes: Based on the decoding dependency graph, the impact of decoding interruption caused by the loss of the video frame is evaluated. Evaluate the impact of video quality degradation caused by the loss of the video frames; The dependency weight score is calculated based on the decoding interruption impact score and the video quality degradation impact score.
7. The method of claim 1, wherein evaluating the data unit based on the transmission dimension specifically includes: Based on the original data size of the data unit, the current network status parameters, and the preset data transmission success rate target, the amount of redundant data required to ensure the success rate of the first transmission is calculated, and the cost of the first transmission is obtained. A transmission cost score is determined based on the initial transmission cost.
8. The method of claim 7, further comprising evaluating the data unit based on the transmission dimension: Based on the remaining effective transmission time of the data unit and the current network status parameters, the amount of retransmitted data required to recover from possible packet loss within the remaining effective transmission time is estimated, and the packet loss recovery cost is obtained. The determination of the transmission cost score based on the initial transmission cost specifically includes: A transmission cost score is determined based on the initial transmission cost and the packet loss recovery cost.
9. The method of claim 1, wherein determining the priority score of each data unit based on the evaluation result specifically includes: For each data unit, a priority score is obtained by weighting the content importance score and content quality score of the content dimension, as well as the dependency weight score and transmission cost score of the transmission dimension.
10. The method of claim 9, wherein, The weighting coefficients used for weighted calculations correspond to the current business scenario.
11. The method of claim 1, wherein the differentiated transmission scheduling of each data unit according to the priority score specifically includes: Based on the current network bandwidth, calculate the maximum amount of data that can be transmitted within a predefined time window; Based on the priority score of each data unit, at least a portion of the data units whose total data volume does not exceed the maximum data volume are selected from the data units to be transmitted and added to the transmission queue.
12. The method of claim 11, wherein the differentiated transmission scheduling of each data unit according to the priority score further comprises: Differential protection is performed on the data units added to the transmission queue; the differential protection includes at least one of the following: A forward error correction coding redundancy ratio is allocated to data units with priority scores higher than a first threshold; the redundancy ratio is positively correlated with the priority score. For data units with a priority score higher than the second threshold, retransmission is performed after packet loss is detected; the number of retransmissions is positively correlated with the priority score. Data units with priority scores higher than the third threshold are placed into a high-priority transmission queue, which has priority scheduling rights or is guaranteed minimum bandwidth when the network is congested.
13. A data transmission apparatus, comprising: The data acquisition module is used to acquire the perception-related data stream to be transmitted; The perception-related data stream includes multiple data units; The data evaluation module is used to evaluate each data unit based on at least one of the content dimension and the transmission dimension to obtain an evaluation result; the content dimension is related to at least one of the content importance and the acquisition quality of the data unit. The transmission dimension is related to at least one of the dependency weight of the data unit and the network transmission cost in the current network state; The priority score calculation module is used to determine the priority score of each data unit based on the evaluation results. The transmission scheduling module is used to perform differentiated transmission scheduling on each of the data units according to the priority score.
14. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 12.