Task processing method and device for cloud collaborative scheduling, equipment, medium and product

By acquiring contextual features in the visual network and using the contextual feature extraction model to select a cloud-based collaborative scheduling method, the problem of imperfect application of the visual network combined with large models is solved, thereby improving task processing efficiency and system adaptability.

CN121125702APending Publication Date: 2025-12-12CHINA MOBILE COMM LTD RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510599990.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

In existing technologies, visual networks are not yet perfect in combining with large models to train and apply industry models. The processing capabilities of small models on the edge are limited. How to achieve cloud-based collaborative scheduling to improve task processing efficiency and scenario adaptability is a key question.

Method used

By obtaining contextual features in the visual network scene, the contextual feature extraction model is used to obtain the context type and prediction confidence. Based on the context type and confidence, the cloud collaborative scheduling method is selected, including edge processing, edge-cloud collaborative processing and privacy data preprocessing, and the confidence and weight are dynamically adjusted for task allocation.

Benefits of technology

It improves task processing efficiency, enhances system flexibility and scenario adaptability, and maximizes dynamic computing power allocation and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121125702A_ABST
    Figure CN121125702A_ABST
Patent Text Reader

Abstract

The invention discloses a task processing method and device for cloud collaborative scheduling, articulated naturality web equipment, a medium and a product. The method comprises the following steps: acquiring a situation feature of an articulated naturality web task in an articulated naturality web scene; inputting the situation feature into a situation feature extraction model to obtain a situation type output by the situation feature extraction model and a prediction confidence coefficient corresponding to the situation type; wherein the prediction confidence is a deterministic evaluation index of the situation feature extraction model for the extracted situation type; and selecting a cloud collaborative scheduling mode to process the articulated naturality web task according to the situation type and the prediction confidence coefficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video communication technology, and in particular to a cloud-based collaborative scheduling task processing method, apparatus, video network device, medium, and product. Background Technology

[0002] As a new type of high-speed, high-definition, real-time video switching network, the video network has significant advantages in video communication and data transmission, but its integration with large models to train and apply industry models is still imperfect. Meanwhile, small edge models offer advantages such as flexible deployment and fast response speed, but their processing capabilities are limited.

[0003] How to achieve cloud-based collaborative scheduling, improve task processing efficiency, and enhance scenario adaptability is an urgent problem to be solved. Summary of the Invention

[0004] This application provides a cloud-based collaborative scheduling task processing method, apparatus, video network device, medium, and product.

[0005] The technical solution of this application embodiment is implemented as follows:

[0006] This application provides a cloud-based collaborative scheduling task processing method, the method comprising:

[0007] Obtain the contextual characteristics of visual network tasks in visual network scenarios;

[0008] The context features are input into the context feature extraction model to obtain the context type output by the context feature extraction model and the prediction confidence level corresponding to the context type; wherein, the prediction confidence level is the deterministic evaluation index of the context feature extraction model for extracting the context type;

[0009] Based on the scenario type and the prediction confidence level, a cloud-based collaborative scheduling method is selected to process the video network task.

[0010] In the above scheme, the step of performing cloud-based collaborative scheduling of the visual network task based on the scenario type and the prediction confidence includes:

[0011] The prediction confidence level is adjusted according to the scenario type to obtain the adjusted confidence level;

[0012] The visual network task is collaboratively scheduled in the cloud based on the scenario type and the adjusted confidence level.

[0013] In the above scheme, the cloud-based collaborative scheduling method includes one of the following scheduling methods:

[0014] When the scenario type belongs to the first type and the adjusted confidence level is within the first threshold range, the cloud-based collaborative scheduling method is the first method: the end device processes the video network task;

[0015] When the scenario type belongs to the second type and the adjusted confidence level is within the second threshold range, the cloud collaborative scheduling method is the second method: the end-side device and the cloud-side device work together to process the video network task, and the cloud-side device receives the processing result uploaded by the end-side device and verifies the processing result;

[0016] When the scenario type belongs to the third type and the adjusted confidence level is within the third threshold range, the cloud collaborative scheduling method is the second method: the end device and the cloud device work together to process the video network task, and the end device performs preprocessing operations on privacy data in the video network task; wherein, the confidence levels included in the first threshold range, the second threshold range and the third threshold range increase sequentially.

[0017] In the above scheme, the step of performing cloud-based collaborative scheduling of the video network task based on the scenario type and the adjusted confidence level includes:

[0018] When there are multiple video network tasks under the same scenario type, the multiple video network tasks are coordinated and scheduled in the cloud according to the adjusted confidence level, the latency weight, computing power weight and network weight corresponding to the same scenario type.

[0019] In the above scheme, obtaining the contextual features of the visual network task in the visual network scene includes:

[0020] Obtain multimodal data collected by end-side devices and / or edge computing nodes in the visual network scenario for the visual network task;

[0021] The contextual features are constructed based on the multimodal data.

[0022] The method in the above scheme further includes:

[0023] Obtain the sample context features of the sample visual network tasks in the aforementioned visual network scenario;

[0024] Based on the time dimension weights and spatial dimension weights of the sample context features, the deterministic evaluation index of the context feature extraction model for extracting sample context types is updated.

[0025] A cloud-based collaborative scheduling task processing device includes:

[0026] The acquisition module is used to acquire the contextual characteristics of visual network tasks in a visual network scenario;

[0027] The processing module is used to input the context features into the context feature extraction model to obtain the context type output by the context feature extraction model and the prediction confidence level corresponding to the context type; wherein, the prediction confidence level is the deterministic evaluation index of the context feature extraction model for extracting the context type;

[0028] The processing module is used to select a cloud-based collaborative scheduling method to process the video network task based on the scenario type and the prediction confidence level.

[0029] A video networking device includes a communication interface and a processor; wherein,

[0030] The communication interface is used to obtain the contextual characteristics of visual network tasks in a visual network scenario;

[0031] The processor is configured to input the context features into a context feature extraction model to obtain the context type output by the context feature extraction model and the prediction confidence level corresponding to the context type; wherein, the prediction confidence level is a deterministic evaluation index of the context feature extraction model for extracting the context type; and select a cloud-based collaborative scheduling method to process the video network task based on the context type and the prediction confidence level.

[0032] This application also provides a storage medium storing a computer program thereon, characterized in that the computer program, when executed by a processor, implements the steps of any of the above methods.

[0033] This application also provides a computer product, including a computer program, which, when executed by a processor, implements the steps of any of the above methods.

[0034] This application provides a cloud-based collaborative scheduling method, apparatus, video network device, medium, and product for task processing. The method includes: obtaining contextual features of a video network task in a video network scenario; inputting the contextual features into a contextual feature extraction model to obtain the context type output by the model and the corresponding prediction confidence level; wherein the prediction confidence level is a deterministic evaluation index of the contextual feature extraction model for extracting the context type; and selecting a cloud-based collaborative scheduling method to process the video network task based on the context type and the prediction confidence level. Thus, by collaboratively deciding on contextual features and confidence levels, a cloud-based collaborative scheduling method is selected to process the video network task, improving task processing efficiency and enhancing system flexibility and scenario adaptability. Attached Figure Description

[0035] Figure 1 This is a flowchart illustrating a cloud-based collaborative scheduling task processing method according to an embodiment of this application.

[0036] Figure 2 This is a schematic diagram of a cloud-based collaborative scheduling process for a task according to an embodiment of this application.

[0037] Figure 3 This is a schematic diagram of the structure of a cloud-based collaborative scheduling task processing device according to an embodiment of this application;

[0038] Figure 4 This is a schematic diagram of the structure of a video networking device according to an embodiment of this application. Detailed Implementation

[0039] The present application will now be described in further detail with reference to the accompanying drawings and embodiments.

[0040] In the current field of artificial intelligence, while large models possess powerful generalization and language understanding capabilities, they often fail to directly meet the high-precision and specialized requirements of specific industry applications due to the unique professional knowledge and language habits of those industries. Traditional fine-tuning training methods use fixed weights, failing to fully consider the characteristics of industry data and the differences in model performance on specific tasks. Furthermore, training with few samples makes it difficult to effectively capture key information, resulting in poor performance of industry-specific models.

[0041] As a new type of high-speed, high-definition, real-time video switching network, the video network has significant advantages in video communication and data transmission, but its integration with large models to train and apply industry models is still imperfect. Meanwhile, small edge models offer advantages such as flexible deployment and fast response speed, but their processing capabilities are limited.

[0042] How to achieve cloud-based collaborative scheduling, improve task processing efficiency, and enhance scenario adaptability is an urgent problem to be solved.

[0043] This application provides a cloud-based collaborative scheduling task processing method, applied to video network devices, such as... Figure 1 As shown, the method includes:

[0044] Step 101: Obtain the contextual features of the visual network task in the visual network scene.

[0045] In practical applications, video network scenarios include, but are not limited to, intelligent security monitoring, industrial monitoring, smart city management, intelligent traffic management, environmental protection and monitoring, and water conservancy and municipal management.

[0046] In a video network scenario, video network devices include end-side devices and servers. Stable connections between end-side devices and servers establish a video network communication environment, ensuring that the video network has high-speed, low-latency, and high-bandwidth characteristics, enabling efficient real-time data transmission.

[0047] In intelligent security monitoring scenarios, video network tasks include, but are not limited to, real-time intrusion detection and alarm, intelligent identification and identity verification through perimeter monitoring systems (such as infrared sensors and high-definition cameras).

[0048] In industrial monitoring scenarios, video network tasks include, but are not limited to, providing high-definition video coverage of production lines, monitoring production safety, and monitoring equipment operating status in real time.

[0049] In smart city management, the tasks of the video network include, but are not limited to, real-time monitoring of public safety incidents and coordination of urban resource allocation.

[0050] The tasks of the visual network in intelligent traffic management include, but are not limited to, using high-definition cameras and artificial intelligence (AI) algorithms to monitor traffic flow and violations, and dynamically adjust traffic light timings.

[0051] The tasks of the visual network in environmental protection and monitoring include, but are not limited to, deploying sensors and cameras to monitor air quality, water quality and noise pollution in real time, and providing timely warnings of environmental anomalies.

[0052] The tasks of the video network in water conservancy and municipal management include, but are not limited to, covering river monitoring, flood warning, and operation and maintenance of municipal facilities.

[0053] In practical applications, the contextual characteristics of visual network tasks in visual network scenarios include, but are not limited to, features obtained from the deep fusion and collaborative analysis of multimodal data of visual network tasks. Multimodal data includes, but is not limited to, one or more of the following: visual modal data, environmental parameter modal data, biometric modal data, device interaction and status module count, audio modal data, spatiotemporal correlation modal data, text and log modal data, and cross-domain fusion modal data.

[0054] Step 102: Input the context features into the context feature extraction model to obtain the context type and the prediction confidence corresponding to the context type output by the context feature extraction model; wherein, the prediction confidence is the deterministic evaluation index of the context feature extraction model for the extracted context type.

[0055] In practical applications, contextual features are input into the contextual feature extraction model for multimodal feature extraction. This yields the context type and its corresponding prediction confidence level, output by the model. The information output by the contextual feature extraction model provides a reliable basis for intelligent decision-making. The prediction confidence level is an evaluation metric for the certainty of the extracted context type, reflecting the model's certainty regarding the classification results.

[0056] Step 103: Select a cloud-based collaborative scheduling method to process the video network task based on the scenario type and prediction confidence.

[0057] In practical applications, different cloud-based collaborative scheduling methods correspond to different task allocation strategies / logics. Based on the scenario type and prediction confidence level, the appropriate task allocation strategy / logic is selected to process the video network tasks.

[0058] This application provides a cloud-based collaborative scheduling method for task processing. The method includes: obtaining contextual features of a video network task in a video network scenario; inputting the contextual features into a contextual feature extraction model to obtain the context type and the corresponding prediction confidence level output by the model; wherein the prediction confidence level is a deterministic evaluation index of the contextual feature extraction model for extracting the context type; and selecting a cloud-based collaborative scheduling method to process the video network task based on the context type and the prediction confidence level. Thus, by collaboratively deciding on contextual features and confidence levels, a cloud-based collaborative scheduling method is selected to process the video network task, improving task processing efficiency and enhancing system flexibility and scenario adaptability.

[0059] In some embodiments, step 101 obtains the contextual features of the visual network task in the visual network scene, including:

[0060] Obtain multimodal data collected by end-side devices and / or edge computing nodes in a visual network scenario for visual network tasks;

[0061] Contextual features are constructed based on multimodal data.

[0062] In practical applications, in video network surveillance scenarios, edge devices (such as surveillance cameras, sensors, etc.) and / or edge computing nodes are used to collect multimodal data (historical sample data and / or real-time sample data) of the target industry, including video, images, audio, and text. For example, in security monitoring scenarios, surveillance videos and personnel entry and exit records are collected; in industrial monitoring scenarios, equipment operation videos and sensor data are collected.

[0063] Based on multimodal data, contextual features are constructed, including: preprocessing the multimodal data to obtain preprocessed data; classifying and labeling the data of different modalities according to industry characteristics and / or task requirements to construct contextual features. Preprocessing includes, but is not limited to, cleaning and denoising the multimodal data collected by the visual network, including denoising and cropping for videos and images; for audio data, performing noise reduction and speech recognition to convert it into text; and for text data, performing operations such as typo correction and format standardization.

[0064] Here, multimodal data is dynamically sensed and collected using edge devices and / or edge computing nodes, and then multimodal data such as vision, audio, and environmental parameters are fused to construct contextual features. The resulting contextual features achieve cross-modal feature complementarity and improve the accuracy of contextual feature analysis.

[0065] In some embodiments, step 103 selects a cloud-based collaborative scheduling method to process the video network task based on the scenario type and prediction confidence, including:

[0066] The prediction confidence level is adjusted according to the scenario type to obtain the adjusted confidence level;

[0067] Video networking tasks are collaboratively scheduled in the cloud based on the scenario type and the adjusted confidence level.

[0068] In practical applications, the adjusted confidence level determines whether visual network tasks will be preferentially allocated to the edge and / or cloud for execution. This application, in the process of adjusting the predicted confidence level according to the scenario type to obtain the adjusted confidence level, can also set a dynamic task adjustment matrix. Based on the visual network scenario characteristics (i.e., scenario type) and the quantification of predicted confidence level, it automatically optimizes task scheduling and adjusts the confidence level to better adapt to different scenario requirements. Here, adjusting the confidence level includes, but is not limited to, adjusting the predicted confidence level to the confidence level threshold range corresponding to the scenario type, thereby enabling rapid decision-making and execution of visual network tasks, achieving dynamic computing power allocation, and improving resource utilization.

[0069] In practical applications, the scenario types include, but are not limited to, one or more of the following: latency-sensitive and privacy-sensitive types.

[0070] For example, based on contextual features (also known as visual network scene features), the corresponding contextual types are divided into A, B, and C categories:

[0071] Category A (High latency sensitive + low privacy): such as traffic light control (latency <10ms), industrial production line quality inspection (requires 200fps response), etc.

[0072] Category B (Medium latency sensitivity + Medium privacy): Community security (300ms tolerance), unmanned retail (user behavior data), etc.

[0073] Category C (Low latency sensitivity + high privacy): Medical image analysis, financial risk control, etc.

[0074] For example, based on the prediction confidence quantification (real-time calculation on the edge), three categories are divided: high, medium, and low.

[0075] Low confidence: New scenario samples (such as abnormal traffic trajectories during rainstorms), cross-modal data (video + sensor fusion), etc.;

[0076] Zhongzhixin: Variations of common scenarios (such as a 20% increase in traffic during morning rush hour), historically similar samples (90% feature overlap), etc.;

[0077] High confidence: Standardized scenarios (such as traffic flow at intersections at 18:00 on weekdays), continuous learning samples on the device side (it has been iterated more than 5 times), etc.

[0078] In some embodiments, the cloud-based collaborative scheduling method includes one of the following scheduling methods:

[0079] When the scenario type belongs to the first type and the adjusted confidence level is within the first threshold range, the cloud-based collaborative scheduling method is the first method: the end device processes the video network task;

[0080] When the scenario type is the second type and the adjusted confidence level is within the second threshold range, the cloud collaborative scheduling method is the second method: the end-side device and the cloud-side device work together to process the video network task, and the cloud-side device receives the processing result uploaded by the end-side device and verifies the processing result;

[0081] When the scenario type is the third type and the adjusted confidence level is within the third threshold range, the cloud collaborative scheduling method is the second method: the end device and the cloud device work together to process the video network task, and the end device performs preprocessing operations on privacy data in the video network task; wherein, the confidence levels included in the first threshold range, the second threshold range and the third threshold range increase sequentially.

[0082] In practical applications, the first type includes, but is not limited to, one or more of high latency sensitivity and low privacy. The second type includes, but is not limited to, one or more of medium latency sensitivity and medium privacy. The third type includes, but is not limited to, one or more of low latency sensitivity and high privacy. For example, the first threshold range is (0, 0.7), the second threshold range is [0.7, 0.85], and the third threshold range is (0.85, 1).

[0083] For example, in scenario type A high real-time, when the adjusted confidence level is within the first threshold range, such as <0.7 → direct output from the edge, the task allocation logic is as follows: if the camera detects running a red light (adjusted confidence level 0.65), the local lightweight model directly triggers the capture (8ms).

[0084] For example, when the scenario type is privacy in category B, the adjusted confidence level is within the second threshold range, such as 0.7-0.85 → edge verification. The task allocation logic is as follows: such as community camera identifying suspicious persons (adjusted confidence level 0.8), the edge cloud returns the comparison result within 300ms.

[0085] For example, when the scenario type is Class C high privacy, if the adjusted confidence level is within the third threshold range, such as >0.85, then the end-side encryption is performed. The task allocation logic is as follows: for example, when a hospital CT device performs initial screening for lung nodules (adjusted confidence level 0.92), the end-side completes 95% of the inference (without uploading the original image).

[0086] The edge task decision-maker schedules task allocation based on the context type and confidence threshold, and calculates the softmax entropy value for each frame to trigger a scheduling signal.

[0087] As can be seen, this application classifies task types based on latency sensitivity, privacy requirements, and confidence quantification, and designs end-edge-cloud collaborative scheduling logic to achieve rapid decision-making and execution of visual network tasks, thereby realizing dynamic computing power allocation and improving resource utilization.

[0088] In some embodiments, the above-mentioned cloud-based collaborative scheduling of visual networking tasks based on scenario type and adjusted confidence level includes:

[0089] When there are multiple video network tasks under the same scenario type, the multiple video network tasks are coordinated and scheduled in the cloud according to the adjusted confidence level, the latency weight, computing power weight and network weight corresponding to the same scenario type.

[0090] In practical applications, the latency weight is set to τ, the computing power weight to ε, and the network weight to ω.

[0091] For example, in a traffic scenario, τ = 0.7 (priority end side), and in a medical scenario, τ = 0.3 (allowing 300ms delay).

[0092] When there are multiple video network tasks (including but not limited to emergency tasks, batch tasks, elastic tasks, etc.) under the same scenario type, the multiple video network tasks are coordinated and scheduled in the cloud according to the adjusted confidence level, the latency weight, computing power weight and network weight corresponding to the same scenario type.

[0093] For example, for urgent tasks (such as Class A high real-time tasks): preemptive scheduling, such as traffic scenarios τ=0.7 (priority end-side), end-side nodes reserve computing power according to time.

[0094] For example, batch tasks (such as Class C high confidence): force ε = 0.6, utilize concentrated processing during off-peak hours such as at night, and use high-performance cloud clusters for inference.

[0095] For example, elastic tasks (such as confidence in Class B): dynamically adjusted according to network weight ω load, prioritizing the terminal side under the 4G network, and partially uploading to the cloud for inference under the 5G slice network.

[0096] As can be seen, this application sets latency weight, computing power weight, and network weight to realize priority edge dynamic scheduling queue.

[0097] In some embodiments, the above method further includes:

[0098] Obtain the sample context features of sample visual network tasks in visual network scenarios;

[0099] Based on the time and space dimension weights of the sample context features, the context feature extraction model is updated with a deterministic evaluation index (also known as the real-time threshold) for extracting the sample context type.

[0100] In practical applications, the deterministic evaluation index of the extracted sample context type is dynamically adjusted and updated based on the time dimension weight (T-Wight) and spatial dimension weight (S-Wight) of the sample context features. This allows for dynamic weight allocation and information fusion through the interaction between the current time embedding and historical time features.

[0101] In practical applications, time and spatial location features are extracted from historically stored video, text and other data. Time dimension weight (T-Wight) and spatial dimension encoding (S-Embed) are set, and real-time thresholds are calculated. Then, a spatiotemporal inverted index is built to quickly query the current threshold.

[0102] For example, a hierarchical time encoding is adopted, which decomposes time into season (spring, summer, autumn, winter and rainy season, snow season), month, weekday (weekday / weekend), day (0-24 hours, 5-minute slices), and holidays according to season (cosine encoding), weekday type (weekday / weekend binary vector), hour (24-dimensional one-hot), and holidays (binary flag). The 128-dimensional time series embedding is generated by LSTM fusion, and the time dimension weights are set.

[0103] The similarity between the current time step and historical patterns is calculated based on an attention mechanism, using the following formula:

[0104]

[0105] Where Q represents the current time embedding, and K / V is the historical time feature matrix, containing key (K) and value (V) vectors for each historical time step, used to provide contextual information. QK T Calculate the similarity between the current time step and all historical time steps to generate an attention weight matrix. (d is the dimension of Q / K) This prevents the inner product from becoming too large, which could cause the softmax gradient to vanish and maintain training stability. Softmax normalization converts the similarity score into a probability distribution, making the sum of the weights equal to 1, reflecting the difference in importance of different historical time steps to the current time step.

[0106] A geographic rasterization code is constructed, which divides the spatial location into grids. Scene features are aggregated in each grid using a graph neural network to generate a 256-dimensional spatial vector, resulting in the spatial dimension code S-Embed. The 256-dimensional vector is reduced to a 1-dimensional scalar through a linear transformation and normalized to a scaling factor S-Wight.

[0107] Dynamic update mechanism: Spatial coding is updated every 5 minutes by streaming real-time events (such as traffic accidents and weather changes).

[0108] The threshold is dynamically adjusted according to the following formula:

[0109] Real-time threshold = Baseline threshold × (T - Wight) × (S - Wight) × Historical performance coefficient

[0110] In this context, T-Weight and T-Wight refer to the same thing.

[0111] In some embodiments, the edge-cloud closed-loop feedback, in addition to using a memory-assisted module to dynamically adjust the threshold to obtain a real-time threshold, can also optimize the confidence of difficult examples. For example, the edge stores 1000+ low-confidence samples and uploads them to the cloud weekly to label difficult examples on a large model. This back-propagation optimizes the confidence calculation on the edge. During cloud labeling, labels are generated through multi-expert voting (3 labelers + large model validation), and back-propagation corrects the confidence estimation bias of the edge model.

[0112] As can be seen, the memory assistance module of this application extracts time and spatial location features from historically stored video, text and other data, sets time dimension weights and spatial dimension encodings, combines a dynamic threshold adjustment algorithm with hierarchical time encoding and geographic raster encoding, calculates the real-time threshold, and then constructs a spatiotemporal inverted index to quickly query the current threshold.

[0113] In a feasible scenario, such as Figure 2 As shown, the task processing flow of cloud-based collaborative scheduling includes:

[0114] Step 201: Setting up the video network environment.

[0115] Here, taking an intelligent security monitoring scenario as an example, a video network communication environment is built to connect devices such as surveillance cameras, access control systems, and servers. The video network collects data such as surveillance video and personnel entry / exit records, and quickly transmits it to the data processing center.

[0116] Step 202: Data collection and processing.

[0117] Here, the data processing center cleans and labels the collected video text data.

[0118] Step 203: Construct the edge-side basic decision-maker.

[0119] This basic decision-maker is also known as the basic threshold decision-maker.

[0120] Here, the edge-side basic threshold decision-maker includes, but is not limited to, the context feature extraction module, the task dynamic adjustment matrix, and the edge-side task decision-maker.

[0121] In practical applications, a lightweight edge-side model (based on convolutional neural networks or transformer networks, such as TinyLlama-1.1B) extracts real-time data features. These features are then combined with multimodal inputs from sensors, video streams, and other sources to construct a task description vector, thus building a contextual feature extraction module. For example, in a smart park scenario, the contextual feature extraction module outputs features such as object position, motion trajectory, and pose angle. The module also outputs prediction confidence.

[0122] Step 204: Construct a dynamic scheduling queue.

[0123] This dynamic scheduling queue is also known as the priority edge dynamic scheduling queue.

[0124] Here, when constructing the priority edge dynamic scheduling queue, the scheduling priority is constructed by using latency weight, computing power weight, and network weight.

[0125] Step 205: End-to-end closed-loop feedback.

[0126] Here, the edge-cloud closed-loop feedback includes the aforementioned difficult example feature extraction optimization and the aforementioned memory assistance module. This closed-loop flexibility improves the model's generalization ability and optimizes the model while completing model deployment (deploying a large model in the cloud to handle complex tasks and deploying a lightweight model on the edge to perform real-time inference).

[0127] The cloud-based collaborative scheduling task processing method provided in this application has the following beneficial effects:

[0128] (1) Task dynamic adjustment matrix and priority edge scheduling queue, efficient transmission and elastic scheduling, cloud network collaborative architecture + dynamic priority matrix, to achieve millisecond-level response, high privacy tasks and maximize resource utilization.

[0129] (2) Multimodal intelligent fusion, lightweight edge model + spatiotemporal dynamic threshold, adapts to the real-time decision-making needs of complex scenarios and speeds up decision-making.

[0130] (3) Closed-loop self-optimization capability, difficult example annotation feedback + spatiotemporal coding technology, continuously improve the generalization of the model.

[0131] This application aims to provide a method for collaborative processing of cloud-side and edge-side models based on visual networks, which solves the problems of weak edge computing power, low training accuracy of edge-side industry models, insufficient integration of visual networks and cloud-side models, lack of effective solutions for collaborative processing of large and small models, and imperfect application of visual networks in monitoring scenarios in related technologies.

[0132] This application also provides a cloud-based collaborative scheduling task processing device, such as... Figure 3 As shown, it includes:

[0133] The module 301 is used to obtain the contextual features of the visual network task in the visual network scene;

[0134] Processing module 302 is used to input context features into the context feature extraction model to obtain the context type and the prediction confidence corresponding to the context type output by the context feature extraction model; wherein, the prediction confidence is the deterministic evaluation index of the context feature extraction model for extracting the context type;

[0135] The processing module 302 is used to select a cloud-based collaborative scheduling method to process the video network task based on the scenario type and prediction confidence.

[0136] In some embodiments, the processing module 302 is used to adjust the prediction confidence based on the scenario type to obtain the adjusted confidence; and to perform cloud-based collaborative scheduling of the visual network task based on the scenario type and the adjusted confidence.

[0137] In some embodiments, the cloud-based collaborative scheduling method includes one of the following scheduling methods:

[0138] When the scenario type belongs to the first type and the adjusted confidence level is within the first threshold range, the cloud-based collaborative scheduling method is the first method: the end device processes the video network task;

[0139] When the scenario type is the second type and the adjusted confidence level is within the second threshold range, the cloud collaborative scheduling method is the second method: the end-side device and the cloud-side device work together to process the video network task, and the cloud-side device receives the processing result uploaded by the end-side device and verifies the processing result;

[0140] When the scenario type is the third type and the adjusted confidence level is within the third threshold range, the cloud collaborative scheduling method is the second method: the end device and the cloud device work together to process the video network task, and the end device performs preprocessing operations on privacy data in the video network task; wherein, the confidence levels included in the first threshold range, the second threshold range and the third threshold range increase sequentially.

[0141] In some embodiments, the processing module 302 is used to perform cloud-based collaborative scheduling of multiple video network tasks under the same scenario type, based on the adjusted confidence level, latency weight, computing power weight, and network weight corresponding to the same scenario type.

[0142] In some embodiments, the acquisition module 301 is used to acquire multimodal data collected by end-side devices and / or edge computing nodes in the visual network scenario for visual network tasks; the processing module 302 is used to construct contextual features based on the multimodal data.

[0143] In some embodiments, the obtaining module 301 is used to obtain sample context features of sample visual network tasks in the visual network scene; the processing module 302 is used to update the deterministic evaluation index of the context feature extraction model for extracting sample context types according to the time dimension weight and spatial dimension weight of the sample context features.

[0144] This application provides a video networking device, such as... Figure 4 As shown, the video networking device 400 includes: a communication interface 401 and a processor 402; wherein,

[0145] Communication interface 401 enables information exchange with end-side devices and servers;

[0146] The processor 402 is connected to the communication interface 401 to enable information interaction with the end-side device and the server, and to execute the methods provided by one or more technical solutions on the video network device side when running computer programs;

[0147] Memory 403 stores computer programs that can run on processor 402.

[0148] Communication interface 401 is used to obtain the contextual characteristics of visual network tasks in a visual network scenario;

[0149] Processor 402 is used to input context features into a context feature extraction model to obtain the context type and the corresponding prediction confidence level output by the context feature extraction model; wherein, the prediction confidence level is a deterministic evaluation index of the context feature extraction model for extracting the context type. Based on the context type and the prediction confidence level, a cloud-based collaborative scheduling method is selected to process the video network task.

[0150] In some embodiments, the processor 402 is configured to adjust the prediction confidence based on the context type to obtain an adjusted confidence; and to perform cloud-based collaborative scheduling of visual networking tasks based on the context type and the adjusted confidence.

[0151] In some embodiments, the cloud-based collaborative scheduling method includes one of the following scheduling methods:

[0152] When the scenario type belongs to the first type and the adjusted confidence level is within the first threshold range, the cloud-based collaborative scheduling method is the first method: the end device processes the video network task;

[0153] When the scenario type is the second type and the adjusted confidence level is within the second threshold range, the cloud collaborative scheduling method is the second method: the end-side device and the cloud-side device work together to process the video network task, and the cloud-side device receives the processing result uploaded by the end-side device and verifies the processing result;

[0154] When the scenario type is the third type and the adjusted confidence level is within the third threshold range, the cloud collaborative scheduling method is the second method: the end device and the cloud device work together to process the video network task, and the end device performs preprocessing operations on privacy data in the video network task; wherein, the confidence levels included in the first threshold range, the second threshold range and the third threshold range increase sequentially.

[0155] In some embodiments, the processor 402 is configured to perform cloud-based collaborative scheduling of multiple video network tasks under the same scenario type, based on the adjusted confidence level, latency weight, computing power weight, and network weight corresponding to the same scenario type.

[0156] In some embodiments, the processor 402 is configured to obtain multimodal data collected by end-side devices and / or edge computing nodes in a visual networking scenario for visual networking tasks; and to construct contextual features based on the multimodal data.

[0157] In some embodiments, the processor 402 is used to obtain sample context features of sample visual network tasks in a visual network scene; and update the deterministic evaluation index of the context feature extraction model for extracting sample context types based on the temporal and spatial dimension weights of the sample context features.

[0158] It should be noted that the specific processing procedures of communication interface 401 and processor 402 can be understood by referring to the above method, and will not be repeated here.

[0159] Of course, in practical applications, the various components in the video network device 400 are coupled together through the bus system 404. It can be understood that the bus system 404 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 404 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in... Figure 4 The general designated all buses as Bus System 404.

[0160] The memory 403 in this embodiment is used to store various types of data to support the operation of the video networking device 400. Examples of such data include any computer program used for operation on the video networking device 400.

[0161] The methods disclosed in the embodiments of this application can be applied to processor 402, or implemented by processor 402. Processor 402 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 402 or by instructions in the form of software. The processor 402 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 402 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in memory 403. Processor 402 reads the information in memory 403 and combines its hardware to complete the steps of the aforementioned method.

[0162] In an exemplary embodiment, the video networking device 400 may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the aforementioned method.

[0163] It is understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or both. Specifically, non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Sync Link Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memories described in the embodiments of this application are intended to include, but are not limited to, these and any other suitable types of memories.

[0164] In an exemplary embodiment, this application also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, such as a memory 403 storing a computer program. This computer program can be executed by a processor 402 of a video networking device 400 to complete the steps of the aforementioned method of the video networking device. The computer-readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM.

[0165] In an exemplary embodiment, this application also provides a computer product including a computer program that can be executed by a processor 402 for a video networking device 400 to perform the steps of the aforementioned method of the video networking device.

[0166] It should be noted that terms such as "first" and "second" are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0167] Furthermore, the technical solutions described in the embodiments of this application can be combined arbitrarily without conflict.

[0168] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application.

Claims

1. A cloud-based collaborative scheduling task processing method, characterized in that, The method includes: Obtain the contextual characteristics of visual network tasks in visual network scenarios; The context features are input into the context feature extraction model to obtain the context type output by the context feature extraction model and the prediction confidence level corresponding to the context type; wherein, the prediction confidence level is the deterministic evaluation index of the context feature extraction model for extracting the context type; Based on the scenario type and the prediction confidence level, a cloud-based collaborative scheduling method is selected to process the video network task.

2. The method according to claim 1, characterized in that, The cloud-based collaborative scheduling of the visual network task based on the scenario type and the prediction confidence includes: The prediction confidence level is adjusted according to the scenario type to obtain the adjusted confidence level; The visual network task is collaboratively scheduled in the cloud based on the scenario type and the adjusted confidence level.

3. The method according to claim 2, characterized in that, The cloud-based collaborative scheduling method includes one of the following scheduling methods: When the scenario type belongs to the first type and the adjusted confidence level is within the first threshold range, the cloud-based collaborative scheduling method is the first method: the end device processes the video network task; When the scenario type belongs to the second type and the adjusted confidence level is within the second threshold range, the cloud collaborative scheduling method is the second method: the end-side device and the cloud-side device work together to process the video network task, and the cloud-side device receives the processing result uploaded by the end-side device and verifies the processing result; When the scenario type belongs to the third type and the adjusted confidence level is within the third threshold range, the cloud collaborative scheduling method is the second method: the end device and the cloud device work together to process the video network task, and the end device performs preprocessing operations on privacy data in the video network task; wherein, the confidence levels included in the first threshold range, the second threshold range and the third threshold range increase sequentially.

4. The method according to claim 2, characterized in that, The cloud-based collaborative scheduling of the video network task based on the scenario type and the adjusted confidence level includes: When there are multiple video network tasks under the same scenario type, the multiple video network tasks are coordinated and scheduled in the cloud according to the adjusted confidence level, the latency weight, computing power weight and network weight corresponding to the same scenario type.

5. The method according to claim 1, characterized in that, The acquisition of contextual features of video network tasks in a video network scenario includes: Obtain multimodal data collected by end-side devices and / or edge computing nodes in the visual network scenario for the visual network task; The contextual features are constructed based on the multimodal data.

6. The method according to claim 1, characterized in that, The method further includes: Obtain the sample context features of the sample visual network tasks in the aforementioned visual network scenario; Based on the time dimension weights and spatial dimension weights of the sample context features, the deterministic evaluation index of the context feature extraction model for extracting sample context types is updated.

7. A cloud-based collaborative scheduling task processing device, characterized in that, include: The acquisition module is used to acquire the contextual characteristics of visual network tasks in a visual network scenario; The processing module is used to input the context features into the context feature extraction model to obtain the context type output by the context feature extraction model and the prediction confidence level corresponding to the context type; wherein, the prediction confidence level is the deterministic evaluation index of the context feature extraction model for extracting the context type; The processing module is used to select a cloud-based collaborative scheduling method to process the video network task based on the scenario type and the prediction confidence level.

8. A video network device, characterized in that, Includes communication interfaces and processors; among which, The communication interface is used to obtain the contextual characteristics of visual network tasks in a visual network scenario; The processor is configured to input the context features into a context feature extraction model to obtain the context type output by the context feature extraction model and the prediction confidence level corresponding to the context type; wherein, the prediction confidence level is a deterministic evaluation index of the context feature extraction model for extracting the context type; and select a cloud-based collaborative scheduling method to process the video network task based on the context type and the prediction confidence level.

9. A storage medium having a computer program stored thereon, characterized in that, When a computer program is executed by a processor, it implements the steps of any one of claims 1 to 6.

10. A computer product comprising a computer program, characterized in that, When a computer program is executed by a processor, it implements the steps of any one of claims 1 to 6.