An AI workflow processing method and system for intelligent hardware
Through the collaborative work of edge devices and cloud service devices, and the use of hybrid modal analysis and hierarchical conflict arbitration models, the integration efficiency and conflict detection problems in multi-person collaborative input are solved, and efficient processing of multimodal data and reliable execution of instructions are achieved, thereby improving collaboration efficiency and the stability of live broadcast services.
Patent Information
- Application Number
- CN202511033069.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-25
AI Technical Summary
Existing technologies have problems with low information integration efficiency, inaccurate conflict detection, and difficulty in multimodal information fusion when processing collaborative input from multiple people, which limits the efficiency and fluency of collaboration.
Through the collaborative work of edge devices and cloud service devices, a hybrid modal parsing model and a hierarchical conflict arbitration network are used to process multi-person collaborative input, generate a standardized integrated data set, and use a multi-dimensional feature extraction and analysis model to generate a priority weight table and a cross-platform streaming routing table. The reinforcement learning algorithm is used to optimize the diversion strategy, and an intelligent retry strategy is generated for streaming interruption scenarios to achieve real-time optimization.
It improves the efficiency of multi-person collaboration and the stability of live broadcast services, realizes the efficient integration and precise processing of multimodal data, and ensures the reliable execution of instructions.
Smart Images

Figure CN120547367B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal data processing, and in particular to an AI workflow processing method and system for intelligent hardware. Background Art
[0002] In the current environment of booming digital collaboration and live broadcasting, edge devices and cloud service devices play a vital role. However, existing technologies still need to be improved when dealing with multi-person collaborative input processing and live broadcast-related command operations.
[0003] From the perspective of multi-person collaborative input processing, the collaborative systems on the market have problems with low integration efficiency and poor accuracy when integrating multi-modal input information such as voice, text, and images. For example, some systems are unable to efficiently handle a large number of concurrent multi-person inputs, resulting in information loss or processing delays, which seriously affects collaboration efficiency. In terms of conflict detection and processing, existing technologies are often unable to accurately identify conflicts between inputs from different users. For example, simple text duplication detection relies only on character matching and fails to deeply consider the similarity at the semantic level, resulting in some repeated inputs with similar semantics but different characters that cannot be effectively recognized, thereby causing a waste of collaborative resources. When fusing multi-modal input information, due to the lack of effective time alignment and semantic association algorithms, it is difficult to establish a close logical connection between voice, text, and images, and it is impossible to provide high-quality basic data for the subsequent generation of workflow-related data, which greatly limits the fluency and collaborative effect of multi-person collaboration. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0005] An AI workflow processing method for intelligent hardware, applied to cloud service equipment, is characterized by comprising: receiving workflow-related data sent by edge devices, the workflow-related data including the integration results of multi-person collaborative input information, work priority processing results, workflow diversion processing information, traffic and computing power allocation information, push interruption emergency processing information, live broadcast effect data and live broadcast setting optimization information; processing the integration results of multi-person collaborative input information to generate an integrated data set in a standardized format; processing the work priority processing results to generate a priority weight coefficient table, a dynamic priority adjustment rule set, and a priority conflict resolution method; The system processes the flow diversion processing information to generate a cross-platform streaming routing table, a flow diversion load balancing configuration file, and a historical diversion data statistical report; processes the traffic and computing power allocation information to generate a resource allocation optimization plan, a resource usage warning threshold configuration, and a cross-device resource scheduling instruction set; processes the flow interruption emergency processing information to generate an intelligent retry strategy configuration file, a backup flow node list, and an interruption cause analysis report; processes the live broadcast effect data and live broadcast setting optimization information to generate a real-time optimization instruction set and an effect optimization evaluation report; and sends the generated target format instruction information to the edge device for the edge device to perform subsequent operations.
[0006] An AI workflow processing method for intelligent hardware, applied to edge devices, is characterized by including: obtaining multi-person collaborative input information, including voice, text, and image information; processing the multi-person collaborative input information to generate workflow-related data, the workflow-related data including the multi-person collaborative input information integration results, work priority processing results, workflow diversion processing information, traffic and computing power allocation information, streaming interruption emergency processing information, and live broadcast effect data and live broadcast setting optimization information; sending the workflow-related data to a cloud service device; receiving target format instruction information sent by the cloud service device; confirming the target application based on the target format instruction information, and sending the target format instruction information to the target application for processing by the target application.
[0007] An AI workflow processing system for intelligent hardware, comprising: an edge device acquiring multi-person collaborative input information, including voice, text, and image information; processing the multi-person collaborative input information to generate workflow-related data, the workflow-related data including multi-person collaborative input information integration results, work priority processing results, workflow diversion processing information, traffic and computing power allocation information, streaming interruption emergency processing information, as well as live broadcast effect data and live broadcast setting optimization information; sending the workflow-related data to a cloud service device; the cloud service device processing the workflow-related data sent by the edge device to generate target format instruction information; the edge device receiving the target format instruction information sent by the cloud service device; confirming a target application based on the target format instruction information, and sending the target format instruction information to the target application for processing by the target application.
[0008] The present invention provides an AI workflow processing method for intelligent hardware. Edge devices collect collaborative input information, such as voice, text, and images, and integrate, prioritize, pre-process, and allocate resources locally using a multimodal fusion model. This data generates workflow-related data and uploads it to a cloud service device. The cloud service device, as the core processing unit, receives data from edge devices and uses a multi-dimensional feature extraction and analysis model to generate a standardized integrated dataset, a priority weight table, and a cross-platform streaming routing table. Specifically, a hybrid modal parsing model and a hierarchical conflict arbitration network are used to process collaborative input. Task priorities are dynamically adjusted using a three-dimensional weight matrix. The platform load assessment model and reinforcement learning algorithm are used to optimize the flow strategy, and an isolation forest algorithm is used to detect resource allocation anomalies. For streaming interruptions, the system generates intelligent retry strategies through feature extraction and impact factor calculation. To optimize live streaming performance, real-time command generation is achieved through risk assessment and parameter modification. The system ensures reliable command transmission using the MQTT protocol and TLS encryption. After receiving commands in the target format, the edge device identifies the target application through semantic matching and anomaly detection, ensuring precise command execution. The system builds a closed-loop workflow of "edge collection - cloud intelligence - edge execution". Through multi-model collaboration and dynamic optimization mechanisms, it improves the efficiency of multi-person collaboration and the stability of live broadcast business. It is suitable for smart hardware scenarios that require real-time multimodal data processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 A flowchart of an AI workflow processing method for intelligent hardware provided by an embodiment of the present invention when applied to a cloud service device;
[0010] Figure 2 A flowchart of an AI workflow processing method for intelligent hardware provided by an embodiment of the present invention when applied to an edge device;
[0011] Figure 3 A schematic diagram of a module of an AI workflow processing system for intelligent hardware provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0012] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention and are not intended to limit the present invention. Figure 1 To describe the AI workflow processing method and system of intelligent hardware according to the exemplary embodiment of the present application. In one embodiment, the present application also proposes an AI workflow processing method of intelligent hardware. Figure 1 As shown:
[0013] In an embodiment of the present application, an AI workflow processing method for intelligent hardware is applied to a cloud service device, such as Figure 1 As shown:
[0014] S101, receiving workflow-related data sent by an edge device.
[0015] In one implementation, the cloud service device receives the full workflow data sent by the edge device in real time through an adaptive communication protocol, establishes an edge-to-cloud data interaction link, and ensures the integrity and real-time nature of data transmission. Specifically, the edge device generates a structured integration of multiple voice, text, and image inputs, including user identification, input timing, and content conflict markers. A standardized interface is reserved to connect to the integrated data set of the edge device, providing a foundation for subsequent semantic mapping and format normalization.
[0016] Edge devices prioritize and resolve conflicts based on task urgency and user permissions. A three-dimensional weight matrix (task type, urgency, and user permissions) is constructed to provide input for the subsequent generation of dynamic priority adjustment rules. After labeling workflows, edge devices generate diversion platforms, push priorities, and historical diversion records. Platform interface parameters are analyzed to construct a three-dimensional routing matrix (platform type, interface parameters, and diversion rules), laying the foundation for generating cross-platform push routing tables.
[0017] Edge devices collect real-time traffic, computing power usage, historical allocation strategies, and resource load status. Business type characteristics and edge computing characteristics (such as real-time load monitoring frequency) are extracted to provide multi-dimensional data support for generating resource allocation optimization solutions. Edge devices also record the time, frequency, platform type, historical retry strategies, and backup node status of streaming interruptions. This interruption feature data is integrated to generate input for the interruption probability prediction model, providing a basis for generating intelligent retry strategy configuration files.
[0018] Edge devices collect audience interaction data, viewing duration, input source switching records, and current live broadcast layout parameters. This is input into the live broadcast optimization model, extracting the interaction-setting mapping relationship to provide data support for generating real-time optimization instruction sets. The cloud-based data processing framework is as follows: Edge device data formats are parsed using standardized protocols (such as collaborative editing protocols and cross-platform streaming protocols) to ensure compatibility. Integrity checks (such as timestamp matching and field missing detection) are performed on received data to trigger an abnormal data marking mechanism. Data is stored in a cloud-based distributed database by data type (collaboration, priority, diversion, etc.) and indexed for subsequent use.
[0019] S102: Process the integration result of the multi-person collaborative input information to generate an integrated data set in a standardized format.
[0020] In one embodiment, the text parsing process is used to perform semantic mapping and format normalization on multi-user input data, and a conflict detection model and a priority arbitration algorithm are introduced to achieve structured integration of input information. The hybrid modal parsing model (Hybrid Modal Parsing Model) integrates natural language processing (NLP) and computer vision (CV) technologies to support multimodal input parsing of speech, text, and images. Specifically, the input layer of the hybrid modal parsing model is multimodal data (speech audio stream, text string, image matrix). The feature extraction layer includes the following: voice wake-up detection (VAD), speech-to-text (ASR, using DeepSpeech2 neural network), sentiment feature extraction (Mel spectrum coefficients), word segmenter (based on BERT-Chinese pre-training model), part-of-speech tagging (BiLSTM+CRF layer), keyword extraction (TF-IDF and TextRank fusion algorithm), target detection (YOLOv5 identifies text areas in images), optical character recognition (OCR, using Tesseract engine), image semantic mapping (CNN+Transformer structure)
[0021] The structured mapping layer includes the following: a field mapper (which converts features into structured fields: user ID, timestamp, keyword vector), a format normalizer (which unifies encoding formats and punctuation rules, and supports automatic conversion between UTF-8 and GBK), and a conflict pre-detection unit (which preliminarily marks data with format anomalies). The model's output layer produces a structured dataset (JSON / Protobuf format). The Hierarchical Conflict Arbitration Network (HCAN) combines a rule engine with deep learning to identify and prioritize conflicts. The data access layer of the hierarchical conflict arbitration model accepts structured input data (user operation logs, workflow instruction sets). The conflict detection layer has the following functions: content duplication detection: semantic similarity calculation (SBERT model, cosine similarity threshold 0.7), edit distance algorithm (Levenshtein Distance, threshold 3), and version conflict marking (based on Git-like difference comparison).
[0022] Logical contradiction detection, event timing verification (workflow dependency DAG graph), command mutual exclusion rule base (e.g., "start live streaming" and "stop streaming" are mutually exclusive), context semantic consistency check (BERT semantic vector comparison). Permission conflict detection, permission matrix (RBAC model, user role-operation permission mapping table), operation permission verifier (based on ABAC attribute-based access control), conflict log recorder (records conflict type, occurrence time, and involved users). Priority arbitration layer, feature vector generation: user permission factor (0-10 points, administrator = 10, ordinary user = 5), task urgency (real-time task = 10, scheduled task = 3), input timeliness (time decay factor λ = 0.95 / minute).
[0023] The arbitration decision network has a multidimensional feature vector (12 dimensions) as input, a two-layer fully connected hidden layer (neurons 64→32, ReLU activation), and a priority score (0-100, Softmax normalization) as output. The conflict resolution strategy library includes: high priority overrides low priority (triggered when the score difference is ≥ 20), most recent priority (ordered by timestamp when the score difference is < 20), and manual intervention flagging (generating an alert when the conflict score is ≥ 80). The final output is the arbitration result (priority ranking table and conflict resolution log). A conflict detection model is introduced to identify duplicate content, logical inconsistencies, or authority conflicts (e.g., modification instructions from different users for the same workflow) in multiple inputs. Based on a priority arbitration algorithm, the input processing priority is determined based on user permission level, task urgency, and other factors. High-priority instructions are automatically marked and prioritized for integration.
[0024] Integrate with collaborative editing protocol standards, build a user permissions matrix, generate a standardized integrated dataset through a content versioning conversion model, and establish a version iteration update mechanism. Integrate with collaborative editing protocol standards (such as the WebSocket real-time collaboration protocol) to build a user permissions matrix that clearly defines different users' permissions for creating, modifying, and deleting workflow content. Through the content versioning conversion model, multiple versions of input content are mapped to a unified data structure, generating a standardized integrated dataset with version numbers and supporting historical version tracking.
[0025] Establish a version iteration update mechanism. After each collaborative input, the current data is automatically compared with the historical version, recording only the incremental changes to reduce data redundancy. Add timestamps and operation logs to the integrated data to ensure that data changes are traceable and meet audit requirements.
[0026] Using the session cycle as the time window, user identification, input timing, content tags, and conflict handling records are integrated to construct a multi-dimensional collaborative feature matrix, achieving full-dimensional association of input information. Using the user session cycle (e.g., the duration of a live broadcast) as the time window, user identification, input timing, content tags (e.g., "Switch input source for live broadcast" and "Push stream to TikTok"), and conflict handling records are integrated. A multi-dimensional collaborative feature matrix is constructed, encompassing the time dimension (input order), the spatial dimension (user geographic location or device identification), and the content dimension (keyword weight), achieving full-dimensional association of input information. Natural language processing (NLP) technology is used to analyze the semantic relevance of input content, linking scattered commands (e.g., "Switch input source" and "Start live broadcast") into a complete workflow logic chain. For low-weight or ambiguous input (e.g., incomplete commands), semantic completion is performed in conjunction with historical collaborative data to improve the integrity of the integrated data.
[0027] Missing content is completed using a conflict resolution model, and key information is extracted through a weighted distribution mechanism. This information is then integrated with the permissions matrix features to generate a standardized, integrated dataset. Machine learning algorithms are used to predict and complete missing input content (e.g., missing fields due to multiple simultaneous entries). Keyword weights are calculated using the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm, and weight distribution is dynamically adjusted based on the user permissions matrix to extract high-value, key information. The completed content is integrated with the permissions matrix features to generate a standardized, integrated dataset using a pre-set data model (e.g., JSON or Protobuf). This dataset includes structured fields (user information, input content, processing status), unstructured attachments (e.g., raw audio clips for speech-to-text conversion), and metadata (data generation time, processing algorithm version) for subsequent workflow scheduling.
[0028] Receive raw input data from edge devices → parse and resolve conflicts → generate standardized data sets → feed back to edge devices or store in cloud databases. Provide integrated input to the work prioritization module as a basis for prioritization; provide a standardized data interface to the workflow diversion module, supporting automatic classification and diversion by content tags (e.g., streaming content to the corresponding platform based on the "TikTok" tag).
[0029] S103: Process the work priority processing result to generate a priority weight coefficient table, a dynamic priority adjustment rule set, and a priority conflict resolution method.
[0030] In one implementation, a priority parsing process performs weight mapping and rule normalization on multi-task priority data. This process introduces a task urgency assessment model and a dynamic priority arbitration algorithm to achieve structured processing of priority policies. The Multi-Dimensional Priority Arbitration Model (MD-PAM), combining a rule engine with deep learning, supports priority calculation based on three dimensions: task urgency, user permissions, and task type. Specifically, it receives multi-task priority data, including task descriptions, urgency flags, user permissions, and historical priority records. The weight mapping layer quantifies urgency using a sigmoid function: f(x) = 1 / (1+e^(-k(x-x0))), where k = 0.5 and x0 is the task deadline difference (in minutes). Real-time tasks (such as live streaming) are marked as urgent. Permission mapping pre-sets the mapping between user roles and permission levels, e.g., administrator = 10, general user = 5, and guest = 1. Dynamic adjustment is supported through configuration files. Task type classification includes 12 built-in basic task types (e.g., live streaming, editing), with initial weights of 7 for live streaming and 5 for editing. User-defined types and weights are supported. Establish a priority conflict resolution rule of "Urgency > Authority > Task Type," uniformly map the three-dimensional weights to the [0, 1] interval, and record the timestamp and reason for weight adjustments. Generate structured priority data, including the task ID, the three-dimensional weight vector, and the priority ranking results.
[0031] Integrating with the workflow priority protocol standard, a three-dimensional weight matrix of task type, urgency, and user permissions is constructed. A standardized weight coefficient table is generated through a priority mapping conversion model, establishing a dynamic priority update mechanism. The Dynamic Weight Generation Model (DWGM) generates standardized weight coefficients based on the three-dimensional matrix of task type, urgency, and user permissions. Specifically, the model parses the workflow priority protocol (supporting ISO / IEC 27001), loads a predefined three-dimensional matrix template, and supports importing external weight configuration files. The matrix is initialized with 12 task types in the row dimension, five levels of urgency (P1-P5, corresponding to weights of 0.8-0.2) in the column dimension, and 10 levels of user permissions in the depth dimension. Task priority processing records from the past 30 days are integrated, missing weights are filled using linear interpolation, outlier weights are detected using the 3σ principle, and Gaussian filtering is used for smoothing. Weight updates are triggered when the match between task priority and execution results is less than 0.7. Parameters are adjusted using gradient descent with a learning rate of η=0.01. The matrix is automatically optimized at 2:00 AM daily, updating only weights that have changed by more than 5%. Manual fine-tuning by administrators is supported. Generates a standardized weight coefficient table, including a three-dimensional matrix, update log, and weight change interface.
[0032] Using the task lifecycle as a time window, the model integrates task type, urgency, user permissions, and historical priority processing records to construct a multi-dimensional priority feature matrix, achieving full-dimensional correlation of priority strategies. The Full Dimensional Correlation Model (FDCM) uses the task lifecycle as a time window and integrates multi-dimensional features to construct a correlation matrix. The task lifecycle is divided into stages such as creation and planning. By default, the time window granularity is 24 hours, and data from the most recent 10 cycles is retained. Basic features such as the one-hot encoding (12 dimensions) of task type, the one-hot vector (5 dimensions) of urgency, and user permissions (0-10) are extracted. These features are combined with temporal features such as task creation time and expected completion time, as well as correlation features such as dependent task ID and collaborating user ID. Conflict features such as the number of historical conflicts and the distribution of conflict types (e.g., permission conflicts account for 65%) are analyzed.
[0033] Features were Z-score normalized (mean μ = 0, standard deviation σ = 1), retaining features with a Pearson correlation coefficient |r| > 0.3. Dimensionality was then reduced to 32 dimensions using PCA. Matrix snapshots were generated hourly, with rows and columns representing task IDs and fused features. Feature importance was ranked using the random forest algorithm, and anomalous features were detected using the isolation forest algorithm. A two-layer LSTM model (64 neurons per layer, Adam optimizer, learning rate 0.001) was used to predict priority trends for the next four hours. A task association network was constructed using the Louvain algorithm, establishing connections when the feature cosine similarity was > 0.6.
[0034] A conflict prediction model complements missing priority rules, and a dynamic weight adjustment mechanism is used to extract key strategies. These strategies are then integrated with the three-dimensional weight matrix to generate a dynamic priority adjustment rule set and conflict resolution solution. The Conflict Pre-Resolution Model (CPRM) predicts potential conflicts and generates resolutions based on historical data. It receives current task priority data, a queue of 100 pending tasks, 100,000 historical conflict cases, and real-time resource usage data. The SBERT model calculates task similarity (threshold 0.7) and combines the resource contention index with the authority cross-coefficient to predict conflict probability using an XGBoost model (100 trees, maximum depth 5, learning rate 0.1). A CNN model (2-layer, 3×3 convolution kernel) is used to classify conflict types (authority, urgency, etc.), achieving a test set accuracy of 92.5%. The rule base is matched, such as prioritizing higher-level permissions in case of conflicting permissions and queueing P1 tasks in case of conflicting urgency. Historical cases with a cosine similarity greater than 0.8 are retrieved and solutions are adjusted based on the current scenario. Solutions are generated using the DQN algorithm (with a 100,000-byte experience replay buffer and a discount factor of 0.99), and efficiency is evaluated using a reward function (-10 to 10 points). Solution execution results are collected, and failures are rolled back and flagged. The rule base is automatically updated monthly based on the latest conflict data. New cases are added to the database daily, and cases with a similarity greater than 0.9 are merged. Cases unused for more than one year are marked as low-frequency. Conflict resolution is generated, including priority adjustment recommendations, resource scheduling plans, and execution steps.
[0035] The cloud service device receives priority data from edge devices, parses and normalizes it using the MD-PAM model, and constructs and dynamically updates a three-dimensional weight matrix using the DWGM model. The FDCM model integrates multidimensional features to construct a correlation matrix and predict priority trends. The CPRM model combines rules and cases to generate conflict resolution solutions, ultimately forming a priority weight coefficient table and a set of dynamically adjusted rules. The results are sent to the edge device for execution, and model parameters (such as urgency function parameters and matrix weights) are optimized based on the feedback, forming a closed-loop optimization mechanism.
[0036] S104: Process the workflow diversion processing information to generate a cross-platform streaming routing table, a diversion load balancing configuration file, and a historical diversion data statistical report.
[0037] In one implementation, protocol mapping and rule normalization are performed on multi-platform diversion data through a stream parsing process, and a platform load assessment model and a dynamic diversion arbitration algorithm are introduced to implement structured processing of diversion strategies. The Multi-Platform Shunt Structuring Model (MPSSM) combines protocol parsing and load assessment to implement standardized processing of diversion strategies. Receive multi-platform diversion data sent by edge devices, including target platform type, push priority, historical diversion records, and current load status. Use a stream parsing process to semantically map the protocols of different platforms (such as the push protocols of TikTok and WeChat Video Account) and convert them into a standardized format. Introduce a platform load assessment model to monitor the current bandwidth, server load, and other indicators of each platform in real time, and use exponential smoothing to predict short-term load trends (smoothing coefficient α = 0.3).
[0038] A dynamic traffic diversion arbitration algorithm was established, determining the traffic diversion strategy based on the weighted order of "platform load < streaming priority < historical success rate" (weight ratio 4:3:3). The diversion rules were normalized, mapping parameters from different platforms (such as streaming addresses and authentication keys) to a unified data structure. This generated structured traffic diversion strategy data, including fields such as target platform, priority weight, and real-time load threshold.
[0039] Integrate with cross-platform streaming protocol standards, construct a three-dimensional routing matrix based on platform type, interface parameters, and traffic diversion rules. Generate a standardized streaming routing table using a traffic diversion benchmarking conversion model, and establish a dynamic update mechanism for diversion strategies. The Cross-Platform Routing Generation Model (CPRGM) standardizes and dynamically maintains the routing table based on this three-dimensional matrix. Integrate with cross-platform streaming protocol standards (such as RTMP and SRT) and analyze the interface parameters of each platform (such as streaming port and buffer size).
[0040] A three-dimensional routing matrix was constructed, consisting of platform type, interface parameters, and traffic diversion rules: the row dimension represents platform type (eight mainstream platforms, including TikTok and Bilibili); the column dimension represents interface parameters (12 basic parameters, such as streaming address and authentication method); and the depth dimension represents the traffic diversion rules (e.g., "high-engagement live streams should be prioritized for streaming to TikTok"). A traffic diversion benchmarking conversion model was used to map the traffic diversion requirements of edge devices to specific nodes in the matrix, generating a standardized streaming routing table. A dynamic update mechanism for the traffic diversion strategy was established, triggering an update when the platform load changes by more than 20% or the historical traffic diversion success rate fluctuates by more than 15%. A reinforcement learning algorithm (Q-Learning) was used to optimize the routing matrix, with a reward function of "diversion success rate × 0.6 + load balancing × 0.4." A standardized streaming routing table was generated, containing the optimal streaming path, backup routes, and update log for each platform.
[0041] Using the workflow lifecycle as a time window, the platform type, interface parameters, diversion rules, and historical diversion processing records are integrated to construct a multi-dimensional diversion feature matrix, achieving full-dimensional correlation of the diversion strategy. The Full-Lifecycle Shunt Correlation Model (FL-SCM) uses the workflow lifecycle as a window to integrate multi-dimensional features. The entire lifecycle of the workflow from creation to completion is used as the time window (the default maximum window is 24 hours), and it is divided into three stages: initialization, execution, and termination. Basic features such as platform type and interface parameters are extracted and combined with real-time load data (bandwidth utilization, CPU occupancy); historical diversion processing records are integrated, including time series features such as the number of successes / failures and average streaming delay; natural language processing technology is introduced to extract keywords (such as "e-commerce live broadcast" and "education courses") from the workflow description text to generate content tag features.
[0042] Features were normalized (Z-score), retaining those with a Pearson correlation coefficient |r| > 0.25. Principal component analysis (PCA) was used to reduce the dimensionality to 20 dimensions, generating a multidimensional traffic flow feature matrix by hour, with rows and columns representing workflow IDs and fused features. A traffic flow association network was constructed based on a graph neural network (GNN), with nodes representing workflow IDs and edges representing feature similarity (connections were established when cosine similarity > 0.6). An LSTM model (2 layers, 32 neurons per layer) was used to predict traffic flow load trends for the next four hours.
[0043] A load prediction model is used to supplement missing traffic diversion rules. Key strategies are extracted through a dynamic traffic adjustment mechanism and integrated with the three-dimensional routing matrix features to generate a traffic diversion load balancing profile and historical traffic diversion data statistics. The Intelligent Shunt Load Optimization Model (ISLOM) combines prediction and dynamic adjustment to generate a load balancing solution. It receives a multi-dimensional traffic diversion feature matrix, real-time platform load data, and a historical traffic diversion rule library (containing 100,000 historical records). The Prophet model is used to predict platform load, accounting for seasonal factors (such as the evening live streaming peak) and holiday effects. Missing traffic diversion rules are supplemented through a load prediction model (XGBoost, with n_estimators=100, max_depth=4), achieving an accuracy of 88%. In conjunction with the dynamic traffic adjustment mechanism, when a platform's load exceeds a threshold (e.g., bandwidth utilization >70%), a traffic diversion policy adjustment is automatically triggered. Key strategies are extracted (e.g., "When Douyin's load is too high, transfer 30% of the streaming tasks to Bilibili") and integrated with the three-dimensional routing matrix features to generate a load balancing profile. The generated load balancing solution is simulated and evaluated with indicators including streaming success rate, average latency, and platform load standard deviation. Based on the execution feedback from edge devices (such as the actual streaming effect), the parameters of the load prediction model are updated weekly (learning rate 0.01).
[0044] The cloud service device receives the diversion data from the edge device, performs protocol mapping and load evaluation through the MPSSM model, and generates a structured diversion strategy. The CPRGM model constructs a three-dimensional routing matrix and dynamically updates it, generating a standardized push routing table to provide path planning for diversion. The FL-SCM model uses the workflow lifecycle as a window, integrates multi-dimensional features to construct an association matrix, and supports full-dimensional analysis of diversion strategies. The ISLOM model combines load prediction and dynamic adjustment to generate load balancing configuration files and historical statistical reports, which are ultimately sent to the edge device for execution. The edge device provides feedback on the execution results, and the cloud service device optimizes the parameters of each model accordingly (such as adjusting the smoothing coefficient of load evaluation and the update threshold of the routing matrix), forming a closed-loop optimization mechanism.
[0045] S105: Process the traffic and computing power allocation information to generate a resource allocation optimization plan, resource usage warning threshold configuration, and a cross-device resource scheduling instruction set.
[0046] In one embodiment, feature extraction is performed on the traffic and computing power allocation information to generate business type features, real-time traffic features, computing power occupancy features, historical allocation features, resource standardization features, and edge computing features, wherein the resource standardization features include unified parameter data after protocol mapping and cloud service resource allocation protocol benchmarking factors, and the edge computing features include real-time load monitoring frequency and node resource balancing indicators. The Multi-Dimensional Resource Feature Extraction Model (MR-FEM) supports the structured extraction of multi-dimensional features such as business traffic and computing power occupancy. Receive traffic and computing power allocation information sent by edge devices, including real-time traffic data, computing power occupancy logs, historical allocation strategies, and edge device load status. Use natural language processing (NLP) to parse the workflow description text, extract business type labels (such as "live streaming" and "file transfer"), and use one-hot encoding to generate a 10-dimensional feature vector;
[0047] Real-time network bandwidth utilization, peak traffic, and traffic fluctuations are collected, and the mean and standard deviation are calculated using a sliding window (5-minute window size). Parameters such as CPU usage, GPU load, and memory bandwidth are monitored, and periodic fluctuation characteristics are extracted using Fourier transforms. Cloud service resource allocation protocols (such as the OpenStack API) are integrated to map parameters of different edge devices (such as bandwidth units and computing power units) to a unified standard. The protocol benchmarking factor defaults to 1.0 and is dynamically adjusted based on device type (edge server factor = 1.2, terminal device factor = 0.8). Real-time load monitoring frequency (default 10 seconds / time) and node resource balancing metrics (the ratio of load standard deviation to mean, with a threshold of 0.3) are extracted for edge nodes. Z-score normalization is used to map all features to a range with a mean of 0 and a standard deviation of 1 to ensure comparability across different dimensions. A structured feature vector is generated, containing 32 features across six categories, for subsequent optimization modeling.
[0048] Feature extraction and processing are performed on optimization and scheduling data to generate load forecasting model features, dynamic adjustment mechanism features, cross-device scheduling features, and multi-objective optimization features. Load forecasting model features include service traffic time series mappings and a multi-dimensional feature matrix of computing power requirements; dynamic adjustment mechanism features include dynamic traffic threshold update strategies and computing power weight parameters; cross-device scheduling features include resource scheduling instruction sets and a device load balancing matrix; and multi-objective optimization features include a comprehensive resource allocation objective function and a cross-device quota allocation matrix. The Resource Optimization Scheduling Feature Model (ROS-FM) focuses on feature extraction for scheduling strategies such as load forecasting and dynamic adjustment. The Prophet time series model is used to predict service traffic time series, accounting for seasonal factors (such as daily live streaming peaks) and holiday effects. Model parameters include an annual seasonal period of 365 days, a weekly seasonal period of 7 days, and a configurable holiday list. A multi-dimensional feature matrix of computing power requirements is constructed, encompassing 12 dimensions, including service type, number of concurrent users, and data processing complexity. Principal component analysis (PCA) is used to reduce the dimensionality to 8.
[0049] The dynamic traffic threshold update strategy uses the exponential moving average (EMA) algorithm with a smoothing coefficient of α = 0.2. If real-time traffic exceeds the threshold of 110% three times in a row, the threshold is automatically raised by 5%. Computing power weight parameters are determined using the Analytic Hierarchy Process (AHP), with default weights of 0.4 for live streaming, 0.3 for data processing, and 0.3 for file storage. The resource scheduling instruction set includes three types of instructions: device wakeup, resource migration, and load balancing. The instruction format complies with the JSON-RPC 2.0 standard. The device load balancing matrix records the real-time load status (load rate, available resources) of each edge node and uses the Hungarian algorithm to solve the optimal allocation solution. The comprehensive resource allocation objective function is: F = w1 × (1 - load imbalance) + w2 × resource utilization + w3 × (1 - scheduling delay), where w1 = 0.4, w2 = 0.3, and w3 = 0.3. The cross-device quota allocation matrix is based on the Nash equilibrium theory from game theory to ensure a balance between fairness and efficiency in resource allocation across devices.
[0050] Based on service type characteristics, real-time traffic characteristics, computing power utilization characteristics, historical allocation characteristics, resource standardization characteristics, and edge computing characteristics, combined with load forecasting model characteristics, dynamic adjustment mechanism characteristics, cross-device scheduling characteristics, and multi-objective optimization characteristics, analysis and processing are performed to generate resource allocation anomaly identification information. This anomaly identification information characterizes the type of resource allocation anomaly, the time of occurrence, and the degree of correlation with multi-source data. This information forms a resource allocation status assessment that includes real-time allocation results, intelligent prediction trends, and the effectiveness of the scheduling policy. This assessment then generates resource allocation optimization plans, resource usage warning threshold configurations, and cross-device resource scheduling instruction sets. The Resource Allocation Anomaly Detection Model (RA-ADM) combines multi-source characteristics to achieve anomaly identification and status assessment. The Isolation Forest algorithm is used to detect resource allocation anomalies, with 100 trees, a subsample size of 256, and an anomaly score threshold of 0.7. Anomaly types include traffic surges (exceeding three standard deviations from the mean), computing power exhaustion (a node's load continuously exceeding 90% for 10 minutes), and allocation imbalances (load differences between devices exceeding 40%).
[0051] This tool integrates real-time allocation results, load forecast trends (for the next 30 minutes), and scheduling strategy effectiveness (historical success rate) to generate a three-dimensional evaluation matrix. Evaluation metrics include resource utilization (target ≥ 80%), load balancing (target ≤ 0.2), and scheduling latency (target ≤ 50ms). It also generates a resource allocation status assessment report, including the anomaly type, occurrence time, impact scope, and recommended solutions.
[0052] The Resource Optimization Decision Model (RO-DM) generates an executable plan based on the status assessment results. The resource allocation optimization plan is solved using a genetic algorithm with a population size of 50, 100 iterations, a crossover probability of 0.8, and a mutation probability of 0.1. The goal is to minimize the weighted sum of load imbalance and scheduling delay. Resource usage warning thresholds are configured based on historical data and business needs, such as setting the traffic threshold to 90% of the historical peak and the computing power threshold to 85% of the node's maximum capacity. Scheduling instruction generation follows the "proximity principle + load balancing" strategy. When a node's load exceeds 80%, cross-device resource migration is triggered, with a migration threshold of a load difference of >15%. The instruction set includes parameters such as the source device ID, target device ID, resource type, and migration amount, and is sent to edge devices via a message queue (such as Kafka).
[0053] Cloud service devices use the MR-FEM model to extract and standardize resource features, generating multidimensional feature vectors. The ROS-FM model extracts features from optimization and scheduling data, providing policy parameters such as load forecasting and dynamic adjustment. The RA-ADM model identifies anomalies based on multi-source features and evaluates resource allocation status, generating anomaly reports and evaluation results. Based on the status evaluation results, the RO-DM model generates resource optimization plans, warning threshold configurations, and cross-device scheduling instructions. Edge devices execute the scheduling instructions and provide feedback. Cloud service devices then optimize model parameters (such as adjusting anomaly detection thresholds and updating computing power weights), forming a closed-loop optimization mechanism.
[0054] S106: Process the streaming interruption emergency handling information to generate an intelligent retry strategy configuration file, a list of standby streaming nodes, and an interruption cause analysis report.
[0055] In one implementation, feature extraction and statistical analysis are performed on streaming interruption emergency response information to generate interruption time, frequency, platform, type, and historical retry strategy information, as well as backup node status information. The Push Interruption Feature Extraction Model (PIFEM) is used to structured extract multi-dimensional features related to streaming interruptions. This model receives streaming interruption emergency response information from edge devices, including interruption logs, historical retry strategies, and backup node status. Interruption time features analyze the specific time and duration of interruptions and extract daily, weekly, or monthly periodic patterns (e.g., frequent interruptions between 7 and 10 p.m.). Interruption frequency features calculate the number of interruptions per unit time and use a sliding window (with a window size of 1 hour) to calculate the mean and standard deviation of the frequency. Interruption platform features identify the target platforms where interruptions occur (e.g., TikTok and Bilibili) and record the proportion of interruptions on each platform. Interruption type features classify interruption types (e.g., network fluctuations, authentication failures, server overloads), and create a type distribution histogram. Historical retry strategy features analyze historical retry intervals, retry limit, and other parameters to evaluate strategy effectiveness. Backup node status features collect metrics such as the online status, bandwidth capacity, and historical success rate of backup streaming nodes. A structured feature vector is generated, comprising six categories of 24-dimensional features, such as interruption time distribution and platform vulnerability index.
[0056] The system processes interruption time, frequency, platform, type, and distribution characteristics of interruptions, historical retry strategies, and backup node status to generate interruption probability predictions, platform vulnerability assessments, interruption type association characteristics, and backup node availability assessments. The Push Interruption Assessment Model (PIAM) generates interruption predictions and platform assessments based on feature vectors. The XGBoost model predicts interruption probability, taking as input features such as interruption time, frequency, and platform. Parameter settings include: number of trees = 100, maximum depth = 5, and learning rate = 0.1. The model outputs the interruption probability for each platform within the next 30 minutes, with a threshold of 0.3 (a probability ≥ 0.3 triggers an alert).
[0057] A vulnerability assessment indicator system was constructed, including the platform's historical outage rate (weighted 40%), load fluctuation (30%), and service provider stability (30%). The analytic hierarchy process (AHP) was used to calculate each platform's vulnerability index (ranging from 0 to 1, with higher values indicating greater vulnerability). Association rule algorithms (such as Apriori) were used to identify correlations between outage types and triggering conditions, with a minimum support of 10% and a minimum confidence of 80%. For example, the confidence level for the association between "TikTok + evening peak" and "network fluctuation outage" reached 85%. An availability assessment model was established, taking as input the node's online time, bandwidth redundancy, and historical switchover success rate, and outputting an availability score (0-100). Nodes with a score below 60 were marked "unavailable" and automatically removed from the backup list.
[0058] Based on interruption probability prediction information, platform vulnerability assessment information, interruption type correlation characteristics, and backup node availability assessment information, anomaly data in the push interruption emergency response information is marked and filtered. This generates a push interruption anomaly data screening result, including the platform, type, link, time of occurrence, and frequency of the abnormal interruption. The Push Interruption Anomaly Detection Model (PIADM) is used to identify abnormal interruptions and quantify their impact. The Isolation Forest algorithm is used to detect abnormal interruptions with parameters: number of trees = 100, subsample size = 256, and anomaly score threshold = 0.7. Anomaly types include sudden high-frequency interruptions (exceeding three standard deviations from the mean) and rare platform interruptions (interruptions occurring on platforms with a historical incidence of less than 1%). The formula for calculating the impact factor of streaming interruptions is as follows: Impact Factor = Platform Vulnerability Index × 0.4 + Interruption Recovery Difficulty Coefficient × 0.3 + Historical Retry Strategy Effectiveness × 0.2 + Future Interruption Probability × 0.1, where the recovery difficulty coefficient is set based on the type of interruption (network fluctuation = 1.0, server overload = 1.5, authentication failure = 0.8).
[0059] The results of the abnormal data screening for push interruptions are integrated and quantified to generate a push interruption impact factor. This factor characterizes the impact of platform type on push interruptions, the difficulty of recovery from interruption types, the effectiveness of historical retry strategies, and future impact trends of push interruptions. This factor then generates an intelligent retry strategy configuration file, a list of backup push nodes, and an interruption cause analysis report. The Push Interruption Decision Model (PIDM) generates an executable plan based on the impact factor. A reinforcement learning (Q-Learning) algorithm is used to optimize retry parameters. The state space is composed of interruption type + platform + time, and the action space is composed of retry intervals (5s / 10s / 30s) and number of retries (1-5). The reward function is (1-interruption probability) × 0.6 + backup node utilization × 0.4, the learning rate is 0.01, and the discount factor is 0.95. For example, for a "TikTok + network fluctuation" interruption, a policy of "first retry interval 10s, maximum 3 retries" is generated.
[0060] Backup nodes are sorted in descending order of availability score, with priority given to nodes with a score of 80 or higher. Each node is assigned parameters: streaming address, authentication information, and recommended bandwidth (dynamically adjusted based on the node's historical performance). A decision tree algorithm is used to trace the root cause of the outage, taking as input an outage feature vector and outputting the three most likely causes (with ≥90% accuracy). The report includes an outage timeline, platform load curve, anomaly indicator annotations, and historical case matching results. The cloud service device uses the PIFEM model to extract streaming outage features and generate a structured vector. The PIAM model uses the feature vector to predict outage probability, assess platform vulnerability, and backup node availability, and outputs the assessment results. The PIADM model screens for abnormal outages and calculates impact factors, identifying high-risk outage scenarios. The PIDM model generates an intelligent retry strategy, a backup node list, and a cause analysis report based on the impact factors, which are then sent to the edge device. The edge device executes the retry strategy and provides feedback. The cloud service device then optimizes model parameters (such as updating the outage prediction threshold for XGBoost and adjusting the reward function weights for Q-Learning), forming a closed-loop optimization mechanism.
[0061] S107: Process the live broadcast effect data and live broadcast setting optimization information to generate a real-time optimization instruction set and an effect optimization evaluation report.
[0062] In one implementation, a live streaming optimization model analyzes and processes live streaming effect evaluation results, live streaming impact factors, and abnormal live streaming identification information to generate a live streaming optimization risk assessment value. Abnormal live streaming identification information includes the type of live streaming anomaly, the time of occurrence, and the degree of correlation with live streaming data. The Live Streaming Effect Multi-Dimensional Evaluation Model (LSE-MEM) is used to analyze live streaming data and generate risk assessment values. It receives live streaming effect data (audience interaction, viewing time, input source switching history), live streaming setting parameters (screen layout, input source configuration), and anomaly logs from edge devices. Live streaming effect features extract metrics such as interaction growth rate, median viewing time, and peak online users, and calculate the dynamic mean using a sliding window (window size of 15 minutes). Live streaming impact factors analyze the correlation between input source quality (resolution, bitrate) and interaction volume, and determine the impact weight using the Pearson correlation coefficient. Abnormal identification features identify live streaming anomaly types (stuttering, black screen, audio and video desynchronization), record the anomaly occurrence timestamp and duration, and construct an abnormal event sequence.
[0063] A risk assessment indicator system was constructed using the Analytic Hierarchy Process (AHP), with weights assigned to: decreased engagement (40%), decreased viewing time (30%), and abnormal frequency (30%). The risk assessment value was calculated using the following formula: Risk Value = 0.4 × Engagement Risk Factor + 0.3 × Viewing Time Risk Factor + 0.3 × Abnormality Risk Factor, where the risk factor ranges from 0 to 1, with higher values indicating higher risk. A live broadcast optimization risk assessment value (ranging from 0 to 100), an abnormality type distribution report, and a risk-time heat map were generated.
[0064] Based on the live broadcast optimization risk assessment, the multi-dimensional adaptive strategy parameter set within the live broadcast optimization model is processed to generate a revised set of optimization parameters. This set includes weight parameters for the live broadcast effect objective function, dynamic parameters for input source switching, and screen layout adjustment parameters. The Multi-Dimensional Adaptive Strategy Optimization Model (MASOM) adjusts live broadcast parameters based on the risk assessment. The default weights for the live broadcast effect objective function are 0.5 for engagement, 0.3 for viewing time, and 0.2 for retention rate, with dynamic adjustment supported. Dynamic parameters for input source switching include switching thresholds (e.g., a 15% drop in engagement triggers a switch) and candidate source priorities (sorted by historical performance). Screen layout adjustment parameters define layout templates (e.g., single-screen vs. multi-screen split-screen) and switching conditions (e.g., peak online user count exceeding a threshold). Parameter adjustments are triggered based on risk assessments: When the risk value exceeds 70, the optimization process begins. A Bayesian optimization algorithm is used to search for the optimal parameter combination, with the objective function F = -risk value + 0.2 × parameter adjustment margin, to avoid over-adjustment. If the decrease in interaction volume is strongly correlated with input source quality (correlation coefficient > 0.6), the input source switching threshold is lowered from 15% to 10%. An optimized parameter revision set is generated, containing each parameter's current value, revised value, and the rationale for the revision.
[0065] The optimized parameter correction set is parsed and converted to generate live streaming dynamic optimization results. These results are used to characterize the input source switching strategy, the live screen layout adjustment plan, and the audience interaction improvement effect. This results in turn generate a real-time optimization instruction set and an optimization effect evaluation report. The Live Streaming Dynamic Optimization Decision Model (LSDODM) converts the parameter correction set into an executable plan. This optimization parameter correction set is mapped to a specific execution strategy. For example, input source switching parameters generate an "Input Source Switching Strategy List," which includes source IDs, switching conditions, and a list of backup sources; screen layout parameters generate a "Layout Adjustment Plan," which includes layout template IDs and trigger conditions (e.g., viewing duration <5 minutes >30%). An LSTM model is trained using historical data to predict optimization results. The inputs are the correction parameters and historical performance data, and the outputs are metrics such as the increase in engagement and the rate of change in viewing duration. The model consists of two LSTM layers, each with 128 neurons, and uses the Adam optimizer with a learning rate of 0.001. The real-time optimization instruction set follows the JSON format and includes instruction type (input source switching, layout adjustment), target parameters, and execution priority. The effect optimization evaluation report includes a comparison of indicators before and after optimization, predicted effects, and confidence scores (based on LSTM prediction results).
[0066] The cloud service device uses the LSE-MEM model to analyze live streaming data, generating a risk assessment and anomaly report. The MASOM model adjusts adaptive strategy parameters based on the risk value and determines the optimal correction plan through Bayesian optimization. The LSDODM model converts the parameter correction set into a specific optimization strategy, predicts the results, and generates an instruction set and an evaluation report. The edge device executes the optimization instructions and provides feedback on the actual results (e.g., a 12% increase in engagement). The cloud service device then updates the LSE-MEM risk assessment weight and MASOM optimization algorithm parameters, forming a closed-loop optimization mechanism. The risk assessment threshold is 70 (out of 100, exceeding this threshold triggers optimization); the number of Bayesian optimization iterations is 50, and the convergence threshold is 0.05; the LSTM prediction confidence threshold is 0.7, with any value below this threshold marked as "requiring manual confirmation"; and the optimization instruction execution priority is as follows: abnormal lag > sudden drop in engagement > reduced viewing time.
[0067] S108: Send the generated target format instruction information to the edge device so that the edge device can perform subsequent operations.
[0068] In one implementation, mapping rules between business results and instruction types are established. For example, a traffic routing table generates a "streaming platform switching instruction," while a resource optimization plan generates a "computing power scheduling instruction." A templated encapsulation approach is used, with the instruction template containing fields such as instruction ID, type, target device, parameter list, execution priority, and timeout. MQTT is used for instruction transmission, supporting QoS level 2 reliability (ensuring that instructions are delivered at least once).
[0069] Dynamically adjust transmission parameters, using TCP persistent connections (latency ≤ 50ms) when network quality is good and switching to UDP with a retransmission mechanism (maximum retransmission count = 5) when network conditions fluctuate. Commands are encrypted using AES-256, with the key periodically negotiated and updated between the cloud service device and the edge device (update cycle = 24 hours). The transmission channel establishes a secure connection via TLS 1.3, supporting two-way authentication (the cloud service device and the edge device each hold their own CA certificates). Real-time monitoring of command transmission status includes metrics such as transmission success rate (target ≥ 99.9%), average latency (target ≤ 100ms), and packet loss rate (target ≤ 0.1%). If an edge device fails to transmit three consecutive commands, the reconnection mechanism is automatically triggered, the device is marked as "high risk," and the retransmission count is temporarily increased to 10.
[0070] like Figure 2 As shown in FIG, an AI workflow processing method for intelligent hardware is applied to edge devices, including:
[0071] S201, obtaining multi-person collaboration input information.
[0072] In one implementation, the Edge Device Multimodal Input Collection Model (ED-MICM) supports real-time acquisition and preprocessing of multimodal information, including voice, text, and images. The voice input module supports both a built-in microphone (sampling rate 44.1kHz, 16-bit bit depth) and an external microphone (supporting XLR / USB interfaces), with automatic gain control (AGC) and noise suppression (NS) features, with a noise suppression threshold set to 40dB. The text input module supports touchscreen keyboards, physical keyboards, and external USB keyboards, using a unified UTF-8 character encoding and a target input latency of ≤50ms. The image input module supports both a built-in camera (resolution 1080p / 30fps) and an external USB camera, using RGB image format, autofocus, and light compensation. A circular buffer is used to store real-time input data. The voice buffer size is 1024 frames (10ms per frame), and the image buffer supports 5-frame deep buffering to prevent data loss. Voice is segmented into 500ms units, text is segmented into paragraphs, and images are segmented into frames. The segment marker contains a timestamp and device ID.
[0073] The Edge Device Input Preprocessing Model (ED-IPM) is responsible for formatting and noise filtering raw data. Speech preprocessing uses the VAD (Voice Activity Detection) algorithm to remove silence segments, with a threshold of -30dB and a detection window of 10ms. Speech noise reduction uses a Weiner filter with a 1-second noise update period, suitable for scenarios with a signal-to-noise ratio (SNR) ≥ 5dB. Text preprocessing automatically corrects common input errors (such as pinyin correction and punctuation completion), using a correction dictionary containing 100,000 common words. The sensitive word filtering module supports custom lexicons and uses an AC automaton matching algorithm with a matching efficiency of ≤ 1ms per 1,000 words. Image preprocessing includes image grayscale conversion and histogram equalization to improve text recognition in low-light environments. Gaussian blurring is used to remove image noise, with a kernel size of 3×3 and a standard deviation of 1.0.
[0074] The Edge Device Multimodal Fusion Model (ED-MFM) enables the association and structuring of data from different modalities. Multimodal data timestamps are synchronized based on the system clock, with a maximum time deviation of 50ms allowed; any deviation beyond this triggers resynchronization. A dynamic time warping (DTW) algorithm is used to align speech and image frames, with a distance threshold of 0.2. Speech-to-text conversion utilizes a lightweight speech recognition model (such as MobileNet-SS, model size ≤10MB) with a word error rate (WER) of ≤8%. Image text recognition utilizes the Tesseract OCR engine, supporting the Chinese character set and achieving a recognition accuracy of ≥95%. Semantic association between text and speech / image is achieved through keyword matching, with a matching threshold of 0.6 (based on the TF-IDF algorithm).
[0075] S202, process the multi-person collaborative input information to generate workflow-related data, which includes the integration results of the multi-person collaborative input information, work priority processing results, workflow diversion processing information, traffic and computing power allocation information, push interruption emergency processing information, live broadcast effect data and live broadcast setting optimization information.
[0076] In one implementation, edge devices are responsible for processing multi-user inputs of voice, text, and images and generating integrated results. They receive voice (sampling rate 44.1kHz), text (UTF-8 encoding), and image (1080p / 30fps) input from multiple users, along with user IDs, device identifiers, and timestamps. The Levenstein distance algorithm is used to detect duplicate text inputs (threshold 3), and semantic similarity is calculated using the SBERT model (threshold 0.7). Voice input is segmented using the VAD algorithm and sorted by user permissions (administrator > general user) in conflict scenarios. Time windowing is performed using a 5-minute session cycle, integrating user input timing with content tags (e.g., "live switch" and "streaming configuration"). Based on TF-IDF weight predictions from historical collaboration data, auto-completion is triggered when the completion probability is ≥ 0.8. A structured integrated result is generated, including user operation logs, conflict resolution records, and version numbers (e.g., v1.0.20250621).
[0077] Prioritization is generated based on urgency and permissions. Urgency is quantified: real-time tasks (e.g., streaming interruptions) are weighted 10, and scheduled tasks are weighted 3. This is converted to a 0-1 value using the sigmoid function f(x) = 1 / (1 + e^(-0.5(x-10))). User permissions are as follows: administrator weight = 1.0, standard user weight = 0.5, and guest weight = 0.2. A three-dimensional matrix is constructed: task type (live streaming / editing / storage) × urgency × permissions, with a default weight ratio of 4:3:3. Weights are updated every 10 minutes based on task execution feedback (learning rate 0.05). A priority ranking table is generated, including task IDs, weight scores, and execution order (e.g., Task A with a priority of 85 takes precedence over Task B with a priority of 72).
[0078] Traffic diversion information is generated based on label classification and initial load assessment. Text classification uses a TextCNN model to identify workflow labels (e.g., "TikTok push" and "Bilibili live broadcast") with an accuracy of ≥90%. Labels for e-commerce live streaming are weighted at 0.8, while labels for general videos are weighted at 0.5. Average platform latency is calculated based on historical streaming data (over the past 24 hours). Platforms exceeding 500ms are marked as high load. Traffic diversion to high-load platforms is ≤30%, with priority given to TikTok (weighted at 0.6) and Bilibili (weighted at 0.3) by default. A traffic diversion preprocessing record is generated, including the target platform, priority, and preliminary load threshold (e.g., a TikTok bandwidth threshold of 2Mbps).
[0079] Resources are allocated based on business type and historical data. Business type characteristics: Live streaming services are allocated bandwidth weights of 0.7 and computing power weights of 0.6; file transfer services are allocated bandwidth weights of 0.3 and computing power weights of 0.4. Real-time load: CPU utilization > 80% triggers a frequency reduction strategy, and remaining memory is < 10% when the cache is cleared. The traffic threshold is set at 80% of the historical peak, and computing power is allocated based on task priority (high-priority tasks account for 40%). When local resources are insufficient, support is requested from neighboring devices (triggered when the load difference > 20%). Local records of resource allocation are generated, including traffic allocation plans (such as allocating 2Mbps bandwidth for Douyin streaming), computing power scheduling instructions (such as allocating a 4-core CPU to live streaming tasks), and resource status feedback (current bandwidth utilization and remaining computing power).
[0080] S203: Send workflow-related data to the cloud service device.
[0081] S204: Receive target format instruction information sent by the cloud service device.
[0082] In one embodiment, after receiving the instruction, the edge device first performs format verification (JSON syntax check, field integrity verification), and then distributes it to the corresponding module (such as streaming instruction → streaming module, resource instruction → resource scheduling module) through the instruction type routing table. If the format is incorrect, an error code (such as 400BadRequest) and error details are returned. After receiving the error feedback, the cloud service device automatically repackages the instruction and reissues it. The execution queue is sorted by instruction priority. High-priority instructions can preempt low-priority tasks (preemption threshold: priority difference ≥ 5); the instruction timeout is set (default = 30 seconds). If the timeout is not completed, the "execution timeout" status is returned, and the current progress is attached (such as "streaming switch progress: 60%"). The execution result feedback includes: instruction ID, execution status (success / failure / partial success), execution time, and result data (such as the connection status after the streaming address is switched); the feedback format is consistent with the instruction format and uses JSON-RPC encapsulation to facilitate unified parsing by cloud service devices.
[0083] S205 , confirming a target application based on the target format instruction information, and sending the target format instruction information to the target application for processing by the target application.
[0084] In one implementation, feature extraction and statistical analysis are performed on target format command information to generate command semantic features, target application matching features, interface protocol compatibility features, command type distribution features, application scenario ratio features, and device interface parameter features. A pre-trained BERT-Chinese model is used to extract a 768-dimensional semantic vector from the command text, identifying key actions (e.g., "switch input source," "start live streaming") and target objects (e.g., "TikTok," "input source 2"). Semantic similarity is calculated using cosine similarity, with a threshold of 0.6. Commands below the threshold are marked as semantically ambiguous.
[0085] Maintain a signature library for target applications (such as the TikTok streaming module and OBS live streaming tools), including each application's interface call rules and parameter requirements. Application matching is calculated using a combination of TF-IDF and rule matching, with a weighting ratio of 3:2. A command is considered compatible with the application when the matching degree is ≥0.7. Commands are parsed for protocol types (such as RTMP and HTTP) and parameter formats (such as JSON and Protobuf), verifying the integrity of protocol fields (missing mandatory fields ≤5%). Device interface parameter extraction: Analyze resource requirements such as CPU and bandwidth from commands and compare them with the current parameters of the edge device (e.g., current bandwidth is 2 Mbps, command requires 3 Mbps).
[0086] The command semantic feature information, target application matching feature information, interface protocol compatibility feature information, command type distribution feature information, application scenario ratio feature information, and device interface parameter feature information are processed to generate command matching probability information, application compatibility assessment information, command type association feature information, and interface parameter adaptation assessment information. The command matching probability model inputs semantic features and application matching features and uses a logistic regression algorithm to output the matching probability (range 0-1). The formula is: P = σ(w1 × semantic vector + w2 × application matching degree + b), where w1 = 0.6 and w2 = 0.4, and σ is a sigmoid function. The probability threshold is set to 0.5, and a value below this triggers secondary verification.
[0087] Interface protocol compatibility: Check the match between the command protocol and the target application's supported protocol (e.g., RTMP protocol support = 1, HTTP = 0.5). Device parameter compatibility: Calculate the difference between the command resource requirements and the device's current resources (e.g., CPU requirement difference = |command requirement 4 cores - device remaining 6 cores| / 6 cores). Compatibility = 1 - difference. Values below 0.3 are marked as parameter incompatibility. Create a correlation matrix between command types and application scenarios (e.g., the probability of the "streaming interruption retry" command on the TikTok platform is 0.7). Use the Apriori algorithm to mine frequent associations (minimum support = 0.2, confidence = 0.8).
[0088] Based on command matching probability information, application compatibility assessment information, command type association characteristics, and interface parameter adaptation assessment information, anomaly data within target format command information is marked and filtered, generating command anomaly data screening results, including the type of abnormal command, application scenario, interface link, time of occurrence, and severity. Semantic anomalies are identified when the semantic matching probability is less than 0.5 or keywords are missing (no target platform name); protocol anomalies are identified when the protocol format is incorrect (JSON syntax error) or required parameters are missing (an empty streaming address); and scenario anomalies are identified when the command type does not match the current scenario (e.g., sending a device sleep command during a live broadcast). Anomaly detection is performed using the isolation forest algorithm, with 50 trees, a subsample size of 128, and an anomaly score threshold of 0.6. Anomaly severity is quantified based on the anomaly type and impact, categorizing anomalies into three levels: minor (missing non-required parameters), severe (protocol incompatibility), and fatal (complete command semantic error).
[0089] The results of the instruction anomaly data screening are integrated and quantified to generate an instruction execution impact factor. The instruction execution impact factor represents the impact of instruction type on execution, the difficulty of application scenario adaptation, device interface compatibility, and future trends in instruction execution impact. The impact factor is calculated as follows: Impact Factor = 0.4 × Match Probability + 0.3 × Adaptability + 0.2 × Anomaly Severity + 0.1 × Historical Execution Risk. Where: Match Probability: Ranges from 0-1, output by the instruction matching probability model, reflecting the semantic and functional match between the instruction and the target application (e.g., the match probability for the "stream to TikTok" instruction is 0.9). Adaptability: Ranges from 0-1, taking a weighted average of interface protocol compatibility (e.g., RTMP protocol adaptability = 1) and device parameter adaptability (e.g., bandwidth requirement match = 0.8). Anomaly Severity: quantifies the impact of the anomaly on execution, with minor anomalies = 0.2, major anomalies = 0.6, and fatal anomalies = 0.9 (e.g., protocol format errors are considered major anomalies, with a corresponding value of 0.6). Historical execution risk: Based on the failure rate of similar instructions in the past 30 days, the risk value is 0.8 when the failure rate is ≥30% and 0.2 when the failure rate is <10%.
[0090] An impact factor of 0.7 or higher will execute normally. Between 0.4 and 0.7 will trigger an alert and recommend manual confirmation. If it's less than 0.4, execution will be automatically rejected (for example, if the command's semantics are ambiguous and the device parameters are incompatible, the impact factor is 0.3, resulting in direct rejection). Command execution parameters are optimized based on the impact factor. For example, when the bandwidth adaptability is 0.5, the streaming bitrate will be automatically reduced by 20% to match the current resources. Real-time monitoring is enabled when executing high-impact-factor commands (≥0.7), collecting device status every 5 seconds. Any anomalies detected will trigger a retry or switch to a backup solution (automatically switching to a backup node if streaming is interrupted).
[0091] like Figure 3 As shown, an AI workflow processing system for intelligent hardware includes:
[0092] The edge device obtains multi-person collaborative input information, including voice, text, and image information; processes the multi-person collaborative input information to generate workflow-related data, which includes the integration results of multi-person collaborative input information, work priority processing results, workflow diversion processing information, traffic and computing power allocation information, push interruption emergency processing information, and live broadcast effect data and live broadcast setting optimization information; sends the workflow-related data to the cloud service device; the cloud service device processes the workflow-related data sent by the edge device to generate target format instruction information; the edge device receives the target format instruction information sent by the cloud service device; confirms the target application based on the target format instruction information, and sends the target format instruction information to the target application for processing by the target application.
[0093] The AI workflow processing system proposed in this application utilizes a collaborative architecture between edge devices and cloud service devices to implement multimodal data processing and intelligent workflow scheduling. Edge devices collect collaborative input information, such as voice, text, and images, and integrate, prioritize, pre-process, and allocate resources locally using a multimodal fusion model. This generates workflow-related data and uploads it to the cloud service devices.
[0094] The cloud service device, serving as the core processing unit, receives data from edge devices and, through multi-dimensional feature extraction and analysis models, generates standardized integrated data sets, priority weight tables, and cross-platform streaming routing tables. Specifically, it utilizes a hybrid modal parsing model and a hierarchical conflict arbitration network to process collaborative input. Task priorities are dynamically adjusted using a three-dimensional weight matrix. Platform load assessment models and reinforcement learning algorithms are used to optimize traffic diversion strategies. Furthermore, an isolation forest algorithm is used to detect resource allocation anomalies. In the event of streaming interruptions, the system generates intelligent retry strategies through feature extraction and impact factor calculation. To optimize live streaming performance, real-time command generation is achieved through risk assessment and parameter correction. The system ensures reliable command transmission using the MQTT protocol and TLS encryption. After receiving commands in the target format, edge devices identify the target application through semantic matching and anomaly detection, ensuring precise command execution. This system establishes a closed-loop workflow of "edge acquisition - cloud intelligence - edge execution." Through multi-model collaboration and dynamic optimization mechanisms, it improves collaborative efficiency and the stability of live streaming services. It is suitable for intelligent hardware scenarios requiring real-time multimodal data processing.
[0095] A computing device comprising a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein, when the computer program instructions are executed by the processor, the device is triggered to execute an AI workflow processing method of any intelligent hardware.
[0096] The methods and / or embodiments in the embodiments of the present application can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. When the computer program is executed by a processing unit, the above-mentioned functions defined in the method of the present application are performed.
[0097] It should be noted that the computer-readable medium described in this application may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this application, a computer-readable medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0098] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0099] It will be apparent to those skilled in the art that the present application is not limited to the details of the exemplary embodiments described above, and that the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the present application is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be encompassed herein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.
Claims
1. An AI workflow processing method for intelligent hardware, applied to cloud service equipment, characterized in that: include: Receive workflow-related data sent by edge devices, including the integration results of multi-person collaborative input information, work priority processing results, workflow diversion processing information, traffic and computing power allocation information, push interruption emergency processing information, live broadcast effect data, and live broadcast setting optimization information; Process the integration results of multi-person collaborative input information to generate an integrated data set in a standardized format; Process the work priority results to generate a priority weight coefficient table, a dynamic priority adjustment rule set, and a priority conflict resolution solution; Process workflow diversion information to generate cross-platform streaming routing tables, diversion load balancing configuration files, and historical diversion data statistics reports; Process traffic and computing power allocation information to generate resource allocation optimization plans, resource usage warning threshold configurations, and cross-device resource scheduling instruction sets; Process the emergency handling information of streaming interruption, generate intelligent retry strategy configuration file, backup streaming node list, and interruption cause analysis report; Process live broadcast effect data and live broadcast setting optimization information to generate real-time optimization instruction sets and effect optimization evaluation reports; The generated target format instruction information is sent to the edge device for the edge device to perform subsequent operations.
2. The AI workflow processing method for intelligent hardware according to claim 1, characterized in that: Process the integration results of multi-person collaborative input information to generate an integrated data set in a standardized format, including: Through the text parsing process, semantic mapping and format normalization are performed on multi-user input data. Conflict detection models and priority arbitration algorithms are introduced to achieve structured integration of input information. Connect with collaborative editing protocol standards, build a user permission matrix, generate standardized integrated data sets through content version benchmarking and conversion models, and establish a version iteration and update mechanism; Using the session cycle as the time window, it integrates user identification, input timing, content tags, and conflict handling records to construct a multi-dimensional collaborative feature matrix, achieving full-dimensional correlation of input information. The conflict resolution model is used to complete missing content, and the weight distribution mechanism is used to extract key information, which is then integrated with the authority matrix features to generate an integrated dataset in a standardized format.
3. The AI workflow processing method for intelligent hardware according to claim 1, characterized in that: Process the work priority results to generate a priority weight coefficient table, a dynamic priority adjustment rule set, and a priority conflict resolution solution, including: Through the priority parsing process, multi-task priority data is weighted and normalized, and a task urgency assessment model and a dynamic priority arbitration algorithm are introduced to achieve structured processing of priority strategies. Connect with workflow priority protocol standards, build a three-dimensional weight matrix of task type, urgency, and user authority, generate a standardized weight coefficient table through a priority benchmarking conversion model, and establish a dynamic priority update mechanism; Taking the task lifecycle as the time window, integrating task type, urgency, user authority and historical priority processing records, a multi-dimensional priority feature matrix is constructed to achieve full-dimensional association of priority strategies; The missing priority rules are supplemented by the conflict prediction model, and the key strategies are extracted by combining the dynamic weight adjustment mechanism. Then, they are integrated with the three-dimensional weight matrix features to generate a dynamic priority adjustment rule set and conflict resolution solution.
4. The AI workflow processing method for intelligent hardware according to claim 1, characterized in that: Process workflow diversion information to generate cross-platform streaming routing tables, diversion load balancing configuration files, and historical diversion data statistics reports, including: Through the flow parsing process, protocol mapping and rule normalization are performed on the multi-platform diversion data. The platform load assessment model and dynamic diversion arbitration algorithm are introduced to achieve structured processing of diversion strategies. Connect to cross-platform streaming protocol standards, build a three-dimensional routing matrix of platform type, interface parameters, and diversion rules, generate a standardized streaming routing table through a diversion benchmark conversion model, and establish a dynamic update mechanism for diversion strategies; Using the workflow lifecycle as the time window, the platform type, interface parameters, diversion rules, and historical diversion processing records are integrated to construct a multi-dimensional diversion feature matrix, thus achieving full-dimensional correlation of diversion strategies. The missing diversion rules are supplemented by the load prediction model, and the key strategies are extracted by combining with the dynamic traffic adjustment mechanism. They are then integrated with the three-dimensional routing matrix features to generate diversion load balancing configuration files and historical diversion data statistical reports.
5. The AI workflow processing method for intelligent hardware according to claim 4, characterized in that: Processes traffic and computing power allocation information to generate resource allocation optimization plans, resource usage warning threshold configurations, and cross-device resource scheduling instruction sets, including: Perform feature extraction and processing on traffic and computing power allocation information to generate business type features, real-time traffic features, computing power usage features, historical allocation features, resource standardization features, and edge computing features. Resource standardization features include unified parameter data after protocol mapping and cloud service resource allocation protocol benchmarking factors. Edge computing features include real-time load monitoring frequency and node resource balancing indicators. Perform feature extraction and processing on the optimization and scheduling data to generate load prediction model features, dynamic adjustment mechanism features, cross-device scheduling features, and multi-objective optimization features. The load prediction model features include the business traffic time series mapping relationship and the multi-dimensional feature matrix of computing power requirements; the dynamic adjustment mechanism features include the dynamic update strategy of traffic thresholds and computing power weight parameters; the cross-device scheduling features include the resource scheduling instruction set and the device load balancing matrix; and the multi-objective optimization features include the comprehensive objective function of resource allocation and the cross-device quota allocation matrix. Based on business type characteristics, real-time traffic characteristics, computing power occupancy characteristics, historical allocation characteristics, resource standardization characteristics, and edge computing characteristics, combined with load prediction model characteristics, dynamic adjustment mechanism characteristics, cross-device scheduling characteristics, and multi-objective optimization characteristics, analysis and processing are performed to generate resource allocation anomaly identification information. Among them, the anomaly identification information is used to characterize the type of resource allocation anomaly, the time when the anomaly occurs, and the degree of correlation with multi-source data, forming a resource allocation status assessment result that includes real-time allocation results, intelligent prediction trends, and scheduling strategy effects, and then generating resource allocation optimization plans, resource usage warning threshold configurations, and cross-device resource scheduling instruction sets.
6. The AI workflow processing method for intelligent hardware according to claim 1, characterized in that: Process the emergency handling information of streaming interruption, generate intelligent retry strategy configuration file, backup streaming node list, and interruption cause analysis report, including: Perform feature extraction and statistical analysis on the emergency handling information of streaming interruptions to generate interruption time feature information, interruption frequency feature information, interruption platform feature information, interruption type distribution feature information, historical retry strategy feature information, and standby node status feature information; Processing interruption time characteristic information, interruption frequency characteristic information, interruption platform characteristic information, interruption type distribution characteristic information, historical retry strategy characteristic information, and standby node status characteristic information to generate interruption probability prediction information, platform vulnerability assessment information, interruption type association characteristic information, and standby node availability assessment information; Based on the interruption probability prediction information, platform vulnerability assessment information, interruption type correlation feature information, and backup node availability assessment information, the abnormal data in the push interruption emergency handling information is marked and filtered, and the abnormal data screening results of the push interruption are generated, including the platform, type, link, abnormal occurrence time and abnormal frequency of the abnormal interruption; The results of abnormal data screening of streaming interruptions are integrated and quantified to generate a streaming interruption impact factor. The streaming interruption impact factor is used to characterize the impact weight of the platform type on the streaming interruption, the difficulty of recovery from the interruption type, the effectiveness of the historical retry strategy, and the impact trend of future streaming interruptions, thereby generating an intelligent retry strategy configuration file, a list of backup streaming nodes, and an interruption cause analysis report.
7. The AI workflow processing method for intelligent hardware according to claim 6, characterized in that: Process live broadcast effect data and live broadcast setting optimization information to generate real-time optimization instruction sets and effect optimization evaluation reports, including: Based on the live broadcast optimization model, the live broadcast effect evaluation results, live broadcast influencing factors, and abnormal live broadcast identification information are analyzed and processed to generate a live broadcast optimization risk assessment value. The abnormal live broadcast identification information includes the type of live broadcast anomaly, the time when the anomaly occurred, and the degree of correlation with the live broadcast data. Based on the live broadcast optimization risk assessment value, the multi-dimensional adaptive strategy parameter set within the live broadcast optimization model is processed to generate an optimization parameter correction set, where the multi-dimensional adaptive strategy parameter set includes the live broadcast effect objective function weight parameter, input source switching dynamic parameter, and screen layout adjustment parameter; The optimization parameter correction set is parsed and converted to generate live dynamic optimization results. The live dynamic optimization results are used to characterize the input source switching strategy, live screen layout adjustment plan and audience interaction enhancement effect, and then generate real-time optimization instruction set and effect optimization evaluation report.
8. An AI workflow processing method for intelligent hardware, applied to edge devices, characterized in that: include: Obtain multi-person collaborative input information, including voice, text, and image information; Processing multi-person collaborative input information to generate workflow-related data, including the integration results of multi-person collaborative input information, work priority processing results, workflow diversion processing information, traffic and computing power allocation information, streaming interruption emergency processing information, live broadcast effect data, and live broadcast setting optimization information; Send workflow-related data to cloud service devices; Receive target format instruction information sent by the cloud service device; Confirm the target application based on the target format instruction information, and send the target format instruction information to the target application for processing by the target application, including feature extraction and statistical analysis of the target format instruction information to generate instruction semantic feature information, target application matching feature information, interface protocol compatibility feature information, instruction type distribution feature information, application scenario proportion feature information, and device interface parameter feature information; process the instruction semantic feature information, target application matching feature information, interface protocol compatibility feature information, instruction type distribution feature information, application scenario proportion feature information, and device interface parameter feature information to generate instruction matching probability information, application compatibility evaluation information, instruction type association feature information, and interface parameter adaptation evaluation information; based on the instruction matching probability information, application compatibility evaluation information, instruction type association feature information, and interface parameter adaptation evaluation information, mark and filter abnormal data in the target format instruction information to generate instruction abnormality data filtering results, including the type of abnormal instruction, application scenario, interface link, time of abnormality occurrence, and degree of abnormality; The results of instruction anomaly data screening are integrated and quantified to generate instruction execution impact factors. The instruction execution impact factors are used to characterize the impact weight of instruction type on execution, the adaptation difficulty of application scenarios, the compatibility of device interfaces, and the impact trend of future instruction execution. The target application is then confirmed based on the target format instruction information, and the target format instruction information is sent to the target application for processing by the target application.
9. An AI workflow processing system for intelligent hardware, characterized in that: include: Edge devices obtain collaborative input information from multiple people, including voice, text, and image information; Processing multi-person collaborative input information to generate workflow-related data, including the integration results of multi-person collaborative input information, work priority processing results, workflow diversion processing information, traffic and computing power allocation information, streaming interruption emergency processing information, live broadcast effect data, and live broadcast setting optimization information; sending workflow-related data to cloud service devices; The cloud service device processes the workflow-related data sent by the edge device and generates target format instruction information; The edge device receives the target format instruction information sent by the cloud service device; Confirm the target application based on the target format instruction information, and send the target format instruction information to the target application for processing by the target application, including feature extraction and statistical analysis of the target format instruction information to generate instruction semantic feature information, target application matching feature information, interface protocol compatibility feature information, instruction type distribution feature information, application scenario proportion feature information, and device interface parameter feature information; process the instruction semantic feature information, target application matching feature information, interface protocol compatibility feature information, instruction type distribution feature information, application scenario proportion feature information, and device interface parameter feature information to generate instruction matching probability information, application compatibility evaluation information, instruction type association feature information, and interface parameter adaptation evaluation information; based on the instruction matching probability information, application compatibility evaluation information, instruction type association feature information, and interface parameter adaptation evaluation information, mark and filter abnormal data in the target format instruction information to generate instruction abnormality data filtering results, including the type of abnormal instruction, application scenario, interface link, time of abnormality occurrence, and degree of abnormality; The results of instruction anomaly data screening are integrated and quantified to generate instruction execution impact factors. The instruction execution impact factors are used to characterize the impact weight of instruction type on execution, the adaptation difficulty of application scenarios, the compatibility of device interfaces, and the impact trend of future instruction execution. The target application is then confirmed based on the target format instruction information, and the target format instruction information is sent to the target application for processing by the target application.
Citation Information
Patent Citations
Cloud application processing method and device, equipment and storage medium
CN114554228A
AI multi-mode fusion interaction method, device, system and equipment
CN120179079A