Distributed reasoning industrial Internet of Things cloud edge collaboration method
By segmenting the deep learning inference model into sub-models and performing format conversion and containerized deployment, combining the WasmEdge runtime and Kafka scheduling center, cross-platform deployment and resource utilization problems are solved, efficient distributed inference is achieved, and the performance of the industrial Internet of Things is improved.
Patent Information
- Application Number
- CN202510596735.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-12
AI Technical Summary
Existing cloud-edge collaborative inference solutions have challenges in cross-platform deployment, resource occupancy and inter-node scheduling efficiency, making it difficult to meet the needs of high concurrency and low latency in heterogeneous edge systems.
The deep learning inference model is divided into multiple sub-models, and encapsulated into a lightweight WASM model through format conversion, operator detection and structured pruning. Combined with containerized deployment and WasmEdge runtime, the Kafka scheduling center is used for resource allocation and distributed inference to achieve multi-node load balancing and reliable scheduling.
It significantly reduces the memory and CPU usage of edge nodes, shortens loading and inference delays, improves the utilization rate of network and computing resources, and improves the throughput and overall inference rate of the industrial Internet of Things.
Smart Images

Figure CN120475025A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial Internet of Things, and in particular to a distributed reasoning industrial Internet of Things cloud-edge collaboration method. Background Art
[0002] With the rapid evolution of the Internet of Things (IoT) and edge computing, massive numbers of intelligent devices (such as cameras and sensors) are generating a large number of inference tasks at the edge of the network, placing higher demands on system real-time performance and resource utilization. However, when deployed on edge devices, traditional deep learning inference frameworks (such as PyTorch and TensorFlow) often face challenges such as large model sizes, complex dependencies, high startup and runtime latency, and high resource consumption. These limitations make it difficult to meet the high-concurrency, low-latency inference requirements in heterogeneous device and resource-constrained scenarios.
[0003] To address these challenges, academia and industry have proposed cloud-edge collaborative distributed inference. This approach distributes inference tasks to multiple edge nodes and coordinates their execution, reducing the load on the central cloud and improving overall system performance. However, existing cloud-edge collaborative inference solutions still face numerous challenges in cross-platform compatibility, lightweight modules, and efficient inter-node communication. First, heterogeneous edge nodes vary in operating system and hardware architecture, making unified deployment complex and incurring significant management overhead. Second, container- or virtual machine-based packaging often results in high memory usage and startup latency, exacerbating the strain on limited edge node resources. Finally, multi-node communication often relies on heavyweight middleware or RPC calls, which increases system coupling and network overhead, weakening the real-time and stability of distributed inference.
[0004] To address these issues, existing research attempts to optimize edge inference performance through model quantization, pruning, microservices encapsulation, and the use of lightweight RPC frameworks. However, these approaches typically only offer improvements in a single dimension, failing to simultaneously address cross-platform deployment simplicity, low runtime resource usage, and efficient inter-node scheduling. In particular, there is a lack of a universal runtime environment that combines container-like isolation and encapsulation with cross-architecture compatibility and low startup overhead to meet the demands of distributed inference in heterogeneous edge systems. Summary of the Invention
[0005] The technical problem to be solved by the present invention is how to simultaneously take into account the simplicity of cross-platform deployment, the reduction of memory and CPU usage of edge nodes during runtime, and the high efficiency of scheduling between edge nodes.
[0006] The present invention provides a distributed reasoning industrial Internet of Things cloud-edge collaboration method, wherein the industrial Internet of Things includes a central control node in the cloud, several heterogeneous edge nodes in the edge, and a Kafka scheduling center. The distributed reasoning industrial Internet of Things cloud-edge collaboration method includes: Step 1: The central control node obtains reasoning task information and loads a deep learning reasoning model. The deep learning reasoning model is divided into multiple sub-models according to the reasoning task information, and each sub-model processes a sub-task. Step 2: Perform format conversion, operator detection, and structured pruning on each sub-model to obtain multiple WASM models. Each WASM model and the WasmEdge runtime are encapsulated as a Docker container image and sent to the Kafka scheduling center. Step 3: The Kafka scheduling center calculates the resource availability index of each edge node and distributes the Docker container image to the edge node that meets the resource requirements according to the resource availability index of each edge node; Step 4: The edge node loads the corresponding version of the WASM model according to the obtained Docker container image and performs the inference task to obtain the intermediate feature vector. The edge node sends the intermediate feature vector to the Kafka scheduling center; Step 5: The central control node obtains the intermediate feature vector from the Kafka scheduling center and performs weighted summation to obtain a fused feature vector; In step 6, the central control node inputs the fused feature vector into the top sub-network of the deep learning inference model to output the final feature vector.
[0007] Compared with the existing technology, the present application has the following advantages: the present application divides the deep learning inference model into multiple sub-models according to the inference task, and combines containerized deployment with WasmEdge runtime to encapsulate each sub-model into a lightweight WASM model through row format conversion, operator detection and structured pruning, which significantly reduces the memory and CPU usage of the edge node, shortens the loading and inference latency, and improves the utilization of network and computing resources; then, each Docker container image is distributed to each edge node according to the resource allocation strategy for distributed inference, realizing multi-node load balancing and reliable scheduling, and improving the throughput and overall inference rate of the industrial Internet of Things.
[0008] In a possible implementation, the subtask processed by each sub-model in step 1 is represented as , where Indicates the subtask number, Indicates the version of the WASM model. represents the input data path, Indicates a timestamp, Indicates the The resource requirement level required by the subtasks processed by each sub-model is calculated as follows: ; Where, Indicates the The set of network layers contained in each sub-model; Indicates the The floating point operations of the network layer; Indicates the Memory required by the network layer; Indicates the weight of calculating floating-point operations. Indicates the The weight of the memory required by the layer network; The deep learning inference model loaded in step 1 is a PyTorch model.
[0009] In a possible implementation, step 2 specifically includes: Step 201: Switch the sub-model to inference mode and export the sub-model into an ONNX model in ONNX format. Step 202: Scan the ONNX model for operator compatibility to obtain the ONNX computation graph. ,in, Represents the set of operator nodes obtained by compatibility scanning; A collection of connections between operator nodes; Step 203: Determine whether each operator node is natively supported by the WasmEdge-ONNX toolchain; the judgment conditions are: ; Where, , Represents an operator node Type; Indicates the operation set version of the operator node, Indicates the native support table of the WasmEdge-ONNX toolchain, Represents an operator node The data type, , Represents an operator node rank, ; Represents an operator node Configuration properties, Indicates that the operator node is of type The set of legal attributes of Represents the shape derivation function determined by the operator node type; When the operator node If the judgment conditions are met, the operator node is marked Natively supported by the WasmEdge-ONNX toolchain, , return to step 203; if the operator node If the judgment condition cannot be met, the operator node Mark the operator node as one to be replaced and proceed to step 204; Step 204: Establishing a mapping process between the operator node to be replaced and the equivalent operator node; including: Let the equivalent operator node template library be a set of two-tuples , Indicates the operator node to be replaced, Represents new operator nodes with equivalent functionality; Using VF2 algorithm, from ONNX calculation graph Match with Isomorphic node instances ; Mapping satisfy ; Then perform the operator node replacement operation, the expression is: ; Represented as functionally equivalent new operator nodes in the ONNX computation graph After replacing the embedded instance in , return to step 203 until all operator nodes are traversed; Step 205: The central control node performs structured pruning on the ONNX model after compatibility scanning and replacement; In step 206, the central control node dynamically quantizes the ONNX model after structured pruning, and the expression is: ; ; Where, represents the raw floating point weights, is the quantized integer weight, is the offset, is the scaling factor; Step 207: Convert the dynamically quantized ONNX model into a WASM model in a bytecode format that complies with the WebAssembly specification; and package the WASM model and the WasmEdge runtime into a Docker container image.
[0010] Compared with the existing technology, the above technical solution can dynamically balance the efficiency relationship between conversion and transmission based on the deep learning inference model and the network and computing resources of the central control node, so that the sub-model can be streamed and converted, thereby improving the utilization of network and computing resources.
[0011] In a possible implementation, before packaging the WASM model and the WasmEdge runtime into a Docker container image in step 207, a security sandbox check is performed on the WASM model.
[0012] Compared with existing technologies, by performing security sandbox verification on the WASM model, the import interface, memory access boundaries and multi-threaded concurrent behavior of the WASM model are verified, ensuring that strict isolation from the host system can be achieved in the WasmEdge runtime to prevent out-of-bounds access or malicious code execution.
[0013] In a possible implementation, the calculation formula for the resource availability index of each edge node calculated by the Kafka scheduling center in step 3 is: ; Where, Indicates the The current CPU usage of edge nodes, Indicates the The current memory usage of edge nodes; is the preset CPU usage threshold. It is the preset memory usage threshold; is the weight factor, satisfying ; In step 3, distributing the Docker container image to the edge nodes that meet the resource requirements according to the resource availability indicators of each edge node specifically includes: The Kafka scheduling center stores the subtasks corresponding to multiple WASM models in the Kafka message queue, and determines the subtask information of each WASM model. Resource availability index of edge nodes Is it satisfied and , then The edge node accepts the The Docker container image corresponding to the WASM model.
[0014] Compared with the existing technology, the above technical solution can allocate network and computing resources for distributed reasoning tasks and execute distributed reasoning tasks on the edge nodes.
[0015] In one possible implementation, in step 4, the edge node loads the corresponding version of the WASM model according to the obtained Docker container image and performs the inference task to obtain the intermediate feature vector At the same time, the edge node also calculates the inference delay ,in, Indicates the loading delay of the WASM model loaded by the edge node. Indicates the computational latency of the WASM model. , is the model computational complexity, is the current CPU main frequency, To allocate the number of threads; Then, the edge node converts the intermediate feature vector Together with the node ID , Task Number , model version number , timestamp and node resource snapshots Encapsulated in the same message, serialized in binary compression format, and sent to the Kafka scheduling center.
[0016] Compared with the existing technology, the above technical solution can efficiently execute through multi-threaded scheduling and WASM bytecode, and the edge node completes model inference and returns results without significantly increasing energy consumption.
[0017] In a possible implementation, the central control node in step 5 obtains the intermediate feature vector from the Kafka scheduling center and generates the intermediate feature vector according to the timestamp. Perform merge sort in each time window A set of intermediate eigenvectors is gathered ; Then calculate the fusion weight of each edge node, the calculation formula is: ; Where, ; Then, the weighted sum of the intermediate feature vector set is performed to obtain the fused feature vector. The calculation formula is: .
[0018] Compared with the existing technology, the weighted fusion operation of the intermediate feature vectors of each edge node can balance the real-time load of each node.
[0019] In a possible implementation, in step 6, the central control node inputs the fused feature vector into the top-level sub-network output of the deep learning inference model to obtain the final feature vector expression: .
[0020] In one possible implementation, in step 5, the central control node obtains the intermediate feature vector from the Kafka scheduling center and calculates the fusion weight of each edge node. The intermediate feature vectors of the expected edge nodes that are not received within milliseconds are removed, and the edge nodes are marked, and then the remaining weights are renormalized; the expression is: , .
[0021] Compared with the existing technology, the timeout threshold is introduced in the central control node. Mechanism, when some nodes are delayed or packet lost, it automatically compensates by renormalizing the weights of the remaining nodes, effectively reducing the overall reasoning error so that the fusion calculation is based only on the set The nodes within the network can effectively avoid the impact of network fluctuations of a few nodes on the overall reasoning. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 This is a system architecture diagram for cloud-edge collaborative distributed reasoning in the industrial Internet of Things according to an embodiment of the present invention; Figure 2 This is a cloud-edge collaborative distributed reasoning process in an actual industrial Internet of Things environment according to an embodiment of the present invention. DETAILED DESCRIPTION
[0023] First, those skilled in the art should understand that these embodiments are merely used to explain the technical principles of the embodiments of the present application and are not intended to limit the scope of protection of the embodiments of the present application. Those skilled in the art may adjust them as needed to suit specific application scenarios.
[0024] In the description of the embodiments of this application, it should be noted that, unless otherwise specified or limited, the terms "connected" and "connection" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium. Those skilled in the art will understand the specific meanings of the above terms in the embodiments of this application based on the specific circumstances.
[0025] In the embodiments of the present application, unless otherwise expressly specified or limited, a first feature being "above" or "below" a second feature may mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. Furthermore, a first feature being "above," "above," and "above" a second feature may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is higher in level than the second feature. A first feature being "below," "below," and "below" a second feature may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is lower in level than the second feature.
[0026] The present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0027] See also Figures 1 and 2 As shown, the embodiments of this application disclose an industrial IoT cloud-edge collaboration method for distributed reasoning, which is suitable for efficient model deployment and reasoning execution in heterogeneous resource-constrained scenarios. This method builds an end-to-end distributed collaboration system from model conversion to system scheduling, from edge distributed reasoning to result return, and can effectively solve the problems of high resource usage, difficult platform adaptation, and slow response speed in existing reasoning methods.
[0028] The industrial Internet of Things includes a central control node in the cloud, several heterogeneous edge nodes on the edge, and a Kafka scheduling center. The industrial Internet of Things in this specific embodiment is applied to industrial quality inspection scenarios based on image recognition. The distributed reasoning industrial Internet of Things cloud-edge collaboration method includes: In step 1, the upper-layer application device initiates an inference task request to the central control node. The central control node obtains the inference task information and loads the ResNet-50 model based on the application scenario and inference task request. The ResNet-50 model includes operator nodes such as convolution (Conv), batch normalization (BatchNorm), activation (ReLU), pooling (Pool), and full connection (FC). Based on the inference task information and resource and network conditions, the central control node divides the ResNet-50 model into multiple sub-models: sub-model A (image front-end feature extractor, including the first 30 layers of Conv+Pool) and sub-model B (back-end classifier, including FC+Softmax). The metadata of sub-module A embeds the recommended number of threads and memory quota.
[0029] The subtask processed by each sub-model is represented as , where Indicates the subtask number, Indicates the version of the WASM model. represents the input data path, Indicates a timestamp, Indicates the The resource requirement level required by the subtasks processed by each sub-model is calculated as follows: ; Where, Indicates the The set of network layers contained in each sub-model; Indicates the The floating point operations of the network layer; Indicates the Memory required by the network layer; Indicates the weight of calculating floating-point operations. Indicates the Layer The memory required by the network layer for weights.
[0030] Step 2: Perform format conversion, operator detection, and structured pruning on sub-model A and sub-model B to obtain multiple WASM models. Each WASM model and the WasmEdge runtime are encapsulated as a Docker container image and sent to the Kafka scheduling center. This specifically includes: Step 201: Switch sub-model A and sub-model B to inference mode, and export sub-model A and sub-model B to a standard ONNX format ONNX model. In this embodiment, parameters can be specified during the export process to ensure that the ONNX model is compatible with the target runtime version. Step 202: Scan the ONNX model for operator compatibility to obtain the ONNX computation graph. ,in, Represents the set of operator nodes obtained by compatibility scanning; A collection of connections between operator nodes; Step 203: Determine whether each operator node is natively supported by the WasmEdge-ONNX toolchain; the judgment conditions are: ; Where, , Represents an operator node Type; Indicates the operation set version of the operator node, Indicates the native support table of the WasmEdge-ONNX toolchain, Represents an operator node The data type, , Represents an operator node rank, ; Represents an operator node Configuration properties, Indicates that the operator node is of type The set of legal attributes of Represents the shape derivation function determined by the operator node type, function Can be determined at compile time; Indicates that the operator node type matches the ONNX version; The data type and rank of the input and output tensors must conform to the supported set ; Indicates that the attribute configuration of the operator node must be in the legal attribute set of type t If there is an attribute Make Such as groups in group convolution Conv or axis in Gather , is also considered as no support; Indicates that the shape of the operator node output must be statically deducible at compile time. It is recorded as static_shape false, the operator node is judged to be unsupported; When the operator node If the judgment conditions are met, the operator node is marked Natively supported by the WasmEdge-ONNX toolchain, , return to step 203; if the operator node If the judgment condition cannot be met, the operator node Mark the operator node as one to be replaced and proceed to step 204; Step 204: Establishing a mapping process between the operator node to be replaced and the equivalent operator node; including: Let the equivalent operator node template library be a set of two-tuples , Indicates the operator node to be replaced, Represents new operator nodes with equivalent functionality; Using VF2 algorithm, from ONNX calculation graph Match with Isomorphic node instances ; Mapping satisfy ; Then perform the operator node replacement operation, the expression is: ; Represented as functionally equivalent new operator nodes in the ONNX computation graph The embedded instance in , and the connection method is reconnected according to the in-and-out edges of the original subgraph; after replacement, , return to step 203 until all operator nodes are traversed; taking GroupNorm as an example, if Contains only a single GroupNorm operator node , its pattern diagram is , replace the diagram Then the three operator node sequence Composition, satisfied after replacement , all input edges are redirected to the first node of the replacement subgraph, and all output edges are derived from the tail node of the replacement subgraph, thereby completing the equivalent replacement of operator functions while maintaining data flow consistency; In step 205, the central control node performs structured pruning on the ONNX model after the compatibility scan and replacement, removes redundant neuron channels, and retains high-importance branches. Structured pruning is an existing technology and will not be described in detail here. In step 206, the central control node dynamically quantizes the ONNX model after structured pruning, and the expression is: ; ; Where, represents the raw floating point weights, is the quantized integer weight, is the offset, is the scaling factor; This converts the original floating-point parameters into low-width integers to reduce the model size and improve running efficiency. This embodiment compresses the model volume by more than 30% while ensuring that the maximum error is controllable, thereby reducing the WASM model loading delay.
[0031] Step 207: Convert the dynamically quantized ONNX model into a WASM model in bytecode format that complies with the WebAssembly specification. In the conversion phase of this embodiment, the program automatically constructs a metadata object. And embed it into the generated WebAssembly module header, where the recommended number of threads is By total number of processor cores and the parallelism coefficient ∈(0,1] and round up to get: Memory quota limit Based on the total available memory and distribution ratio ∈(0,1] and round down to the integer, expressed in bytes: To ensure the module integrity during deployment, the system generates the compiled bytecode sequence wasm_bytes= Calculate SHA-256 to get the module hash ; Next, the WASM model is sandboxed to verify the WASM model's import interface, memory access boundaries, and multi-threaded concurrent behavior, ensuring strict isolation from the host system in the WasmEdge runtime to prevent out-of-bounds access or malicious code execution.
[0032] The central control node then packages the WASM model and the WasmEdge runtime into a Docker container image. This Docker container image pre-integrates the wasi-nn interface and execution environment, enabling plug-and-play deployment across various operating system architectures. This Docker container image is then distributed to all edge nodes registered in the system via a private image. Once distributed, each node pulls and initializes the local image, ready to load and execute when a task is triggered.
[0033] Step 3: The Kafka scheduling center calculates the resource availability index of each edge node and distributes the Docker container image to the edge node that meets the resource requirements based on the resource availability index of each edge node; specifically, it includes: In step 301, the Kafka scheduling center calculates the resource availability index of each edge node using the following formula: ; Where, Indicates the The current CPU usage of edge nodes, Indicates the The current memory usage of edge nodes; is the preset CPU usage threshold. It is the preset memory usage threshold; is the weight factor, satisfying ; Step 302: The Kafka dispatch center stores the subtasks corresponding to the multiple WASM models into the Kafka message queue, and determines the subtask information of each WASM model. Resource availability index of edge nodes Is it satisfied and , then The edge node accepts the The Docker container image corresponding to the WASM model.
[0034] Step 4: The edge node loads the corresponding version of the WASM model according to the obtained Docker container image and performs the inference task to obtain the intermediate feature vector. The edge node sends the intermediate feature vector to the Kafka scheduling center. Specifically, the process includes: The edge node loads the corresponding version of the WASM model according to the obtained Docker container image and performs the inference task to obtain the intermediate feature vector At the same time, the edge node also calculates the inference delay ,in, Indicates the loading delay of the WASM model loaded by the edge node. Indicates the computational latency of the WASM model. , is the model computational complexity, is the current CPU main frequency, To allocate the number of threads; Then, the edge node converts the intermediate feature vector Together with the node ID , Task Number , model version number , timestamp and node resource snapshots Encapsulated in the same message, serialized in binary compression format, and sent to the Kafka scheduling center.
[0035] In this embodiment, each edge node receives and loads the corresponding version of the WASM model and executes the corresponding subtasks when the WasmEdge is running. Among them, an edge node maps the input image to the local data buffer, performs forward convolution and pooling, and outputs the intermediate feature tensor. ; Another edge node tensor performs full connection and Softmax operation, outputting the category probability vector [x, y, z] (corresponding to "no defect", "crack", and "scratch" respectively); the above two edge nodes connect their respective results to the task number , model version number , timestamp and node resource snapshots Encapsulated in the same message and published to the Kafka scheduling center.
[0036] Step 5: The central control node obtains the intermediate feature vector from the Kafka scheduling center and performs weighted summation to obtain a fused feature vector; specifically, the following steps are performed: The central control node obtains the intermediate feature vector from the Kafka scheduling center and sends it according to the timestamp. Perform merge sort in each time window A set of intermediate eigenvectors is gathered ; Then calculate the fusion weight of each edge node, the calculation formula is: ; Where, ; Then, the weighted sum of the intermediate feature vector set is performed to obtain the fused feature vector. The calculation formula is: ; In addition, in order to deal with the situation where node messages are late or lost, the central control node will The intermediate feature vectors of the expected edge nodes that are not received within milliseconds are removed, and the edge nodes are marked, and then the remaining weights are renormalized; the expression is: , .
[0037] In step 6, the central control node inputs the fused feature vector into the top sub-network of the deep learning inference model to obtain the final feature vector, which is expressed as: .
[0038] The central control node converts the final feature vector After being packaged with relevant metadata, it is published to the Kafka dispatch center, allowing upper-layer application devices or decision modules to obtain refined inference results in real time. Through this complete asynchronous, reliable, and adaptively weighted transmission and fusion process, this application not only ensures low latency and high throughput for distributed inference, but also achieves robust and efficient cloud-edge collaborative inference in industrial IoT scenarios with heterogeneous nodes and unstable networks.
[0039] In order to ensure the long-term stable operation of the system, a sixth step is introduced: version management and security isolation mechanism. Each converted model file and WASM module is assigned a version number. With hash code The central control node records this information in a version table. When a grayscale update is needed, the system proportionally forwards tasks to the new version module nodes, expanding its scope of application once the metrics meet the requirements. Furthermore, the execution of WASM modules is protected by the WasmEdge sandbox mechanism. All memory accesses, system calls, and resource scheduling must pass through a virtual interface, preventing modules from unauthorized access to underlying system resources.
[0040] This application implements end-to-end distributed reasoning for industrial image defects. The cloud-edge collaborative framework can also be seamlessly extended to various AI reasoning tasks such as posture recognition, anomaly detection, and speech recognition. Simply replace the corresponding model and split the sub-modules as needed to reuse the same deployment and scheduling mechanism to meet the real-time and scalability requirements in various scenarios.
[0041] In summary, the cloud-edge collaborative distributed reasoning method for the Industrial Internet of Things (IIoT) provided by this invention forms a complete closed loop in terms of model conversion optimization, resource scheduling, cross-platform deployment, and secure execution. This significantly improves the deployment efficiency and responsiveness of distributed reasoning in actual industrial applications. In particular, in IoT scenarios with limited edge resources and strong node heterogeneity, the system architecture and method provided by this invention possess excellent scalability, versatility, and engineering practical value.
[0042] According to the above description in conjunction with the accompanying drawings, those skilled in the art will also understand that the embodiments of the present invention can also be implemented by software programs. Therefore, the present invention also provides a computer program product. The computer program product can be used to implement the present invention in conjunction with the accompanying drawings. Figure 1 The described method for cloud-edge collaborative computing offloading in the Industrial Internet of Things.
[0043] In the description of the embodiments of the present application, it should be noted that in the description of the present application, terms such as "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or component must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present application.
[0044] In the description of the present application, the description with reference to the terms "one embodiment", "some embodiments", "in the present embodiment", "specific example", or "some examples" means that the specific features, mechanisms, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, mechanisms, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.
[0045] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A distributed reasoning industrial Internet of Things cloud-edge collaboration method, wherein the industrial Internet of Things includes a central control node in the cloud, several heterogeneous edge nodes on the edge, and a Kafka scheduling center, characterized in that: The distributed reasoning industrial IoT cloud-edge collaboration method includes: Step 1: The central control node obtains reasoning task information and loads a deep learning reasoning model. The deep learning reasoning model is divided into multiple sub-models according to the reasoning task information, and each sub-model processes a sub-task. Step 2: Perform format conversion, operator detection, and structured pruning on each sub-model to obtain multiple WASM models. Each WASM model and the WasmEdge runtime are encapsulated as a Docker container image and sent to the Kafka scheduling center. Step 3: The Kafka scheduling center calculates the resource availability index of each edge node and distributes the Docker container image to the edge node that meets the resource requirements according to the resource availability index of each edge node; Step 4: The edge node loads the corresponding version of the WASM model according to the obtained Docker container image and performs the inference task to obtain the intermediate feature vector. The edge node sends the intermediate feature vector to the Kafka scheduling center; Step 5: The central control node obtains the intermediate feature vector from the Kafka scheduling center and performs weighted summation to obtain a fused feature vector; In step 6, the central control node inputs the fused feature vector into the top sub-network of the deep learning inference model to output the final feature vector.
2. The distributed reasoning industrial Internet of Things cloud-edge collaboration method according to claim 1 is characterized in that: The subtask processed by each sub-model in step 1 is expressed as , where Indicates the subtask number, Indicates the version of the WASM model. represents the input data path, Indicates a timestamp, Indicates the The resource requirement level required by the subtasks processed by each sub-model is calculated as follows: ; Where, Indicates the The set of network layers contained in each sub-model; Indicates the The floating point operations of the network layer; Indicates the Memory required by the network layer; Indicates the weight of calculating floating-point operations. Indicates the The weight of the memory required by the layer network; The deep learning inference model loaded in step 1 is a PyTorch model.
3. The distributed reasoning industrial Internet of Things cloud-edge collaboration method according to claim 1 is characterized in that: The step 2 specifically includes: Step 201: Switch the sub-model to inference mode and export the sub-model into an ONNX model in ONNX format. Step 202: Scan the ONNX model for operator compatibility to obtain the ONNX computation graph. ,in, Represents the set of operator nodes obtained by compatibility scanning; A collection of connections between operator nodes; Step 203: Determine whether each operator node is natively supported by the WasmEdge-ONNX toolchain; the judgment conditions are: ; Where, , Represents an operator node Type; Indicates the operation set version of the operator node, Indicates the native support table of the WasmEdge-ONNX toolchain, Represents an operator node The data type, , Represents an operator node rank, ; Represents an operator node Configuration properties, Indicates that the operator node is of type The set of legal attributes of Represents the shape derivation function determined by the operator node type; When the operator node If the judgment conditions are met, the operator node is marked Natively supported by the WasmEdge-ONNX toolchain, , return to step 203; if the operator node If the judgment condition cannot be met, the operator node Mark the operator node as one to be replaced and proceed to step 204; Step 204: Establishing a mapping process between the operator node to be replaced and the equivalent operator node; including: Let the equivalent operator node template library be a set of two-tuples , Indicates the operator node to be replaced, Represents new operator nodes with equivalent functionality; Using VF2 algorithm, from ONNX calculation graph Match with Isomorphic node instances ; Mapping satisfy ; Then perform the operator node replacement operation, the expression is: mapping satisfy ; Then perform the operator node replacement operation, the expression is: ; Represented as functionally equivalent new operator nodes in the ONNX computation graph After replacing the embedded instance in , return to step 203 until all operator nodes are traversed; Represented as functionally equivalent new operator nodes in the ONNX computation graph After replacing the embedded instance in , return to step 203 until all operator nodes are traversed; Step 205: The central control node performs structured pruning on the ONNX model after compatibility scanning and replacement; In step 206, the central control node dynamically quantizes the ONNX model after structured pruning, and the expression is: ; ; Where, represents the raw floating point weights, is the quantized integer weight, is the offset, is the scaling factor; Step 207: Convert the dynamically quantized ONNX model into a WASM model in a bytecode format that complies with the WebAssembly specification; and package the WASM model and the WasmEdge runtime into a Docker container image.
4. The distributed reasoning industrial Internet of Things cloud-edge collaboration method according to claim 3 is characterized in that: Before packaging the WASM model and the WasmEdge runtime into a Docker container image in step 207, a security sandbox check is performed on the WASM model.
5. The distributed reasoning industrial Internet of Things cloud-edge collaboration method according to claim 2 is characterized in that: In step 3, the Kafka scheduling center calculates the resource availability index of each edge node using the following formula: ; Where, Indicates the The current CPU usage of edge nodes, Indicates the The current memory usage of edge nodes; is the preset CPU usage threshold. It is the preset memory usage threshold; is the weight factor, satisfying + =1; In step 3, distributing the Docker container image to the edge nodes that meet the resource requirements according to the resource availability indicators of each edge node specifically includes: The Kafka scheduling center stores the subtasks corresponding to multiple WASM models in the Kafka message queue, and determines the subtask information of each WASM model. Resource availability indicators of edge nodes Is it satisfied and , then The edge node accepts the The Docker container image corresponding to the WASM model.
6. The distributed reasoning industrial Internet of Things cloud-edge collaboration method according to claim 1 is characterized in that: In step 4, the edge node loads the corresponding version of the WASM model according to the obtained Docker container image and performs the inference task to obtain the intermediate feature vector At the same time, the edge node also calculates the inference delay ,in, Indicates the loading delay of the WASM model loaded by the edge node. Indicates the computational latency of the WASM model. , is the model computational complexity, is the current CPU main frequency, To allocate the number of threads; Then, the edge node converts the intermediate feature vector Together with the node ID , Task Number , model version number , timestamp and node resource snapshots Encapsulated in the same message, serialized in binary compression format and sent to the Kafka scheduling center.
7. The distributed reasoning industrial Internet of Things cloud-edge collaboration method according to claim 6 is characterized in that: The central control node in step 5 obtains the intermediate feature vector from the Kafka scheduling center and calculates the intermediate feature vector according to the timestamp. Perform merge sort in each time window A set of intermediate eigenvectors is gathered ; Then calculate the fusion weight of each edge node, the calculation formula is: ; Where, ; Then, the weighted sum of the intermediate feature vector set is performed to obtain the fused feature vector. The calculation formula is: 。 8. The distributed reasoning industrial Internet of Things cloud-edge collaboration method according to claim 7 is characterized in that: In step 6, the central control node inputs the fused feature vector into the top sub-network output of the deep learning inference model to obtain the final feature vector expression: .
9. The distributed reasoning industrial Internet of Things cloud-edge collaboration method according to claim 7 is characterized in that: In step 5, the central control node obtains the intermediate feature vector from the Kafka scheduling center and calculates the fusion weight of each edge node. If the central control node The intermediate feature vectors of the expected edge nodes that are not received within milliseconds are removed, and the edge nodes are marked, and then the remaining weights are renormalized; the expression is: , .
Citation Information
Cited By
Distributed data processing method and system for color sorting equipment
CN121166384A
Flexible operator replacement experiment platform
CN121189409A
A flexible operator replacement experiment platform
CN121189409B
Inference method and device, inference cluster, storage medium and program product
CN121480725A