A method and device for IoT scene perception based on cloud-edge collaboration

By adopting cloud-edge collaboration methods in IoT scenario perception, instances are used to perceive and process IoT scenario data, the problem of insufficient perception real-time and accuracy in the prior art is solved, and more efficient perception processing and better adaptability are achieved.

CN114299370BActive Publication Date: 2025-05-16BEIJING UNIV OF POSTS & TELECOMM +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111478787.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-06
Publication Date
2025-05-16
Estimated Expiration
2041-12-06

AI Technical Summary

Technical Problem

The existing IoT scenario perception solutions are difficult to meet practical application needs in terms of real-time and accuracy, and edge server computing capabilities are limited, making it difficult to handle complex computing tasks.

Method used

The Internet of Things scene perception method based on cloud-edge collaboration is adopted, and instance perception is obtained by obtaining initial scene data, static instances, dynamic instances and abnormal instances corresponding to image data are determined, and this information is input into the deep neural network model deployed in the edge server for processing, combining the cloud server for local and global scene fusion.

Benefits of technology

It effectively reduces the delay in the perception processing of IoT scenarios, improves the adaptability and perception accuracy for high dynamic scenarios, and meets the needs of real-time and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114299370B_ABST
    Figure CN114299370B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for scene perception of the Internet of Things based on cloud-edge collaboration. The method includes: obtaining initial scene data to be perceived; the initial scene data includes image data and sensor data; performing instance perception on the initial scene data based on a dynamic instance perception model, and determining static instances, dynamic instances, and abnormal instances corresponding to the image data; inputting the perception results of the static instances, dynamic instances, abnormal instances, and the sensor data into a local multi-instance scene fusion model for processing, and obtaining a local scene output by the local multi-instance scene fusion model; the dynamic instance perception model is a deep neural network model that is pre-trained based on cloud-edge collaboration and deployed in an edge server. The method provided by the present invention can effectively reduce the processing delay of scene perception of the Internet of Things through a dynamic instance perception model, and improve the adaptability and perception accuracy of high dynamic scenes in the Internet of Things scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cloud service technology, and more particularly to a method and device for Internet of Things scene perception based on cloud-edge collaboration. In addition, the present invention also relates to an electronic device and a processor-readable storage medium. Background Art

[0002] In recent years, with the vigorous development of the fifth generation mobile communication technology (5G), cloud computing technology, and artificial intelligence technology, the concept of "Internet of Everything" has received more and more attention and development in the industry. Scene information perception is an important cornerstone for realizing the Internet of Everything technology, involving the collection and processing of video image data and various sensor data in specific scenes. The perceived scene information will be used for the upper-level applications of the Internet of Things system and the decision-making of managers. However, due to the rapid changes of various types of information in actual scenes, higher requirements are placed on the real-time and accuracy of perception.

[0003] At present, there are usually two methods to achieve scene perception of smart IoT: cloud computing and edge computing. Among them, the cloud computing model, as the most commonly used scene information perception method, meets the needs of scene information perception to a certain extent, but it is difficult to meet the actual application needs in terms of real-time, accuracy and resource utilization; the edge computing model processes data at the edge of the network, has a lower processing delay, and can also reduce the network load, but the computing power of the edge server is limited and cannot meet the needs of complex computing tasks in scene perception tasks. Therefore, how to design a real-time and accurate IoT scene perception solution based on cloud-edge collaboration, so that it can not only ensure high-quality perception of different types of information in the scene, but also meet the real-time requirements, has very important practical significance. Summary of the invention

[0004] To this end, the present invention provides an Internet of Things scene perception method and device based on cloud-edge collaboration to solve the defects of high limitations of Internet of Things scene perception solutions in the prior art, resulting in poor perception accuracy and real-time performance of specific instances in Internet of Things scenes.

[0005] In a first aspect, the present invention provides an IoT scene perception method based on cloud-edge collaboration, comprising: obtaining initial scene data to be perceived; wherein the initial scene data includes image data and sensor data;

[0006] Based on the dynamic instance perception model, the initial scene data is instance-aware, and the static instance, dynamic instance and abnormal instance corresponding to the image data are determined; wherein the dynamic instance perception model is a deep neural network model that is pre-trained based on cloud-edge collaboration and deployed in the edge server;

[0007] The static instance, the dynamic instance, the abnormal instance and the perception results of the sensor data are input into a local multi-instance scene fusion model for processing to obtain a local scene output by the local multi-instance scene fusion model.

[0008] Further, the static instance, the dynamic instance, the abnormal instance and the perception result of the sensor data are input into the local multi-instance scene fusion model for processing to obtain the local scene output by the local multi-instance scene fusion model, specifically including:

[0009] The static instance, the dynamic instance, the abnormal instance and the perception results of the sensor data are input into a local multi-instance scene fusion model, the static instance, the dynamic instance and the abnormal instance are combined and processed to obtain a corresponding three-dimensional scene, and the sensor data is matched to the three-dimensional scene to obtain the local scene.

[0010] Furthermore, instance perception is performed on the initial scene data based on a dynamic instance perception model to determine static instances, dynamic instances, and abnormal instances corresponding to the image data. Specifically, the initial scene data is input into the dynamic instance perception model for instance perception to obtain static instances, dynamic instances, and abnormal instances corresponding to the image data.

[0011] Furthermore, the IoT scene perception method based on cloud-edge collaboration also includes: fusing at least one of the local scenes on the cloud server side to realize the synthesis of the global scene based on the local scenes.

[0012] Furthermore, the IoT scene perception method based on cloud-edge collaboration, before obtaining the initial scene data to be perceived, further includes:

[0013] The dynamic instance perception model deployed in the edge server is pre-trained based on the deep neural network model pre-deployed in the cloud server, so as to share some target parameters of each dynamic instance perception model in each edge server during the training process of the dynamic instance perception model.

[0014] In a second aspect, the present invention further provides an IoT scene perception device based on cloud-edge collaboration, comprising: a data extraction unit, configured to obtain initial scene data to be perceived; wherein the initial scene data includes image data and sensor data;

[0015] An instance perception unit, configured to perform instance perception on the initial scene data based on a dynamic instance perception model, and determine static instances, dynamic instances, and abnormal instances corresponding to the image data; wherein the dynamic instance perception model is a deep neural network model that is pre-trained based on cloud-edge collaboration and deployed in an edge server;

[0016] A scene synthesis unit is used to input the static instance, the dynamic instance, the abnormal instance and the perception results of the sensor data into a local multi-instance scene fusion model for processing to obtain a local scene output by the local multi-instance scene fusion model.

[0017] Furthermore, the scene synthesis unit is specifically used for:

[0018] The static instance, the dynamic instance, the abnormal instance and the perception results of the sensor data are input into a local multi-instance scene fusion model, the static instance, the dynamic instance and the abnormal instance are combined and processed to obtain a corresponding three-dimensional scene, and the sensor data is matched to the three-dimensional scene to obtain the local scene.

[0019] Furthermore, the instance perception unit is specifically used to: input the initial scene data into the dynamic instance perception model to perform instance perception, and obtain static instances, dynamic instances and abnormal instances corresponding to the image data.

[0020] Furthermore, the scene synthesis unit is specifically used to: fuse at least one of the local scenes on the cloud server side to synthesize a global scene based on the local scenes.

[0021] Furthermore, the IoT scene perception device based on cloud-edge collaboration, before acquiring the initial scene data to be perceived, further includes: a model training unit;

[0022] The model training unit is used to pre-train the dynamic instance perception model deployed in the edge server based on the deep neural network model pre-deployed in the cloud server, so as to share some target parameters of each local multi-instance scene fusion model in each edge server during the dynamic instance perception model training process.

[0023] In a third aspect, the present invention also provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the Internet of Things scene perception method based on cloud-edge collaboration as described in any one of the above items are implemented.

[0024] In a fourth aspect, the present invention also provides a processor-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, the steps of the Internet of Things scene perception method based on cloud-edge collaboration as described in any one of the above items are implemented.

[0025] The IoT scene perception method based on cloud-edge collaboration provided by the present invention performs instance perception through initial scene data to determine the static instance, dynamic instance and abnormal instance corresponding to the image data, and inputs the perception results of the static instance, the dynamic instance, the abnormal instance and the sensor data into a local multi-instance scene fusion model deployed in an edge server for processing, thereby obtaining the local scene output by the local multi-instance scene fusion model, which can effectively reduce the IoT scene perception processing delay and improve the adaptability and perception accuracy to high dynamic scenes in the IoT scene. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0027] Figure 1 It is a flow chart of a method for IoT scene perception based on cloud-edge collaboration provided by an embodiment of the present invention;

[0028] Figure 2 It is a schematic diagram of a deep neural network model training process based on cloud-edge collaboration provided by an embodiment of the present invention;

[0029] Figure 3 is a schematic diagram of a training differentiation process of an abnormal instance perception model provided by an embodiment of the present invention;

[0030] Figure 4 is a local scene perception flow chart provided by an embodiment of the present invention;

[0031] Figure 5 It is a schematic diagram of the structure of a deep neural network model in a cloud server and an edge server provided in an embodiment of the present invention;

[0032] Figure 6 It is a schematic diagram comparing the inference delay and transmission delay of two computing solutions provided in an embodiment of the present invention for processing key point perception tasks;

[0033] Figure 7It is a schematic diagram of comparison of loss_keypoint when ShareLayer is set to the first 0, 4, and 8 layers of keypoint_head when the loading batch is 3 provided by an embodiment of the present invention;

[0034] Figure 8 It is a structural diagram of an Internet of Things scene perception device based on cloud-edge collaboration provided by an embodiment of the present invention;

[0035] Fig. 9 It is a schematic diagram of the physical structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0036] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0037] Based on the IoT scene perception method based on cloud-edge collaboration described in the present invention, the following describes in detail an embodiment of the method. Figure 1 As shown, it is a flow chart of the method for IoT scene perception based on cloud-edge collaboration provided by an embodiment of the present invention, and the specific implementation process includes the following steps:

[0038] Step 101: Acquire initial scene data to be sensed; wherein the initial scene data includes image data and sensor data.

[0039] In an embodiment of the present invention, before executing this step, it is necessary to pre-train the dynamic instance perception model deployed in the edge server based on the deep neural network model pre-deployed in the cloud server, and share some target parameters of each dynamic instance perception model in each edge server during the dynamic instance perception model training process to improve the convergence speed of the dynamic instance perception model in high dynamic scenarios.

[0040] like Figure 2 As shown, the present invention provides a neural network model training method based on cloud-edge collaboration for the task of human posture key point recognition in dynamic instance perception, which is used to train a deep neural network model for the Internet of Things scenario on each edge server, namely, a dynamic instance perception model; at the same time, the dynamic instance perception model deployed in the edge server is pre-trained based on the deep neural network model pre-deployed in the cloud server to reduce the load on the edge server.

[0041] In the embodiment of the present invention, the deep neural network model deployed on the cloud server is called CloudNet, and the lightweight deep neural network model deployed on the edge server is called EdgeNet. The present invention trains EdgeNet based on the scene information collected in real time to obtain a dynamic instance perception model for practical application.

[0042] Specifically, the keypoint_head (head network of feature points) of the ROI head (head network of the region of interest) in EdgeNet is divided into two parts: a shared layer ShareLayer and an adaptive layer AdaptiveLayer. Among them, ShareLayer is the high m layer of keypoint_head, which is used to extract the common features of instances in each sub-scene and is trained jointly by each EdgeNet; AdaptiveLayer is the rest of the keypoint_head, which is used to adapt to the different features of instances in each sub-scene and is different in each EdgeNet. Among them, ROI (region of interest) refers to the region of interest in machine vision and image processing. In acquiring the label data required for EdgeNet training, the present invention can automatically generate the label data required for training through CloudNet. This is because CloudNet has a large number of neural network layers and can obtain higher recognition accuracy. Compared with traditional labeling methods, the present invention can greatly reduce the workload and speed up the model training and updating by automatically generating label data with the help of CloudNet.

[0043] In the implementation process of the present invention, the EdgeNet training process includes three parts: EdgeNet initialization process, EdgeNet learning process, and Sharelayer update process.

[0044] Among them, for the EdgeNet initialization process. Let W c is the parameter of CloudNet, W c-b are some target parameters shared with EdgeNet in the CloudNet backbone network, W c-rs is the parameter of ShareLayer in ROI head of CloudNet; W e is the parameter of EdgeNet, W e-b is the parameter of the backbone network in EdgeNet, W e-rs , W e-ra They are the parameters of ShareLayer and AdaptiveLayer in the ROI head of EdgeNet. Given a training set EdgeNet will be trained to optimize the following loss function:

[0045]

[0046] In the initialization phase of EdgeNet, CloudNet is first trained with a large amount of data and sent to the edge server. After receiving CloudNet, the edge server first converts the W c-b , W c-rs As EdgeNet e-b , W e-rs Next, we will e-ra Initialize to a random value and then W e-b , W e-rs , W e-ra Combined into W e , optimize We to complete the training of EdgeNet. In order to save computing resources, the present invention freezes the parameters of the backbone network part of EdgeNet during the training process and only fine-tunes the ROI head part. W e The updated value, the connection ∪ represents the connection between two parameter sets.

[0047] The specific steps are as follows: Input: W c , CloudNet; Output: EdgeNet, Label-builder. Algorithm 1 includes: Step 1: Use the cloud server to send CloudNet (parameter value W c ) is sent to the edge server; Step 2: CloudNet is processed by the edge server to generate W c , W c-b , W c-rs Three copies; Step 3: The edge server connects the backbone network with ShareLayer and AdaptiveLayer in ROI head to build EdgeNet; Step 4: e-ra Initialize to a random value; Step 5: W e =W c-b ∪W e-rs ∪W e-ra ; Step 6: Step 7: Return the value of EdgeNet The value of Label-builder is W c .

[0048] Among them, for the EdgeNet learning process. The present invention uses the scene information collected in real time to further train EdgeNet, and generates the label data required for training based on the Label-builder obtained in the EdgeNet initialization phase. Assume that the label data generated by the Label-builder is accurate. After the edge server receives the scene data, it will use EdgeNet to obtain the perception results and save a copy in the database. When the edge server is idle, it will use the Label-builder to generate the label data of the scene data, and train EdgeNet based on the label data of the scene data. is the parameter of EdgeNet after the initialization phase of EdgeNet. When the number of instances in the newly collected scene accumulates to M, EdgeNet training begins. The newly collected M instances and their labels are recorded as EdgeNet will be trained to optimize the following loss function:

[0049]

[0050] The specific steps are as follows:

[0051] Input: EdgeNet, Target Label-builder; Output: Updated EdgeNet.

[0052] Step 1: The edge server uses Label-builder to obtain the scene instance Tags Step 2: for Updated value; Step 3: Return the value of EdgeNet

[0053] In the present invention, a shared layer ShareLayer (W e-rs ), which is jointly trained by each EdgeNet to further accelerate the convergence speed of EdgeNet training and extract common features of instances in different scenarios. During real-time training, after a certain round of training, each EdgeNet on the edge server will extract the parameters of Sharelayer in its ROI head and upload it to the cloud server. The cloud server uses the preset Fedavg algorithm to summarize the gradients generated by the Sharelayer part of each EdgeNet training to update the Sharelayer. The parameters after Sharelayer update. The detailed process of Sharelayer update is as follows:

[0054] Input: Sharelayer of N EdgeNets, training batch size, n i , CloudNet's Sharelayer, W c-rs ; Output: Updated ShareLayer.

[0055] Step 1: Each edge server separates the parameters of Sharelayer from EdgeNet (i.e. ) and upload to the cloud server; Step 2: Calculate the gradient Step 3: Calculate the weighted average gradient and in Step 4: Step 5 returns the updated ShareLayer as After the Sharelayer update is completed, the cloud server will send the updated ShareLayer parameters to each edge server, and each edge server will call steps 3-7 of the above algorithm 1 to train EdgeNet again, thereby completing an EdgeNet initialization-EdgeNet learning-Sharelayer update cycle, and finally obtain a dynamic instance perception model that meets the application conditions.

[0056] In this step, scene instance extraction can be achieved. Specifically, in the IoT scenario, the initial scene information includes two categories: sensor data (such as voltage, temperature, humidity, etc.) and image data (such as personnel, inspection robots, faulty equipment, etc.), among which image data is divided into three categories: static instances, dynamic instances, and abnormal instances. In the scene instance extraction stage, the initial scene information is first classified to obtain static instances, dynamic instances, abnormal instances, and sensor data.

[0057] The sensor data can be acquired by a preset sensor data acquisition device, and the sensor data can be processed by using a preset N-shotK-way small sample fusion algorithm model. First, K support set classes are constructed, each with N samples (S1, ..., S N ), the purpose of the perception algorithm model is to determine which support set class the collected sensor data should belong to, and the optimization objective is as follows:

[0058]

[0059] Where: The sensor data in the IoT scenario are fused into K categories through the N-shot K-way small sample learning process, which reduces the computational complexity of subsequent analysis in upper-layer applications.

[0060] The image data is obtained by a preset image acquisition device and includes multiple instance types. The present invention divides the image data into three categories: static instances, dynamic instances, and abnormal instances. The present invention uses the Mask-RCNN model to detect and extract three types of target instances from the image data and generate masks so that the subsequent steps can use the corresponding perception algorithm model (i.e., instance perception model) to complete the perception processing.

[0061] Step 102: Perform instance perception on the initial scene data based on a dynamic instance perception model to determine static instances, dynamic instances, and abnormal instances corresponding to the image data; wherein the dynamic instance perception model is a deep neural network model that is pre-trained based on cloud-edge collaborative training and deployed in the edge server.

[0062] In the specific implementation process of this step, the initial scene data can be input into a preset dynamic instance perception model for instance perception, and static instances, dynamic instances and abnormal instances corresponding to the image data can be obtained. The dynamic instance perception model is trained using preset annotated data as training samples.

[0063] It should be noted that the static instance refers to a relatively fixed object in the IoT scene, such as a factory building, a generator set, a pipeline, etc. In an embodiment of the present invention, when implementing the scene instance perception process, a 3D model of a static instance can be preset in the edge server. After the static instance in the scene is identified by Mask-RCNN, the preset 3D model is directly called to participate in the fusion of the local scene. The dynamic instance refers to an object that changes in real time in the IoT scene, such as a worker, etc., which can be synthesized by first detecting key points and then synthesizing a 3D model according to the key point parameters. Taking the generation of a human 3D model as an example, first, the keypoint_head branch of Mask RCNN is used to detect the parameters of the key points of the human body, and then the SMPL (A Skinned Multi-Person Linear Model) parameterized human body model is further used to generate a 3D model of the human body based on the detected parameters, and finally the perception of the human body instance (i.e., dynamic instance) is completed. The abnormal instance is abnormal information in the IoT scene, such as illegal intrusion, faulty equipment, etc., which is characterized by unpredictable features and its 3D model can only be directly perceived by the scene image.

[0064] The present invention perceives abnormal instances through Mesh-RCNN, identifies abnormal instances in the initial image based on Mask-RCNN and segments its sub-images, and then further generates mesh data of the instance based on the sub-images by Mesh-RCNN. The present invention trains Mesh-RCNN models for different types of abnormal instances respectively. When perceiving abnormal instances, the corresponding Mesh-RCNN model is selected from the model library according to the corresponding label data to generate mesh data of the abnormal instance.

[0065] like Figure 3 As shown in the figure, it is the training differentiation process of Mesh-RCNN. First, Mask-RCNN is trained based on the labeled data so that it can detect several specific types of abnormal instances, and they are marked with labels respectively. Under the initial conditions, the Mesh-RCNN model corresponding to each label data is a general model, which is initialized based on the pre-trained model deployed on the cloud server. When a certain number of instances corresponding to a type of label data are collected cumulatively, the mesh annotation of the instance will be obtained through relevant annotation technology, and the Mesh-RCNN model corresponding to the training label will be based on this. When it is necessary to detect new abnormal instances, Mask-RCNN can be trained again based on the labeled data so that it can detect new abnormal instances, and then the above Mesh-RCNN training process is repeated again.

[0066] Step 103: Input the static instance, the dynamic instance, the abnormal instance and the perception results of the sensor data into the local multi-instance scene fusion model for processing to obtain the local scene output by the local multi-instance scene fusion model.

[0067] like Figure 4 As shown, in this step, the static instance, the dynamic instance, the abnormal instance and the perception results of the sensor data are input into the local multi-instance scene fusion model, the static instance, the dynamic instance and the abnormal instance are combined and processed to obtain the corresponding three-dimensional scene, and the sensor data is matched to the three-dimensional scene to obtain the local scene. Furthermore, at least one of the local scenes can be fused on the cloud server side to realize the synthesis of the global scene based on the local scene. Specifically, after the initial scene information is obtained by the acquisition device in step 101, the local scene is obtained by performing instance extraction, instance perception and scene fusion in three stages based on the model pre-deployed in the edge server. After the local scene is obtained based on the perception of each edge server, the local scenes are finally synthesized on the cloud server side to obtain the global scene.

[0068] In the process of scene synthesis, the 3D models of static instances, dynamic instances and abnormal instances will be combined based on the edge server, and the sensor data will be matched to the three-dimensional scene, finally completing the perception of the local scene. It should be noted that in the specific implementation process, the synthesized local scene will retain two copies: one is stored in the edge server to provide real-time services to local users, and the other is uploaded to the cloud server to synthesize the global scene.

[0069] Specifically, the feature extraction process can be denoted as F, the instance perception process can be denoted as P, the instance synthesis process can be denoted as B, the initial data can be denoted as Ori, the sensor data in the IoT scenario can be denoted as Sen, the dynamic instance can be denoted as Dym, the static instance can be denoted as Sta, the abnormal instance can be denoted as Gen, the local scene perception result can be denoted as Loc, and the global scene perception result can be denoted as Gol. From the perspective of data flow, the multi-instance perception fusion process of the local scene can be expressed as follows:

[0070]

[0071] The global scene synthesis process on the cloud server side can be expressed as follows:

[0072]

[0073] The following is an explanation with specific examples:

[0074] This specific process verifies the deep neural network model training method based on cloud-edge collaboration in the task of human posture key point recognition. The cloud-edge collaborative training method of Mask RCNN is constructed based on the Dectectron2 framework available in target detection. CloudNet and EdgeNet are based on the keypoint_rcnn_R_50_FPN_3x model provided by Dectectron2. The detailed network structure is as follows: Figure 5 As shown in the figure. The network structure of CloudNet is the same as the original model. The backbone network of EdgeNet retains the three parts of stem, res2, and res3 in the backbone network of the original model, and the rest are the same as the original model. The ShareLayer for extracting common parameter features is the first 4 layers of FPN, Box Head, Box_predictor, and keypoint_head of Mask RCNN, and the AdaptiveLayer is the last 4 layers of keypoint_head. The preset image data with human posture key point annotations is used as real-time scene data, and is loaded in batches during training to simulate the process of scene data collection in reality.

[0075] The specific process of the present invention mainly compares two aspects: the average delay of scene perception and the training effect of the cloud-edge collaborative training method.

[0076] Figure 6 As shown in the figure, it shows the processing delay and transmission delay of the human posture key point detection task executed on the cloud server and edge server. The inference delay is measured by the simulation system, and the transmission delay is summarized by the online data report. The data in the figure is the transmission delay in the WiFi environment. Compared with the processing method of single cloud computing, sinking the scene instance perception task to the edge server reduces the total processing delay by about 39.8%. Among them, the inference delay is reduced by 38.6% and the transmission delay is reduced by 26.9%.

[0077] In order to verify the effect of the cloud-edge collaborative training method proposed in the present invention on EdgeNet training, the process of collecting scene data and training EdgeNet in practice is simulated. Under initial conditions, the cloud server has deployed a pre-trained CloudNet (deep neural network model). The edge server first initializes EdgeNet (initial dynamic instance perception model) based on 100 preset image data; then, loads images in batches of 100 images, and after each batch of loading is completed, the edge server will retrain EdgeNet; after the edge servers have completed 2 EdgeNet retrainings, they will jointly update ShareLayer with the cloud server and send the update results to the edge servers. After loading a batch of images again, EdgeNet will be trained based on the merged ShareLayer. In this way, after three batches of images are loaded, a round of EdgeNet initialization-EdgeNet learning-Sharelayer update cycle is completed. The present invention simulates a total of 12 loading batches, that is, 4 of the above processing cycles are completed, and finally a dynamic instance perception model is obtained.

[0078] The present invention verifies the change of loss_keypoint with the number of iterations when ShareLayer is the first 0, 4, and 8 layers of keypoint_head respectively. Among them, ShareLayer is the first 4 layers of keypoint_head, which is the method proposed by the present invention. Figure 7 As shown in the figure, it shows the comparison of loss_keypoint in the process of training EdgeNet using the above three methods when loading for the third time. It can be seen that setting ShareLayer_num = 4 obtains the minimum loss_keypoint, and the final converged loss_keypoint is reduced by 7.428% compared with ShareLayer_num = 0 and 4.856% compared with ShareLayer_num = 8.

[0079] In view of the problem that the existing cloud computing model and edge computing model are difficult to adapt to the real-time and accuracy requirements of IoT scene perception, the present invention proposes an IoT scene perception method based on cloud-edge collaboration. First, the characteristics and perception requirements of scene data such as sensor data and image data in the IoT scene are analyzed, and a scene information perception method that distinguishes dynamic instances, static instances and abnormal instances is designed to support local scene information edge perception and global scene cloud synthesis. Secondly, for the high-precision network recognition model of high-frequency changing dynamic instances in the scene (i.e., dynamic instance perception model), a deep neural network model training method based on cloud-edge collaboration is designed, which assists the training of the neural network model on the edge server through the cloud, and uses the cloud to share some parameters of each edge neural network model to improve the model convergence speed, so as to obtain the above-mentioned dynamic instance perception model, effectively reducing the perception processing delay and model training time.

[0080] The IoT scene perception method based on cloud-edge collaboration described in the embodiment of the present invention can effectively reduce the IoT scene perception processing delay and improve the adaptability and perception accuracy to high dynamic scenes in the IoT scene.

[0081] Corresponding to the above-mentioned method for IoT scene perception based on cloud-edge collaboration, the present invention also provides an IoT scene perception device based on cloud-edge collaboration. Since the embodiment of the device is similar to the above-mentioned method embodiment, the description is relatively simple. For relevant details, please refer to the description of the above-mentioned method embodiment. The embodiment of the IoT scene perception device based on cloud-edge collaboration described below is only illustrative. Please refer to Figure 8 As shown, it is a structural diagram of an Internet of Things scene perception device based on cloud-edge collaboration provided by an embodiment of the present invention.

[0082] The IoT scene perception device based on cloud-edge collaboration described in the present invention specifically includes:

[0083] The data extraction unit 801 is used to obtain the initial scene data to be sensed; wherein the initial scene data includes image data and sensor data;

[0084] An instance perception unit 802 is used to perform instance perception on the initial scene data based on a dynamic instance perception model to determine static instances, dynamic instances, and abnormal instances corresponding to the image data; wherein the dynamic instance perception model is a deep neural network model that is pre-trained based on cloud-edge collaboration and deployed in an edge server;

[0085] The scene synthesis unit 803 is used to input the static instance, the dynamic instance, the abnormal instance and the perception result of the sensor data into the local multi-instance scene fusion model for processing to obtain the local scene output by the local multi-instance scene fusion model.

[0086] Furthermore, the scene synthesis unit is specifically used for:

[0087] The static instance, the dynamic instance, the abnormal instance and the perception results of the sensor data are input into a local multi-instance scene fusion model, the static instance, the dynamic instance and the abnormal instance are combined and processed to obtain a corresponding three-dimensional scene, and the sensor data is matched to the three-dimensional scene to obtain the local scene.

[0088] Furthermore, the instance perception unit is specifically used to: input the initial scene data into the dynamic instance perception model to perform instance perception, and obtain static instances, dynamic instances and abnormal instances corresponding to the image data.

[0089] Furthermore, the scene synthesis unit is specifically used to: fuse at least one of the local scenes on the cloud server side to synthesize a global scene based on the local scenes.

[0090] Furthermore, the IoT scene perception device based on cloud-edge collaboration, before acquiring the initial scene data to be perceived, further includes: a model training unit;

[0091] The model training unit is used to pre-train the dynamic instance perception model deployed in the edge server based on the deep neural network model pre-deployed in the cloud server, so as to share some target parameters of each dynamic instance perception model in each edge server during the dynamic instance perception model training process.

[0092] By adopting the Internet of Things scene perception device based on cloud-edge collaboration described in the embodiment of the present invention, instance perception is performed through initial scene data to determine the static instance, dynamic instance and abnormal instance corresponding to the image data, and the perception results of the static instance, the dynamic instance, the abnormal instance and the sensor data are input into the dynamic instance perception model deployed in the edge server for processing, so as to obtain the local scene output by the dynamic instance perception model, which can effectively reduce the Internet of Things scene perception processing delay and improve the adaptability and perception accuracy to high dynamic scenes in the Internet of Things scene.

[0093] Corresponding to the above-mentioned method for IoT scene perception based on cloud-edge collaboration, the present invention also provides an electronic device. Since the embodiment of the electronic device is similar to the above-mentioned method embodiment, the description is relatively simple. For relevant details, please refer to the description of the above-mentioned method embodiment. The electronic device described below is only exemplary. Fig. 9As shown, it is a schematic diagram of the physical structure of an electronic device disclosed in an embodiment of the present invention. The electronic device may include: a processor 901, a memory 902 and a communication bus 903, wherein the processor 901 and the memory 902 complete mutual communication through the communication bus 903, and communicate with the outside through the communication interface 904. The processor 901 can call the logic instructions in the memory 902 to execute the IoT scene perception method based on cloud-edge collaboration, the method comprising: obtaining the initial scene data to be perceived; wherein the initial scene data contains image data and sensor data; based on the dynamic instance perception model, the initial scene data is instance-perceived to determine the static instance, dynamic instance and abnormal instance corresponding to the image data; wherein the dynamic instance perception model is a deep neural network model pre-trained based on cloud-edge collaboration and deployed in the edge server; the perception results of the static instance, the dynamic instance, the abnormal instance and the sensor data are input into the local multi-instance scene fusion model for processing, and the local scene output by the local multi-instance scene fusion model is obtained.

[0094] In addition, the logic instructions in the above-mentioned memory 902 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: a storage chip, a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a disk or an optical disk, and other media that can store program codes.

[0095] On the other hand, an embodiment of the present invention further provides a computer program product, the computer program product comprising a computer program stored on a processor-readable storage medium, the computer program comprising program instructions, and when the program instructions are executed by a computer, the computer can execute the IoT scene perception method based on cloud-edge collaboration provided by the above-mentioned method embodiments. The method comprises: obtaining initial scene data to be perceived; wherein the initial scene data comprises image data and sensor data; performing instance perception on the initial scene data based on a dynamic instance perception model, and determining static instances, dynamic instances, and abnormal instances corresponding to the image data; wherein the dynamic instance perception model is a deep neural network model that is pre-trained based on cloud-edge collaboration and deployed in an edge server; the perception results of the static instance, the dynamic instance, the abnormal instance, and the sensor data are input into a local multi-instance scene fusion model for processing, and a local scene output by the local multi-instance scene fusion model is obtained.

[0096] On the other hand, an embodiment of the present invention further provides a processor-readable storage medium, on which a computer program is stored, and when the computer program is executed by the processor, it is implemented to execute the IoT scene perception method based on cloud-edge collaboration provided by the above embodiments. The method includes: obtaining initial scene data to be perceived; wherein the initial scene data contains image data and sensor data; performing instance perception on the initial scene data based on a dynamic instance perception model, and determining static instances, dynamic instances, and abnormal instances corresponding to the image data; wherein the dynamic instance perception model is a deep neural network model that is pre-trained based on cloud-edge collaboration and deployed in an edge server; the perception results of the static instance, the dynamic instance, the abnormal instance, and the sensor data are input into a local multi-instance scene fusion model for processing, and a local scene output by the local multi-instance scene fusion model is obtained.

[0097] The processor-readable storage medium can be any available medium or data storage device that can be accessed by the processor, including but not limited to magnetic storage (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical storage (such as CD, DVD, BD, HVD, etc.), and semiconductor storage (such as ROM, EPROM, EEPROM, non-volatile memory (NANDFLASH), solid-state drive (SSD)), etc.

[0098] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0099] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for IoT scene perception based on cloud-edge collaboration, characterized in that: include: Acquire initial scene data to be sensed; wherein the initial scene data includes image data and sensor data; Based on the dynamic instance perception model, the initial scene data is instance-aware, and the static instance, dynamic instance and abnormal instance corresponding to the image data are determined; wherein the dynamic instance perception model is a deep neural network model that is pre-trained based on cloud-edge collaboration and deployed in the edge server; The static instance, the dynamic instance, the abnormal instance and the perception results of the sensor data are input into a local multi-instance scene fusion model for processing to obtain a local scene output by the local multi-instance scene fusion model; the static instance, the dynamic instance, the abnormal instance and the perception results of the sensor data are input into a local multi-instance scene fusion model for processing to obtain a local scene output by the local multi-instance scene fusion model, specifically including: combining the static instance, the dynamic instance and the abnormal instance to obtain a corresponding three-dimensional scene, and matching the sensor data to the three-dimensional scene to obtain the local scene.

2. The method for IoT scene perception based on cloud-edge collaboration according to claim 1 is characterized in that: The initial scene data is instance-perceived based on a dynamic instance perception model to determine static instances, dynamic instances, and abnormal instances corresponding to the image data. Specifically, the initial scene data is input into the dynamic instance perception model for instance perception to obtain static instances, dynamic instances, and abnormal instances corresponding to the image data.

3. The method for IoT scene perception based on cloud-edge collaboration according to claim 1 is characterized in that: Also includes: At least one of the local scenes is fused on the cloud server to synthesize a global scene based on the local scenes.

4. The method for IoT scene perception based on cloud-edge collaboration according to claim 1 is characterized in that: Before obtaining the initial scene data to be perceived, it also includes: The dynamic instance perception model deployed in the edge server is pre-trained based on the deep neural network model pre-deployed in the cloud server, so as to share some target parameters of each dynamic instance perception model in each edge server during the training process of the dynamic instance perception model.

5. An IoT scene perception device based on cloud-edge collaboration, characterized in that: include: A data extraction unit, used to obtain initial scene data to be sensed; wherein the initial scene data includes image data and sensor data; An instance perception unit, configured to perform instance perception on the initial scene data based on a dynamic instance perception model, and determine static instances, dynamic instances, and abnormal instances corresponding to the image data; wherein the dynamic instance perception model is a deep neural network model that is pre-trained based on cloud-edge collaboration and deployed in an edge server; A scene synthesis unit is used to input the static instance, the dynamic instance, the abnormal instance and the perception results of the sensor data into a local multi-instance scene fusion model for processing, so as to obtain a local scene output by the local multi-instance scene fusion model; the scene synthesis unit is specifically used to: combine the static instance, the dynamic instance and the abnormal instance to obtain a corresponding three-dimensional scene, and match the sensor data to the three-dimensional scene to obtain the local scene.

6. The IoT scene perception device based on cloud-edge collaboration according to claim 5 is characterized in that: The instance perception unit is specifically used to: input the initial scene data into the dynamic instance perception model to perform instance perception, and obtain static instances, dynamic instances and abnormal instances corresponding to the image data.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the Internet of Things scene perception method based on cloud-edge collaboration as described in any one of claims 1 to 4 are implemented.

8. A processor-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the Internet of Things scene perception method based on cloud-edge collaboration as described in any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Beyond visual range sensing method and system, terminal and storage medium

    CN110210280A

  • Scene interaction method and device and electronic equipment

    CN111274910A