Cloud edge collaboration method based on multi-source heterogeneous information fusion

Through multimodal image processing, Faster R-CNN network and reinforcement learning decision model, the low-conflict and high-reliability problems of cloud-edge collaborative decision-making in multi-source heterogeneous information fusion are solved, and information integration and decision optimization in complex environments are achieved.

CN120766078AActive Publication Date: 2025-10-10EAST CHINA UNIV OF SCI & TECH +1

Patent Information

Application Number
CN202510877955.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-10
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

In the big data environment, how to effectively mine information in the fusion of multi-source heterogeneous information, build an efficient collaborative optimization mechanism, ensure low conflict and high reliability of the system in a complex environment, especially how to identify conflicts between different sides and conduct global information integration in cloud-edge collaborative decision-making.

Method used

A cloud-edge collaborative method is constructed by adopting multimodal image processing and fusion, multi-label training of Faster R-CNN network, evidence distance matrix calculation, cloud-side reinforcement learning decision model and reinforcement learning transfer strategy, and feature extraction and decision optimization through weighted averaging, multi-label cross entropy loss, Dempster rule and deep Q network.

Benefits of technology

It achieves low conflict and high reliability of cloud-edge collaboration in complex environments, and ensures the accuracy and speed of the system under multi-view image input through evidence fusion strategy and reinforcement learning optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766078A_ABST
    Figure CN120766078A_ABST
Patent Text Reader

Abstract

The invention provides a cloud edge collaboration method based on multi-source heterogeneous information fusion, and the method comprises the steps: processing an image, constructing a Faster R-CNN network, outputting basic belief distribution, fusing multi-view basic belief distribution through a supervised method and an unsupervised method, extracting target state features, and building a reinforcement learning model for reasoning and updating. And fusing cloud side classification head migration and a side model, and completing edge reasoning updating under multi-view image input. According to the invention, during fusion, feature extraction and efficient integration are carried out on multi-modal data collected by different sensors, an evidence fusion strategy is introduced, the reliability of evidence is ensured, an evidence correction strategy is introduced for the possible high conflict problem of evidence after fusion of multiple sides, and a conflict identification and dynamic management mechanism is combined, so that the reliability of the evidence is ensured. The generation of an anti-intuition conclusion is effectively avoided, global optimization is performed through reinforcement learning, dynamic interaction and intelligent feedback between cloud edges are realized, and low conflict and high reliability of the system in a complex environment are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of "cloud-edge" collaboration technology, and in particular to a cloud-edge collaboration method based on multi-source heterogeneous information fusion. Background Art

[0002] In the era of big data, there are more and more types of data and the scale of data is increasing. The need to extract information with high utilization value from massive and complex data is becoming more and more urgent. The information extracted by using a single data source and method may have certain deviations. However, using data from multiple sources for information fusion can fully explore the inherent characteristics and laws contained in the information.

[0003] Multi-source heterogeneous information fusion aims to extract implicit and high-value information from many different data sources and various types of data. However, in the big data environment, the data has large structural differences, wide sources, and strong real-time characteristics. How to extract effective information from multi-dimensional and massive big data, establish a reliable conflict management mechanism and collaborative decision-making method, and ensure the accuracy and speed of reasoning has become the key to multi-source heterogeneous information fusion research.

[0004] "Cloud-edge" collaborative decision-making methods that utilize multi-source heterogeneous information fusion involve key technologies such as information fusion at the data level, evidence correction at the feature level, and conflict management at the decision-making level. These methods require mining the deep features of massive signals from diverse sources. However, current research still faces several difficulties and challenges. On the one hand, the system's adaptability to complex environments and the identification of conflicts between different sides require further exploration. On the other hand, cloud-edge collaborative decision-making requires global information integration and decision optimization, and the construction of an efficient collaborative optimization mechanism also presents challenges. Therefore, it is crucial to design a cloud-edge collaborative method based on multi-source heterogeneous information fusion. Summary of the Invention

[0005] In order to overcome the shortcomings of the existing technology, the purpose of the present invention is to provide a cloud-edge collaboration method based on multi-source heterogeneous information fusion.

[0006] To achieve the above object, the present invention provides the following solutions:

[0007] The present invention also provides a cloud-edge collaboration method based on multi-source heterogeneous information fusion, including:

[0008] Step 1: Obtain multimodal images, process and fuse them, decompose the multimodal images into basic parts and detailed content, fuse the basic parts using a weighted average strategy, and extract features of the detailed content using the VGG-19 network. Then, generate the fused detailed content by selecting a strategy, and finally reconstruct the fused image;

[0009] Step 2: Build a Faster R-CNN network and a multi-label training dataset. Replace the single-label classification head of Faster R-CNN with a multi-label Sigmoid layer. Use a multi-label binary cross entropy loss combined with a bounding box regression loss to train the network and output a multi-label membership probability vector for the candidate object region.

[0010] Step 3: The basic belief distribution from multiple perspectives is fused through supervised and unsupervised methods. The supervised weights are based on the historical accuracy of the sensor. The unsupervised weights are constructed by constructing an evidence distance matrix, calculating the distance between each piece of evidence, determining the best BPA and the worst BPA, and using Deng relative entropy to measure the best-other vector and the other-worst vector. The evidence weights are solved through a constrained optimization problem, and the acceptance threshold is determined through sensitivity analysis. The evidence is discounted based on the fusion of the two weights and the final belief distribution is generated through the Dempster rule.

[0011] Step 4: Build a cloud-side reinforcement learning decision model. Based on the prediction results of the edge and cloud models, extract the state features of each target category, construct a state space, and use a deep Q-network to build a reinforcement learning model. Train the model to obtain a migration decision strategy based on performance gain.

[0012] Step 5: Fusion the cloud-side classification head migration and the edge-side model. Based on the migration action output by the reinforcement learning system, fine-tune the parameters of the Faster R-CNN model classification head of the selected category to form a migration patch. The patch is then replaced with the corresponding position of the original model to build a fusion model. Edge inference updates are then completed under multi-view image input.

[0013] Preferably, in step 1, the multimodal image is decomposed into a basic part and detailed content, the basic part is fused using a weighted average strategy, the detailed content is extracted using a VGG-19 network, the fused detailed content is generated by selecting a strategy, and finally a fused image is reconstructed, specifically:

[0014] Acquire multimodal images, including visible light images, infrared images, and SAR images;

[0015] For each input image I K , K is the visible light image, infrared image and SAR image, and the basic part and detailed content are obtained by solving the optimization problem, where the optimization problem is:

[0016]

[0017] Where, As the basic part, g x =[-1,1], g y =[-1,1] T is the gradient operator, λ=5, For detailed content;

[0018] The weighted average strategy is used to integrate the basic part, which is:

[0019]

[0020] Where, (x,y) is the pixel coordinate, They are the basic parts of visible light images, infrared images, and SAR images respectively;

[0021] The pre-trained VGG-19 network is used to extract multi-layer features of detail content, including relu1_1 layer to relu4_1 layer, and the activity level map is calculated by l1-norm. The weight map is generated by Softmax, and the weighted fusion of detail content is performed after upsampling. Finally, the fusion detail F is generated by the maximum selection strategy. d ;

[0022] The final fused image is obtained by directly adding the obtained fused basic part and fused detail part:

[0023] F=F b +F d .

[0024] Preferably, in step 2, a multi-label training dataset is constructed, specifically:

[0025] Get the label vector Y = [y1,y2,...,y L ], where y l ∈{0,1} indicates whether it belongs to the lth category label, and a multi-label training dataset is constructed.

[0026] Preferably, in step 2, the loss function of the Faster R-CNN network is:

[0027]

[0028] Where, L BCE is the predicted label probability vector, Y is the true label vector, and B are the predicted and true bounding box coordinates respectively;

[0029] The Faster R-CNN network outputs instance number, predicted category, bounding box coordinates, and confidence levels for all classification labels. The detection results of each sensor are converted into a unified format: [instance index, category, bounding box, probability array].

[0030] Preferably, in step 3, the basic belief distribution of multiple perspectives is fused through supervised and unsupervised methods. The supervised method is based on the historical accuracy of the sensor; the unsupervised method constructs an evidence distance matrix, calculates the distance between each piece of evidence, determines the best BPA and the worst BPA, uses Deng relative entropy to measure the best-other vector and the other-worst vector, solves the evidence weight through constrained optimization problem, determines the acceptance threshold through sensitivity analysis, discounts the evidence based on the fusion of the two weights, and generates the final belief distribution through Dempster rule, specifically:

[0031] Construct the Jousselme evidence distance matrix based on the multi-label membership probability vector;

[0032] Calculate the distance CD (m) between each piece of evidence j ), select the one with the largest CD as the worst belief allocation m W , and m W The one with the farthest distance is the best belief allocation m B ;

[0033] Based on Deng relative entropy, the best-other vector and the other-worst vector are calculated, and the evidence weight ω is solved through constrained optimization problem. j , determine the acceptance threshold ξ through sensitivity analysis max , if ξ≤ξ max Then accept the weights, otherwise re-adjust the comparison matrix;

[0034] Determine the supervised weight based on the historical accuracy of the sensor, and perform weighted fusion of the two weights;

[0035] According to the fused weighted discounted belief distribution, the final belief distribution is obtained through Dempster rule fusion.

[0036] Preferably, in step 4, a cloud-side reinforcement learning decision model is constructed. Based on the prediction results of the edge model and the cloud model, the state features of each target category are extracted, the state space is constructed, and a deep Q network is used to establish a reinforcement learning model. The training obtains a migration decision strategy based on performance gain, specifically:

[0037] Based on the prediction results of the edge model and the cloud model, the state feature vector of each target category is extracted, and the state feature vector includes the accuracy Acc of the edge category. edge , cloud category accuracy ACC cloud And the difference between the two Δ=A cloud -A edge , forming the state vector S = [A edge ,A cloud ,Δ];

[0038] Build a reinforcement learning agent with a deep Q network as the core, and define the action set A = {0, 1}, where 0 means skipping the migration and 1 means triggering the current category migration operation;

[0039] Set the reward mechanism, after the migration is executed, the accuracy change of this category based on the edge model δ = A cloud -A edge Calculate the reward R, that is:

[0040]

[0041] Where θ is the minimum precision gain threshold, δ' is the accuracy of the last effective migration, and A cloud is the cloud-side accuracy of the current model, A edge is the side accuracy of the current model;

[0042] Generate a strategy, update and iterate through the Q function, and output the migration decision strategy π(s)=argmax based on the state S a Q(s,a), which indicates whether to perform transfer on the target category.

[0043] Preferably, in step 5, the cloud-side classification head migration and the edge-side model are integrated, and the parameters of the FasterR-CNN model classification head of the selected category are fine-tuned according to the migration action output by the reinforcement learning system to form a migration patch. The patch is then replaced to the corresponding position of the original model to build a fusion model, and the edge inference update is completed under multi-view image input, specifically as follows:

[0044] Based on the generated migration strategy output, select the category with action 1 and trigger the migration training of the corresponding classification head;

[0045] After each migration training is completed, the accuracy of the category on the cloud model and the edge model is calculated respectively, and recorded as A cloud ,A edge , and calculate the accuracy difference δ = A cloud -A edge If δ≤θ, the migration is considered to be fully completed, the migration training patch of this category is saved, and the migration operation of this category will not be triggered in subsequent training cycles;

[0046] During the migration training process, all parameters in the Faster R-CNN network except the classification head corresponding to the current target category are frozen. Only the weight vector head.score.weight and bias item head.score.bias corresponding to the category are back-propagated and optimized. After the migration is completed, the trained patch parameters are stored as a lightweight update module dedicated to the category and replaced with the corresponding position of the edge model to complete the local update and model fusion deployment.

[0047] When all target categories have completed the migration and their cloud-edge accuracy difference is δ≤θ, the entire training process is automatically terminated.

[0048] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0049] The present invention provides a cloud-edge collaboration method based on multi-source heterogeneous information fusion, which includes acquiring multimodal images, processing and fusing them, decomposing the multimodal images into basic parts and detailed contents, fusing the basic parts with a weighted average strategy, extracting features of the detailed contents using a VGG-19 network, generating fused detailed contents by selecting a strategy, and finally reconstructing a fused image, constructing a FasterR-CNN network and a multi-label training dataset, replacing the single-label classification head of FasterR-CNN with a multi-label Sigmoid layer, training the network with a multi-label binary cross entropy loss combined with a bounding box regression loss, outputting a multi-label membership probability vector of the target candidate area, and fusing the basic belief distribution of multiple perspectives through supervised and unsupervised methods. The supervised weights are based on the sensor history. Historical accuracy; unsupervised weights construct an evidence distance matrix, calculate the distance between each piece of evidence, determine the best BPA and the worst BPA, use Deng relative entropy to measure the best-other vector and other-worst vector, solve the evidence weights through constrained optimization problems, determine the acceptance threshold through sensitivity analysis, fuse the two weights, discount the evidence according to the fused weights, and generate the final belief distribution through the Dempster rule. A cloud-side reinforcement learning decision model is constructed. Based on the prediction results of the edge model and the cloud model, the state features of each target category are extracted, the state space is constructed, and a deep Q network is used to establish a reinforcement learning model. The training obtains a migration decision strategy based on performance gain, integrates the cloud-side classification head migration and the edge-side model, and fine-tunes the classification head parameters of the selected category of the Faster R-CNN model according to the migration action output by the reinforcement learning system to form a migration patch. The patch is then replaced to the corresponding position of the original model to build a fusion model, and the edge reasoning update is completed under multi-view image input. In the information fusion stage, the present invention extracts features and efficiently integrates multimodal data collected by different sensors, introduces an evidence fusion strategy to ensure the reliability of evidence, and introduces an evidence correction strategy to address the high conflict problem that may exist in evidence after multiple side fusions. Combined with conflict identification and dynamic management mechanisms, it effectively avoids the generation of counter-intuitive conclusions. Finally, global optimization is performed through reinforcement learning to achieve dynamic interaction and intelligent feedback between the "cloud and edge", ensuring low conflict and high reliability of the system in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0051] Figure 1 A cloud-edge collaboration method based on multi-source heterogeneous information fusion provided by the embodiment of the present application is shown in the flowchart of the multi-source heterogeneous information fusion method.

[0052] Figure 2 A cloud-edge collaboration method based on multi-source heterogeneous information fusion provided by the embodiment of the present application is shown in the flowchart of the multi-source heterogeneous information fusion method.

[0053] Figure 3 A cloud-edge collaboration method based on multi-source heterogeneous information fusion provided by the embodiment of the present application is shown in the flowchart of the multi-source heterogeneous information fusion method.

[0054] Figure 4 A cloud-edge collaboration method based on multi-source heterogeneous information fusion provided by the embodiment of the present application is shown in the flowchart of the multi-source heterogeneous information fusion method. DETAILED DESCRIPTION

[0055] The technical solutions in the embodiments of the present application will be described clearly and completely with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments only constitute some embodiments of the present application, and all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0056] The purpose of the present application is to provide a cloud-edge collaboration method based on multi-source heterogeneous information fusion. In the information fusion stage, the multi-modal data collected by different sensors are subjected to feature extraction and efficient integration. The evidence fusion strategy is introduced to ensure the reliability of the evidence. In view of the high conflict problem that may exist in the evidence after the multi-edge fusion, the evidence correction strategy is introduced. Combined with the conflict identification and dynamic management mechanism, the generation of counterintuitive conclusions is effectively avoided. Finally, the global optimization is carried out through reinforcement learning to realize the dynamic interaction and intelligent feedback between the "cloud" and the "edge", and to ensure the low conflict and high reliability of the system in complex environment.

[0057] In order to make the above-mentioned purposes, features and advantages of the present application more apparent and easy to understand, the present application will be further described in detail with reference to the drawings and specific embodiments.

[0058] Figure 1 A cloud-edge collaboration method based on multi-source heterogeneous information fusion provided by the embodiment of the present application is shown in the flowchart of the multi-source heterogeneous information fusion method, as shown in Figure 1 The present application provides a cloud-edge collaboration method based on multi-source heterogeneous information fusion, which comprises:

[0059] Step 1: Acquire a multimodal image, deploy feature extraction and image recognition models on the side, process and fuse them, decompose the multimodal image into a base part and detailed content, fuse the base part using a weighted average strategy, and extract features from the detailed content using the VGG-19 network. A strategy is selected to generate the fused detailed content, and finally reconstruct the fused image;

[0060] Step 2: Build a Faster R-CNN network and a multi-label training dataset. Replace the single-label classification head of Faster R-CNN with a multi-label Sigmoid layer. Use a multi-label binary cross entropy loss combined with a bounding box regression loss to train the network and output a multi-label membership probability vector for the candidate object region.

[0061] Step 3: Deploy supervised and unsupervised methods on the cloud side to fuse the basic belief allocation from multiple perspectives. The supervised method is based on the historical accuracy of the sensor. The unsupervised method constructs an evidence distance matrix, calculates the distance between each piece of evidence, determines the best and worst BPA, uses Deng relative entropy to measure the best-other vector and the other-worst vector, solves the evidence weights through a constrained optimization problem, determines the acceptance threshold through sensitivity analysis, discounts the evidence based on the fusion of the two weights, and generates the final belief allocation through the Dempster rule.

[0062] Step 4: Build a cloud-side reinforcement learning decision model. Based on the prediction results of the edge and cloud models, extract the state features of each target category, construct a state space, and use a deep Q-network to build a reinforcement learning model. Train the model to obtain a migration decision strategy based on performance gain.

[0063] Step 5: Fusion the cloud-side classification head migration and the edge-side model. Based on the migration action output by the reinforcement learning system, fine-tune the parameters of the Faster R-CNN model classification head of the selected category to form a migration patch. The patch is then replaced with the corresponding position of the original model to build a fusion model. Edge inference updates are then completed under multi-view image input.

[0064] like Figure 2 As shown in Figure 1, in step 1, the multimodal image is decomposed into basic parts and detailed contents. The basic parts are fused using a weighted average strategy, and the detailed contents are extracted using the VGG-19 network. The fused detailed contents are generated by selecting a strategy, and finally the fused image is reconstructed. Specifically:

[0065] Acquire multimodal images, including visible light images, infrared images, and SAR images;

[0066] For each input image I K, K is the visible light image, infrared image and SAR image, and the basic part and detailed content are obtained by solving the optimization problem, where the optimization problem is:

[0067]

[0068] Where, As the basic part, g x =[-1,1], g y =[-1,1] T is the gradient operator, λ=5, For detailed content;

[0069] The weighted average strategy is used to integrate the basic part, which is:

[0070]

[0071] Where, (x,y) is the pixel coordinate, They are the basic parts of visible light images, infrared images, and SAR images respectively;

[0072] The pre-trained VGG-19 network is used to extract multi-layer features of detail content, including relu1_1 layer to relu4_1 layer, and the activity level map is calculated by l1-norm. The weight map is generated by Softmax, and the weighted fusion of detail content is performed after upsampling. Finally, the fusion detail F is generated by the maximum selection strategy. d ;

[0073] The final fused image is obtained by directly adding the obtained fused basic part and fused detail part:

[0074] F=F b +F d .

[0075] In step 2, a multi-label training dataset is constructed, specifically:

[0076] Get the label vector Y = [y1,y2,...,y L ], where y l ∈{0,1} indicates whether it belongs to the lth category label, and a multi-label training dataset is constructed.

[0077] In step 2, the loss function of the Faster R-CNN network is:

[0078]

[0079] Where, L BCE is the predicted label probability vector, Y is the true label vector, and B are the predicted and true bounding box coordinates respectively;

[0080] The Faster R-CNN network outputs the instance number, predicted category, bounding box coordinates, and confidence levels for all classification labels. The detection results of each sensor are converted into a unified format: [instance index, category, bounding box, probability array].

[0081] For different sensor postures and heights, the sensor with the largest number of detected instances is selected as the reference perspective, and the detection bounding box is rotated and scaled to map it to the reference coordinate system;

[0082] The target pool is initialized with the instances detected by the benchmark sensor. The IoU matrix between the detection results of the benchmark sensor and the current sensor is calculated as the matching cost. The optimal matching pair is solved by linear allocation using the Hungarian algorithm. The IoU threshold is set to be greater than 0.5 to filter out valid matches. The successfully matched targets are fused, and new targets that are not matched are added to the target pool as new instances.

[0083] like Figure 3 As shown in Figure 3, in step 3, to address the high conflict problem that may exist between the evidence of multiple side sensors, an evidence correction strategy is introduced. Combining conflict identification with a dynamic management mechanism, it effectively avoids the generation of counterintuitive conclusions. The basic belief allocation (BPA) of multiple perspectives is integrated through supervised and unsupervised methods. The supervised method is based on the historical accuracy of the sensor; the unsupervised method constructs an evidence distance matrix, calculates the distance between each piece of evidence, determines the best BPA and the worst BPA, and uses Deng relative entropy to measure the best-other vector and the other-worst vector. The evidence weight is solved through a constrained optimization problem, the consistency ratio is calculated to verify the reliability of the weight, the evidence is discounted according to the weight, and the final belief allocation is generated through the Dempster rule. Specifically,

[0084] Construct a single element focal element m based on the probability array of the same instance detected by multiple sides of the multi-label membership probability vector i and m j The Jousellme distance matrix between:

[0085]

[0086] in, D is the Jaccard matrix:

[0087] Calculate the distance CD (m) between each piece of evidence j ),for:

[0088]

[0089] Choose the one with the largest CD as the worst belief allocation m W , which means that this BPA contributes the most to the disorder of the system, and is the same as m W The one with the farthest distance is the best belief allocation m B ,

[0090] Using Deng relative entropy to measure the distribution of two basic beliefs m i and m j The information difference between:

[0091]

[0092] Satisfies non-negativity (σ≥0) and asymmetry (σ(m i ||m j )≠σ(m j ||m i ));

[0093] Construct a relative entropy comparison matrix, including m B to other BPA vectors and other BPA to m W The vector is:

[0094] M B =(σ(m B ||m1),σ(m B ||m2),...,σ(m B ||m n ))

[0095] M W =(σ(m1||m W ),σ(m2||m w ),...,σ(m n ||m w )) T

[0096] Solving the weight of evidence ω through constrained optimization problem j ,for:

[0097] minξ

[0098]

[0099] ∑ω j =1,ω j ≥0

[0100] Determine the acceptance threshold ξ through sensitivity analysis max , if ξ≤ξ max Then accept the weights, otherwise re-adjust the comparison matrix;

[0101] The accuracy of the sensor under different viewing angles is calculated respectively, the supervised weight is determined according to the accuracy, and the supervised weight and the unsupervised weight are fused, that is:

[0102] Fusion weight = alpha * supervised weight + (1-alpha) * omega j

[0103] Take alpha = 0.5 temporarily, discount the belief distribution according to the fusion weight, and obtain the final belief distribution by fusing through the Dempster rule.

[0104] As Figure 4 shown, in step 4, a cloud-side reinforcement learning decision model is constructed, based on the prediction results of the edge model and the cloud model, the state features of each target category are extracted, the state space is constructed, and a deep Q network is used to establish a reinforcement learning model, and a migration decision strategy based on performance gain is obtained by training, specifically:

[0105] Based on the prediction results of the edge model and the cloud model result.xlsx, (excel fields include aircraft: detected aircraft, satellite_id: detected aircraft corresponding satellite, soft_label_vector_edge: edge category soft label, confidence_edge: edge category confidence, type_prediction_edge: edge prediction aircraft category, confidence_cloud: cloud side confidence, type_prediction_cloud: cloud side prediction category), the state feature vector of each target category is extracted, the state feature vector includes the accuracy of the edge category Acc edge , the accuracy of the cloud category ACC cloud and the difference Δ = A cloud -A edge , the state vector S = [A edge , A cloud , Δ] is constructed.

[0106] A reinforcement learning agent with deep Q network as the core is constructed, the action {0, 1} is selected, a random number <epsilon: random action, otherwise: select action argmax(Q(s, a)) according to current Q network, execute migration and evaluate new accuracy.

[0107] Set the reward mechanism, after migration execution, based on the accuracy change of the edge model of this category δ = A cloud -A edge Calculate the reward R, that is:

[0108]

[0109] Where θ is the minimum precision gain threshold, δ' is the accuracy of the last effective migration, and A cloud is the cloud-side accuracy of the current model, A edge is the side accuracy of the current model;

[0110] Generate a strategy, update and iterate through the Q function, add (state, action, reward, next_state) to the experience pool, update the DQN strategy, and output the migration decision strategy π(s) = argmax based on the state S a Q(s,a), which indicates whether to perform transfer on the target category.

[0111] In step 5, the cloud-side classification head migration and the edge-side model are integrated. Based on the migration action output by the reinforcement learning system, the parameters of the Faster R-CNN model classification head of the selected category are fine-tuned to form a migration patch. The patch is then replaced with the corresponding position of the original model to build a fusion model. The edge inference update is completed under multi-view image input. Specifically:

[0112] Based on the generated migration strategy output, select the category with action 1 and trigger the migration training of the corresponding classification head;

[0113] After each migration training is completed, the accuracy of the category on the cloud model and the edge model is calculated respectively, and recorded as A cloud ,A edge , and calculate the accuracy difference δ = A cloud -A edge If δ≤θ, the migration is considered to be fully completed, the migration training patch of this category is saved, and the migration operation of this category will not be triggered in subsequent training cycles;

[0114] Construct a dataset and load the FasterR-CNN pre-trained model, freeze all parameters except for this class, and only train the classification head head.score.weight and head.score.bias of this class;

[0115] During the transfer training process, all network parameters in the Faster R-CNN network except the classification head (classifier) ​​parameters corresponding to the current target category T are frozen. Only the classification weight vector and bias term of this category are fine-tuned to improve the detection performance of this category in the edge model. The fine-tuning training only uses the predicted candidate regions (ROIs) with a high overlap (IoU>0.5) with the annotation box as positive samples. The optimization objective function consists of two parts:

[0116] L total =L cls +λL reg

[0117] Where, L cls : Classification loss, using the cross entropy function (CrossEntropyLoss) to calculate the category discrimination error of the positive sample ROI area, L leg : Regression loss, using the smooth L1 loss function (SmoothL1 Loss), is used for the bounding box coordinate regression of the positive sample ROI; the regression loss weight coefficient λ = 0.1 is used to balance the classification and position regression goals; the screening condition for the positive sample ROI area is that the overlap between the IoU and the ground truth box is > 0.5; fine-tuning only acts on the classification head parameters corresponding to the current category, that is:

[0118] W T =head.score.weight[T]

[0119] After the migration is complete, the trained patch parameters are stored as a lightweight update module dedicated to this category and written to the corresponding location of the edge model to complete the local update and model fusion deployment.

[0120] During the training process, the set of all target categories is: C = {T1, T2, .... T N For each category T i ∈C, calculate the accuracy difference between its cloud model and edge model after each round of training:

[0121]

[0122] in, Category T i Prediction accuracy of cloud-based models; Category T i Prediction accuracy on marginal models; δ Ti Indicates the performance difference between the edge model and the cloud model. Define the training completion status indicator function:

[0123]

[0124] Among them, θ is an adjustable accuracy convergence threshold (0.02). When all target categories meet the accuracy difference that does not exceed the threshold, that is: It is determined that all target categories have been transferred and the training process is automatically terminated.

[0125] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0126] The present invention uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.

Claims

1. A cloud-edge collaboration method based on multi-source heterogeneous information fusion, characterized by: include: Step 1: Obtain multimodal images, process and fuse them, decompose the multimodal images into basic parts and detailed content, fuse the basic parts using a weighted average strategy, and extract features of the detailed content using the VGG-19 network. Then, generate the fused detailed content by selecting a strategy, and finally reconstruct the fused image; Step 2: Build a Faster R-CNN network and a multi-label training dataset. Replace the single-label classification head of Faster R-CNN with a multi-label Sigmoid layer. Use a multi-label binary cross entropy loss combined with a bounding box regression loss to train the network and output a multi-label membership probability vector for the candidate object region. Step 3: The basic belief distribution from multiple perspectives is fused through supervised and unsupervised methods. The supervised weights are based on the historical accuracy of the sensor. The unsupervised weights are constructed by constructing an evidence distance matrix, calculating the distance between each piece of evidence, determining the best BPA and the worst BPA, and using Deng relative entropy to measure the best-other vector and the other-worst vector. The evidence weights are solved through a constrained optimization problem, and the acceptance threshold is determined through sensitivity analysis. The evidence is discounted based on the fusion of the two weights and the final belief distribution is generated through the Dempster rule. Step 4: Build a cloud-side reinforcement learning decision model. Based on the prediction results of the edge and cloud models, extract the state features of each target category, construct a state space, and use a deep Q-network to build a reinforcement learning model. Train the model to obtain a migration decision strategy based on performance gain. Step 5: Fusion the cloud-side classification head migration and the edge-side model. Based on the migration action output by the reinforcement learning system, fine-tune the parameters of the Faster R-CNN model classification head of the selected category to form a migration patch. The patch is then replaced with the corresponding position of the original model to build a fusion model. Edge inference updates are then completed under multi-view image input.

2. The method according to claim 1, characterized in that In step 1, the multimodal image is decomposed into a basic part and detailed content. The basic part is fused using a weighted average strategy, and the detailed content is extracted using the VGG-19 network. The fused detailed content is generated by selecting a strategy, and finally the fused image is reconstructed. Specifically: Acquire multimodal images, including visible light images, infrared images, and SAR images; For each input image I K , K is the visible light image, infrared image and SAR image, and the basic part and detailed content are obtained by solving the optimization problem, where the optimization problem is: Where, As the basic part, g x =[-1,1], g y =[-1,1] T is the gradient operator, λ=5, For detailed content; The weighted average strategy is used to integrate the basic part, which is: Where, (x,y) is the pixel coordinate, They are the basic parts of visible light images, infrared images, and SAR images respectively; The pre-trained VGG-19 network is used to extract multi-layer features of detail content, including relu1_1 layer to relu4_1 layer, and the activity level map is calculated by l1-norm. The weight map is generated by Softmax, and the weighted fusion of detail content is performed after upsampling. Finally, the fusion detail F is generated by the maximum selection strategy. d ; The final fused image is obtained by directly adding the obtained fused basic part and fused detail part: F=F b +F d 。 3. The method according to claim 2, characterized in that In step 2, a multi-label training dataset is constructed, specifically: Get the label vector Y = [y1,y2,...,y L ], where y l ∈{0,1} indicates whether it belongs to the lth category label, and a multi-label training dataset is constructed.

4. The method according to claim 3, characterized in that In step 2, the loss function of the Faster R-CNN network is: Where, L BCE is the predicted label probability vector, Y is the true label vector, and B are the predicted and true bounding box coordinates respectively; The Faster R-CNN network outputs instance number, predicted category, bounding box coordinates, and confidence levels for all classification labels. The detection results of each sensor are converted into a unified format: [instance index, category, bounding box, probability array].

5. The method according to claim 4, characterized in that In step 3, the basic belief distribution of multiple perspectives is fused through supervised and unsupervised methods. The supervised weight is based on the historical accuracy of the sensor; the unsupervised weight is determined by constructing an evidence distance matrix, calculating the distance between each piece of evidence, determining the best BPA and the worst BPA, and using Deng relative entropy to measure the best-other vector and the other-worst vector. The evidence weight is solved through a constrained optimization problem, and the acceptance threshold is determined through sensitivity analysis. The evidence is discounted based on the fusion of the two weights and the final belief distribution is generated through the Dempster rule. Specifically, Construct the Jousselme evidence distance matrix based on the multi-label membership probability vector; Calculate the distance CD (m) between each piece of evidence j ), select the one with the largest CD as the worst belief allocation m W , and m W The one with the farthest distance is the best belief allocation m B ; Based on Deng relative entropy, the best-other vector and the other-worst vector are calculated, and the evidence weight ω is solved through constrained optimization problem. j , determine the acceptance threshold ξ through sensitivity analysis max , if ξ≤ξ max Then accept the weights, otherwise re-adjust the comparison matrix; Determine the supervised weight based on the historical accuracy of the sensor, and perform weighted fusion of the two weights to obtain a new weight; The belief distribution is discounted according to the new weights and fused through the Dempster rule to obtain the final belief distribution.

6. The method according to claim 5, characterized in that In step 4, a cloud-side reinforcement learning decision model is constructed. Based on the prediction results of the edge model and the cloud model, the state features of each target category are extracted, the state space is constructed, and a deep Q-network is used to establish a reinforcement learning model. The training obtains a migration decision strategy based on performance gain. Specifically: Based on the prediction results of the edge model and the cloud model, the state feature vector of each target category is extracted, and the state feature vector includes the accuracy Acc of the edge category. edge , cloud category accuracy ACC cloud And the difference between the two Δ=A cloud -A edge , forming the state vector S = [A edge ,A cloud ,Δ]; Build a reinforcement learning agent with a deep Q network as the core, and define the action set A = {0, 1}, where 0 means skipping the migration and 1 means triggering the current category migration operation; Set the reward mechanism, after the migration is executed, the accuracy change of this category based on the edge model δ = A cloud -A edge Calculate the reward R, that is: Where θ is the minimum precision gain threshold, δ' is the accuracy of the last effective migration, and A cloud is the cloud-side accuracy of the current model, A edge is the side accuracy of the current model; Generate a strategy, update and iterate through the Q function, and output the migration decision strategy π(s)=argmax based on the state S a Q(s,a), which indicates whether to perform transfer on the target category.

7. The method according to claim 6, characterized in that In step 5, the cloud-side classification head migration and the edge-side model are integrated. Based on the migration action output by the reinforcement learning system, the parameters of the Faster R-CNN model classification head of the selected category are fine-tuned to form a migration patch. The patch is then replaced with the corresponding position of the original model to build a fusion model. The edge inference update is completed under multi-view image input. Specifically: Based on the generated migration strategy output, select the category with action 1 and trigger the migration training of the corresponding classification head; After each migration training is completed, the accuracy of the category on the cloud model and the edge model is calculated respectively, and recorded as A cloud ,A edge , and calculate the accuracy difference δ = A cloud -A edge If δ≤θ, the migration is considered to be fully completed, the migration training patch of this category is saved, and the migration operation of this category will not be triggered in subsequent training cycles; During the migration training process, all parameters in the Faster R-CNN network except the classification head corresponding to the current target category are frozen. Only the weight vector head.score.weight and bias term head.score.bias corresponding to the category are back-propagated and optimized. After the migration is completed, the trained patch parameters are stored as a lightweight update module dedicated to the category and replaced with the corresponding location in the edge model to complete the local update and model fusion deployment. When all target categories have completed the migration and their cloud-edge accuracy difference is δ≤θ, the entire training process is automatically terminated.

Citation Information

Patent Citations

  • Self-service terminal login state monitoring and active safety protection method

    CN117992941A

  • Visible light, SAR and electronic reconnaissance fusion remote sensing target detection method and device

    CN118887559A

  • Thermal power plant fan equipment fault diagnosis method based on fusion fault diagnosis model

    CN119066614A

  • Model update system, model update method, and related device

    US20220284352A1

Cited By

  • Park governance-oriented cloud-edge collaborative multi-modal reasoning method and system

    CN121146100A

  • Building material quality intelligent judgment device and method fused with edge calculation

    CN121563326A