A Smart City Decision-making Method and System Based on Edge Computing
By training the image encoder and classifier in the edge computing unit of the smart city monitoring device, extracting and fusing road image features, the problem of inaccurate identification of road occupation operations in the prior art is solved, and higher recognition accuracy and decision-making support capabilities are achieved.
Patent Information
- Application Number
- CN202510429456.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-04-08
AI Technical Summary
When identifying road-occupying operations, the prior art fails to make full use of the information in the initial image, resulting in inaccurate identification results and inability to provide accurate information for smart city decision-making.
Using an edge computing method, the accuracy of road occupation business identification is improved by training the image encoder and classifier in the edge computing unit of the monitoring device.
By integrating real-time feature maps of multiple monitoring devices, more accurate road-occupying business identification results are generated, and the information accuracy of smart city management decisions is improved.
Smart Images

Figure CN119942466B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of smart cities, and in particular, to a smart city decision-making method and system based on edge computing. Background Art
[0002] With the continuous advancement of urban intelligence, artificial intelligence technology is used to empower the decision-making of smart cities, automatically identify illegal occupation of roads, and improve the efficiency of urban appearance environment management.
[0003] Currently, the patent application document with the publication number CN114155217A discloses a vision-based illegal operation detection method. It performs object detection on the initial image to extract all objects in the initial image; classifies the extracted objects; filters out interference objects according to the classification and outputs the objects of illegal operations; performs road segmentation on the initial image to divide the in-store operation area and the area of illegal occupation of roads; combines the area where the objects of illegal operations are located with the result mask of road segmentation to output the objects of in-store operations and the objects of illegal occupation of roads.
[0004] Although the above method can eliminate the influence of interference objects on the identification of illegal occupation of roads, it ignores other effective information related to the identification of illegal occupation of roads in the initial image, does not fully utilize the information in the initial image, resulting in inaccurate identification results of illegal occupation of roads and being unable to provide accurate information for the decision-making of smart cities. Summary of the Invention
[0005] In order to solve the above technical problems in the prior art, this application provides a smart city decision-making method based on edge computing to improve the accuracy of the identification results of illegal occupation of roads and provide accurate information for the management decision-making of smart cities.
[0006] The present invention provides a smart city decision-making method based on edge computing for identifying illegal occupation of roads on a target road. A plurality of monitoring devices including edge computing units are deployed on the target road. The method includes: in an edge computing unit, a trained image encoder extracts features from the real-time road images collected by the monitoring device to obtain the real-time feature map of the monitoring device, where the parameters of the image encoders in different edge computing units are shared; in response to a target monitoring device receiving the real-time feature maps of other monitoring devices, based on the set association vector of the target monitoring device, weighted summation is performed on each real-time feature map to obtain a real-time fusion feature. The target monitoring device is any one of the monitoring devices, and the set association vector includes the degree of association between the illegal occupation of road identification result of the target monitoring device and each real-time feature map; inputting the real-time fusion feature into the trained classifier in the target edge computing unit to output the illegal occupation of road identification result of the target monitoring device, where the target edge computing unit is the edge computing unit in the target monitoring device, and the parameters of the classifiers in different edge computing units are not shared; formulating management measures based on the illegal occupation of road identification result.
[0007] Deploy a plurality of monitoring devices on the target road to collect road images of different areas on the target road. The edge computing units of the monitoring devices store a trained image encoder and a classifier. The parameters of the image encoders in different edge computing units are shared to ensure the same feature extraction process for road images in different areas, and obtain the feature maps corresponding to the road images in each area on the target road; further, for a target monitoring device, fuse the feature maps according to the degree of association between the illegal occupation of road identification result of the target monitoring device and all the feature maps, and input the fused features into the classifier stored in the edge computing unit of the target monitoring device to obtain the illegal occupation of road identification result of the corresponding area of the target monitoring device; the fused features include all environmental features related to the illegal occupation of road identification result of the target monitoring device on the target road, improving the accuracy of the illegal occupation of road identification result of the target monitoring device.
[0008] In some embodiments, the image encoder is a convolutional neural network, and the classifier includes a fully connected neural network and a classification function, and the classification function is a softmax function.
[0009] In some embodiments, the real-time fusion feature satisfies the relationship:
[0010] ;
[0011] where, is the th real-time feature map, is the number of real-time road images, is the number of real-time feature maps corresponding to a real-time road image, is the degree of association between the identification result of illegal occupation of roads by the target monitoring device and the th real-time feature map, is the real-time fusion feature.
[0012] All real-time feature maps are fused according to the degree of association between the identification result of illegal occupation of roads by the target monitoring device and each real-time feature map to obtain a real-time fusion feature; the real-time fusion feature includes all environmental features related to the identification result of illegal occupation of roads by the target monitoring device on the target road. Classification based on the real-time fusion feature can improve the accuracy of the identification result of illegal occupation of roads by the target monitoring device.
[0013] In some embodiments, the method for obtaining the set association vector of the target monitoring device includes: for a set of training data, the feature maps of all road images are obtained according to the trained image encoder, the mean value of all feature maps is input into the trained classifier in the target edge computing unit to obtain the identification result of the target road image; calculate the cross-entropy loss between the identification result of the target road image and the label, and calculate the gradient magnitude of each feature map based on the cross-entropy loss; normalize the gradient magnitudes of all feature maps to obtain the normalized vector corresponding to the training data, and calculate the average value of the normalized vectors corresponding to all training data to obtain the set association vector of the target monitoring device.
[0014] In some embodiments, the gradient magnitude satisfies the relational expression:
[0015] ;
[0016] wherein, is the th feature map in the training data, and are the identification result and label of the target road image respectively, is the cross-entropy loss of the target road image , is the th gradient magnitude of the feature map.
[0017] In some embodiments, the training method for the image encoder and classifier in the target edge computing unit includes: using the road images of each monitoring device at any historical moment and the labels of the target monitoring device as a set of training data, where the labels include occupying the road for business and not occupying the road for business; inputting the road images into the image encoder to obtain the feature maps of each road image, obtaining the fusion result of all the feature maps, and inputting the fusion result of all the feature maps into the classifier to output the recognition result of the target monitoring device; using the cross-entropy between the recognition result and the label, and the sum of the correlations between multiple feature maps of one road image as the loss function value, and updating the image encoder and the classifier using the gradient descent method; iteratively training the image encoder and the classifier, and completing the training in response to the loss function value being less than the set value.
[0018] Since different monitoring devices are located in different areas and the environments of different areas are different, to avoid inaccurate recognition results of occupying the road for business caused by environmental differences between areas, during the training process, the parameters of the classifiers in different edge computing units are not shared to ensure the accuracy of the recognition results of each target monitoring device.
[0019] In some embodiments, the loss function value satisfies the relational expression:
[0020] ;
[0021] where is the cross-entropy loss of the target road image , is the label of the target road image , is the recognition result of the target road image , is the th feature map of the road image , is the th feature map of the road image , represents calculating the and Hadamard product of, is the number of feature maps corresponding to one road image, is the number of road images in a set of training data, represents the gradient of the cross-entropy loss of the target road image with respect to the th feature map, is the loss function value.
[0022] During the training process of the image encoder and classifier, it is necessary to constrain the independence between multiple feature maps corresponding to a road image, improve the diversity of the feature maps. At the same time, calculate the gradient magnitude of each feature map based on the cross-entropy loss, and use the gradient magnitude of the feature map to accurately measure the degree of association between the feature map and the recognition result of the target road image. Then, perform weighted summation on all feature maps according to the degree of association to achieve the fusion of all feature maps, and further improve the accuracy of the recognition result of the target monitoring device's illegal occupation of the road.
[0023] In some embodiments, obtaining the fusion result of all feature maps includes: setting an initial weighted vector, where the initial weighted vector includes the initial weights of each feature map, and the initial weights of each feature map are the same; calculating the fusion result of all feature maps based on the initial weighted vector, and the fusion result satisfies the relational expression: ; where is the th feature map, is the number of road images in a set of training data, is the number of feature maps corresponding to a road image, represents the total number of all feature maps in a set of training data, is the initial weight, is the fusion result of all feature maps.
[0024] One feature map corresponds to one type of feature in a road image. Weight all feature maps to obtain a fusion result, and the fusion result includes all types of features at different positions on the target road.
[0025] In some embodiments, obtaining the fusion result of all feature maps further includes: performing weighted fusion on each feature map according to the weighted vector of the current training to obtain the fusion result of all feature maps, inputting the fusion result into the classifier in the target edge computing unit to obtain the recognition result of the target road image in the current training; calculating the cross-entropy loss of the target road image in the current training, and calculating the gradient magnitude of each feature map based on the cross-entropy loss; normalizing the gradient magnitudes of all feature maps to obtain the weighted vector for the next training; where, in response to the current training being the first training, the weighted vector of the current training is the initial weighted vector.
[0026] The present invention also provides an edge-computing-based smart city decision-making system, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, the edge-computing-based smart city decision-making method of the present invention is implemented.
[0027] The above-mentioned smart city decision-making method based on edge computing provided by the embodiments of the present application deploys multiple monitoring devices on the target road to collect road images of different regions on the target road, and the trained image encoder and classifier are stored in the edge computing unit of the monitoring device. The parameters of the image encoder in different edge computing units are shared to ensure the same feature extraction process for road images in different regions, and feature maps corresponding to the road images of each region on the target road are obtained. Further, for a target monitoring device, the feature maps are fused according to the correlation degree of all feature maps based on the identification result of illegal occupation of roads by the target monitoring device, and the fused features are input into the classifier stored in the edge computing unit of the target monitoring device to obtain the identification result of illegal occupation of roads in the corresponding region of the target monitoring device. The fused features include all environmental features related to the identification result of illegal occupation of roads by the target monitoring device on the target road, improving the accuracy of the identification result of illegal occupation of roads by the target monitoring device.
[0028] Further, since the regions where different monitoring devices are located are different and the environments of different regions are different, to avoid inaccurate identification results of illegal occupation of roads caused by environmental differences between regions, the parameters of the classifiers in different edge computing units are not shared. Description of the Drawings
[0029] Figure 1 is a flowchart of the smart city decision-making method based on edge computing according to the embodiments of the present application;
[0030] Figure 2 is a schematic flowchart of outputting the identification result of illegal occupation of roads by the target monitoring device;
[0031] Figure 3 is a flowchart of the training method of the image encoder and classifier in the target edge computing unit provided by the preferred embodiment of the present application;
[0032] Figure 4 is a schematic diagram of outputting the identification result of the target road image in the training data provided by the preferred embodiment of the present application;
[0033] Figure 5 is a block diagram of the smart city decision-making system based on edge computing according to the embodiments of the present application. Detailed Embodiments
[0034] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application.
[0035] According to the first aspect of the present application, the present application provides a smart city decision-making method based on edge computing for identifying unlicensed business operations on a target road, where the target road is any road that requires identification of unlicensed business operations. A plurality of monitoring devices are deployed along the target road, and the field of view of all monitoring devices can cover the entire target road, and the position of each monitoring device remains fixed; an edge computing unit is deployed inside each monitoring device, and the edge computing unit has the functions of data storage, data calculation, and data transmission, that is, all monitoring devices can realize data transmission through the edge computing unit.
[0036] Please refer to Figure 1 As shown, it is a flowchart of the smart city decision-making method based on edge computing provided by a preferred embodiment of the present application. According to different requirements, the order of steps in the flowchart can be changed, and some steps can be omitted.
[0037] S11. In an edge computing unit, the trained image encoder extracts features from the real-time road image collected by the monitoring device to obtain the real-time feature map of the monitoring device, where the parameters of the image encoder in different edge computing units are shared.
[0038] In one embodiment, the road image collected by a monitoring device corresponds to an area on the target road, and one monitoring device corresponds to one unlicensed business operation identification result. That is to say, when the unlicensed business operation identification result of a monitoring device is unlicensed business operation, the area with unlicensed business operation on the target road can be determined according to the serial number information of the monitoring device.
[0039] A trained image encoder and a trained classifier are stored in the edge computing unit of a monitoring device.
[0040] Among them, the trained image encoder extracts features from the road image collected by the monitoring device to obtain the feature map corresponding to the monitoring device. One monitoring device corresponds to at least one feature map, and the number of feature maps is related to the network structure of the image encoder. The image encoder can be any existing convolutional neural network such as ResNet, ShuffleNet, or VGGNet. Exemplarily, when the image encoder adopts VGG16, VGG16 performs multiple convolutional operations on the road image with size information of 224×224×3 to achieve feature extraction, and the size information of the output result is 7×7×512. Then the number of feature maps corresponding to the monitoring device is 512, and the size of each feature map is 7×7.
[0041] It should be noted that at the same moment, the road maps collected by different monitoring devices reflect the environmental information of different regions on the target road. The edge computing unit of each monitoring device extracts features from the road images of different regions to obtain the feature maps of different regions on the target road. To maintain the consistency of the feature extraction process for different regions on the target road, the parameters of the image encoder in different edge computing units are shared. Exemplarily, the parameters of the image encoder are the trainable parameters in VGG16.
[0042] That is to say, at any moment, the edge computing unit of each monitoring device will perform the same feature extraction process on the collected road images to obtain the feature maps of each region on the target road at any moment. Exemplarily, if there are 10 monitoring devices on the target road and the image encoder uses VGG16, each monitoring device will obtain 512 feature maps at one moment, and a total of 5120 feature maps will be obtained for the target road. The 5120 feature maps can reflect the environmental information of all regions of the target road at this moment.
[0043] In this way, at the current moment, the real-time road images collected by the monitoring device are input into the trained image encoder in the corresponding edge computing unit, and the real-time feature maps of each region on the target road can be obtained, providing a data basis for subsequent obtaining the occupation of business operation recognition results of the target monitoring device, where the target monitoring device is any one of the monitoring devices on the target road.
[0044] S12. In response to the target monitoring device receiving the real-time feature maps of other monitoring devices, perform weighted summation on all the real-time feature maps based on the set association vector of the target monitoring device to obtain the real-time fusion feature, where the target monitoring device is any one of the monitoring devices on the target road, and the set association vector includes the association degree between the occupation of business operation recognition result of the target monitoring device and each real-time feature map.
[0045] In one embodiment, the edge computing unit deployed inside the target monitoring device is used as the target edge computing unit. The target monitoring device receives the real-time feature maps of other monitoring devices through the target edge computing node, and all the real-time feature maps can reflect the environmental information of all regions on the target road at the current moment.
[0046] In an alternative embodiment, the set association vector includes the association degree between the occupation of business operation recognition result of the target monitoring device and each real-time feature map; since different monitoring devices are located in different regions on the target road, the set association vectors corresponding to different monitoring devices are also different.
[0047] Exemplarily, at the current moment, a total of 5,120 real-time feature maps are obtained for the target road, and the set association vector includes 5,120 values, where one value represents the degree of association between a real-time feature map and the recognition result of illegal occupation of roads by vendors of the target monitoring device.
[0048] Among them, the real-time fusion feature satisfies the relational expression:
[0049] ;
[0050] Among them, is the th real-time feature map, is the number of real-time road images, is the number of real-time feature maps corresponding to one real-time road image, is the degree of association between the recognition result of illegal occupation of roads by vendors of the target monitoring device and the th real-time feature map, is the real-time fusion feature.
[0051] In one embodiment, the method for obtaining the set association vector of the target monitoring device includes: obtaining at least one set of training data, where the training data includes road images of all monitoring devices and labels of the target road image at any historical moment, and the labels are illegal occupation of roads by vendors or non-illegal occupation of roads by vendors; for one set of training data, obtaining the feature maps of all road images according to the trained image encoder, calculating the mean value of all feature maps, and inputting it into the trained classifier in the target edge computing unit to obtain the recognition result of the target road image; calculating the cross-entropy loss between the recognition result of the target road image and the label, and calculating the gradient magnitude of each feature map based on the cross-entropy loss, where the gradient magnitude satisfies the relational expression:
[0052] ;
[0053] Among them, is the th feature map in the training data, and are respectively the recognition result and the label of the target road image , is the cross-entropy loss of the target road image , is the th gradient magnitude of the feature map; normalizing the gradient magnitudes of all feature maps to obtain the normalized vector corresponding to the training data, and calculating the average value of the normalized vectors corresponding to all training data to obtain the set association vector of the target monitoring device.
[0054] Among them, the gradient magnitude of each feature map is used to reflect the degree of association between the feature map and the recognition result of the target road image. The larger the gradient magnitude, the more effective the information of the feature map is for obtaining the correct recognition result of the target road image, that is, the greater the degree of association between the feature map and the recognition result of the occupation of road by unlicensed vendors of the target monitoring device.
[0055] In this way, all real-time feature maps are fused according to the degree of association between the recognition result of the occupation of road by unlicensed vendors of the target monitoring device and each real-time feature map to obtain a real-time fusion feature; the real-time fusion feature includes all environmental features related to the recognition result of the occupation of road by unlicensed vendors of the target monitoring device on the target road. Classifying according to the real-time fusion feature can improve the accuracy of the recognition result of the occupation of road by unlicensed vendors of the target monitoring device.
[0056] S13. Input the real-time fusion feature into the classifier trained in the target edge computing unit to output the recognition result of the occupation of road by unlicensed vendors of the target monitoring device, where the target edge computing unit is an edge computing unit deployed inside the target monitoring device, and the parameters of the classifiers in different edge computing units are not shared.
[0057] In one embodiment, the classifier trained in the target edge computing unit is used to obtain the recognition result of the target monitoring device. The classifier performs dimensionality transformation on the real-time fusion feature to output the recognition result of the occupation of road by unlicensed vendors of the target monitoring device, and the recognition result of the occupation of road by unlicensed vendors is occupation of road by unlicensed vendors or non-occupation of road by unlicensed vendors; where the classifier includes a fully connected neural network and a classification function, and the classification function is the softmax function.
[0058] It should be noted that since different monitoring devices are located in different regions (the environments in different regions are different), when different monitoring devices perform the recognition of the occupation of road by unlicensed vendors, the corresponding classification logics will also be different. To ensure that each monitoring device can obtain an accurate recognition result of the occupation of road by unlicensed vendors, the parameters of the classifiers in different edge computing units are not shared. The parameters of the classifier are the trainable parameters in the fully connected neural network.
[0059] Exemplarily, please refer to Figure 2, which is a schematic flowchart of the process for outputting the occupation of road operation recognition results of the target monitoring device according to an embodiment of the present application. 10 monitoring devices are deployed on the target road, and different monitoring devices can collect road images of different regions. If the target monitoring device is monitoring device 5, in order to output the occupation of road operation recognition result of the target monitoring device, the target edge computing node receives the real-time feature maps of other monitoring devices except monitoring device 5, and fuses all the real-time feature maps according to the set association vector of monitoring device 5 to obtain a real-time fused feature; the real-time fused feature is input into the classifier in the target edge computing node to obtain the occupation of road operation recognition result of monitoring device 5, and the occupation of road operation recognition result of monitoring device 5 can reflect whether there is an occupation of road operation behavior in the area corresponding to monitoring device 5.
[0060] It can be understood that at the same moment, each monitoring device will obtain the occupation of road operation recognition result of the corresponding area.
[0061] S14, formulating management measures based on the occupation of road operation recognition result.
[0062] In one embodiment, in response to the occupation of road operation recognition result being an occupation of road operation, the location information of the monitoring device corresponding to the occupation of road operation recognition result is sent to the management personnel. The management personnel can timely receive the location information of the occupation of road operation behavior in the target road and timely formulate management measures to stop the occupation of road operation behavior.
[0063] In this way, the recognition of the occupation of road operation on the target road is realized. Based on the monitoring devices deployed on the target road and the edge computing units in the monitoring devices, the location information of the occupation of road operation behavior on the target road can be quickly and accurately obtained, and the location information is timely sent to the management personnel.
[0064] In one embodiment, in order to realize the recognition of the occupation of road operation on the target road, it is necessary to train the image encoder and classifier of the edge computing unit; the parameters of the image encoders in different edge computing units are shared, and the parameters of the classifiers in different edge computing units are not shared. Taking the target edge computing unit as an example, the training methods of the image encoder and classifier are introduced in detail. The target edge computing unit is the edge computing unit deployed inside the target monitoring device, and the target monitoring device is any one of the multiple monitoring devices deployed on the target road. Please refer to Figure 3 As shown, it is a flowchart of the training methods of the image encoder and classifier in the target edge computing unit provided by a preferred embodiment of the present application. The training methods include steps S21 to S24.
[0065] S21. Collect the road images of all monitoring devices at the same historical moment, and obtain the labels of the target road images, thereby obtaining a set of training data. The labels include illegal occupation of business areas and non-illegal occupation of business areas. The target road image is the road image collected by the target monitoring device.
[0066] In one embodiment, at any historical moment, collect the road images of all monitoring devices at the historical moment, use the road image collected by the target monitoring device as the target road image, and the target road image corresponds to a label, which is either illegal occupation of business areas or non-illegal occupation of business areas; in this way, a set of training data for the target edge computing unit is obtained.
[0067] Exemplarily, there are 10 monitoring devices on the target road, and the target monitoring device is monitoring device 5; at the historical moment The collected training data is , where represents the road image of the 2nd monitoring device at the historical moment ; represents the road image of the 1st monitoring device at the historical moment ; represents the road image of the 3rd monitoring device at the historical moment ; represents the road image of the 4th monitoring device at the historical moment ; represents the road image of the 6th monitoring device at the historical moment ; represents the road image of the 7th monitoring device at the historical moment ; represents the road image of the 8th monitoring device at the historical moment ; represents the road image of the 9th monitoring device at the historical moment ; represents the road image of the 10th monitoring device at the historical moment ; represents the target road image corresponding to the target monitoring device at the historical moment ; represents the label of the target monitoring device at the historical moment .
[0068] In this way, according to the road images of all monitoring devices at multiple historical moments, multiple sets of training data can be collected.
[0069] S22. Input all the road images in a set of training data into the image encoder to obtain the feature maps of each road image, and obtain the fusion result of all the feature maps, and input it into the classifier in the target edge computing unit to output the recognition result of the target road image.
[0070] In one embodiment, all road images in a set of training data include the target road image of the target monitoring device and the road images of other monitoring devices outside the target monitoring device. Please refer to Figure 4 , which is a schematic diagram of outputting the recognition result of the target road image in the training data provided by the preferred embodiment of the present application. All road images in a set of training data are sent into the image encoder in the corresponding edge computing unit to output the feature map of each road image. One road image corresponds to a feature map, where . It can be understood that the parameters of the image encoders in different edge computing units are shared, that is, both the target road image and the other road images outside the target road image have undergone the same feature extraction operation. Further, obtain the fusion result of all feature maps and input it into the classifier in the target edge computing unit to output the recognition result of the target road image.
[0071] In an alternative embodiment, the obtaining the fusion result of all feature maps includes: setting an initial weighted vector, the initial weighted vector includes the initial weight of each feature map, and the initial weights of each feature map are the same; calculating the fusion result of all feature maps based on the initial weighted vector, and the fusion result satisfies the relational expression:
[0072]
[0073] where, is the th feature map, is the number of road images in a set of training data, is the number of feature maps corresponding to one road image, represents the number of all feature maps in a set of training data, is the initial weight, is the fusion result of all feature maps.
[0074] In another alternative embodiment, the obtaining the fusion result of all feature maps includes: performing weighted fusion on all feature maps according to the weighted vector of the current training to obtain the fusion result of all feature maps; inputting the fusion result into the classifier in the target edge computing unit to obtain the recognition result of the target road image in the current training; calculating the cross-entropy loss of the target road image in the current training and calculating the gradient magnitude of each feature map based on the cross-entropy loss; normalizing the gradient magnitudes of all feature maps to obtain the weighted vector for the next training; in response to the current training being the first training, the weighted vector of the current training is the initial weighted vector.
[0075] Thus, based on a set of training data, the recognition result of the target road image is obtained, and the recognition result is the recognition result of occupying the road for business predicted by the classifier.
[0076] S23. Calculate the loss function value, and use the gradient descent method to update the image encoder and the classifier to complete one training.
[0077] In one embodiment, the loss function value satisfies the relational expression:
[0078]
[0079] where is the cross-entropy loss of the target road image is the cross-entropy loss of the target road image is the label of the target road image is the recognition result of the target road image is the recognition result of the target road image is the recognition result of the target road image is the nth feature map of the road image is the nth feature map, is the mth feature map of the road image is the mth feature map, represents the calculation of and the Hadamard product of, is the number of feature maps corresponding to a road image, is the number of road images in a set of training data, represents the gradient of the cross-entropy loss of the target road image with respect to the nth feature map, is the gradient of the cross-entropy loss of the target road image is the loss function value.
[0080] where is the cross-entropy loss, which is used to constrain that the recognition result of the target road image is the same as the label of the target road image so that the classifier and the image encoder learn the mapping relationship between the road image and the recognition result of occupying the road for business, and obtain the correct recognition result of the target road image is obtained.
[0081] where is used to reflect the correlation between the nth feature map and the mth feature map of the road image The closer its value is to 0, the smaller the correlation between the nth feature map and the mth feature map corresponding to the road image is. The closer its value is to 0, the smaller the correlation between the nth feature map and the mth feature map corresponding to the road image is. The closer its value is to 0, the smaller the correlation between the nth feature map and the mth feature map corresponding to the road image is. The closer its value is to 0, the smaller the correlation between the nth feature map and the mth feature map corresponding to the road image is. The closer its value is to 0, the smaller the correlation between the nth feature map and the mth feature map corresponding to the road image is. As part of the loss function value, it is used to constrain the road image corresponding The feature maps are independent of each other, that is, one feature map corresponds to one kind of feature, ensuring that the image encoder can learn different kinds of features in a road image and improving the diversity of the feature maps.
[0082] Among them, It is used to make the gradient of the output of the classifier with respect to the irrelevant feature maps become sparse, thereby reducing the degree of dependence of the recognition result of the target road image on the irrelevant feature maps; is the magnitude of the gradient of the cross-entropy loss of the target road image with respect to the th feature map, which is used to characterize the relevance of the feature map to the recognition result of the target road image That is, it can be used to characterize the recognition effectiveness of the feature map for the recognition result of the target road image and realize the accurate quantification of the recognition effectiveness of the recognition result of the target road image for each feature map.
[0083] In one embodiment, the image encoder and the classifier are updated using the gradient descent method to complete one training.
[0084] It can be understood that due to the parameter sharing of the image encoder in different edge computing units and the non-sharing of the parameters of the classifier in different edge computing units, when training the image encoder and the classifier in the target edge computing unit, the classifier of the target edge computing unit and the image encoders in all edge computing units can be updated.
[0085] S24, iteratively train the image encoder and the classifier, and in response to the loss function value being less than the set value, obtain the trained image encoder and classifier in the target edge computing unit.
[0086] In one embodiment, new training data is continuously collected, and the image encoder and the classifier are iteratively updated until the loss function value is less than the set value and then stop, to obtain the trained image encoder and classifier in the target edge computing unit. Among them, the set value is taken as 0.01.
[0087] In this way, the training of the image encoder and the classifier in the target edge computing unit is completed according to the method from step S21 to step S24; the training of the image encoder and the classifier in the edge computing unit corresponding to each monitoring device can be obtained in the same way. Since the parameters of the image encoders in different edge computing units are shared and the parameters of the classifiers in different edge computing units are not shared, a total of 1 image encoder and N classifiers need to be trained for the N monitoring devices on the target road.
[0088] The present application also provides a smart city decision-making system based on edge computing. Figure 5 It is a block diagram of the smart city decision-making system based on edge computing according to an embodiment of the present application. As Figure 5 shown, the device 50 includes a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, a smart city decision-making method based on the first aspect of the present application is implemented. The device 50 also includes other components well known to those skilled in the art such as a communication bus and a communication interface, and their settings and functions are known in the art, so they will not be described in detail here.
[0089] In the present application, the foregoing memory can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or device. For example, the computer-readable storage medium can be any suitable magnetic storage medium or magneto-optical storage medium, such as, resistive random access memory RRAM (Resistive Random Access Memory), dynamic random access memory DRAM (Dynamic Random Access Memory), static random access memory SRAM (Static Random-Access Memory), enhanced dynamic random access memory EDRAM (Enhanced Dynamic Random Access Memory), high-bandwidth memory HBM (High-Bandwidth Memory), hybrid memory cube HMC (Hybrid Memory Cube), etc., or any other medium that can be used to store the required information and can be accessed by an application, a module, or both. Any such computer storage medium can be part of the device or accessible or connectable to the device. Any application or module described in the present application can be implemented using computer-readable / executable instructions that can be stored or otherwise held by such a computer-readable medium.
[0090] It should be noted that, for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application shall be subject to the appended claims.
Claims
1. A smart city decision-making method based on edge computing, characterized in that: For identifying road occupation business on a target road, where a plurality of monitoring devices including edge computing units are deployed on the target road, the method comprises: In one edge computing unit, a trained image encoder extracts features from a real-time road image collected by a monitoring device to obtain a real-time feature map of the monitoring device, wherein parameters of the image encoders in different edge computing units are shared; In response to the target monitoring device receiving the real-time feature map of other monitoring devices, a set association vector of the target monitoring device is obtained: The road images of each monitoring device at any historical moment and the labels of the target monitoring device are used as a set of training data, wherein the labels include road occupation and non-road occupation; For a set of training data, the feature maps of all road images are obtained according to the trained image encoder, and the mean of all feature maps is input into the trained classifier in the target edge computing unit to obtain the recognition result of the target road image; Calculating the cross entropy loss of the recognition result and the label corresponding to the target road image, and calculating the gradient size of each feature map based on the cross entropy loss; Normalizing the gradient sizes of all feature maps to obtain normalized vectors corresponding to the training data, calculating the average value of the normalized vectors corresponding to all training data, and obtaining the set association vector of the target monitoring device; Based on the set association vector of the target monitoring device, weighted summation is performed on each real-time feature map to obtain a real-time fusion feature, wherein the target monitoring device is any one of the monitoring devices, and the set association vector includes the association degree between the road occupation business identification result of the target monitoring device and each real-time feature map; The real-time fusion features are input into the trained classifier in the target edge computing unit to output the road-blocking business identification result of the target monitoring device, wherein the target edge computing unit is an edge computing unit within the target monitoring device, and the parameters of the classifiers in different edge computing units are not shared; management measures are formulated based on the road-blocking business identification result.
2. A smart city decision-making method based on edge computing as claimed in claim 1, characterized in that: The image encoder is a convolutional neural network, the classifier includes a fully connected neural network and a classification function, and the classification function is a softmax function.
3. The smart city decision-making method based on edge computing as claimed in claim 1, characterized in that: The real-time fusion feature satisfies the relationship: ; in, For the Real-time feature map, is the number of real-time road images, is the number of real-time feature maps corresponding to a real-time road image, The road occupation identification result of the target monitoring device is The correlation degree of the real-time feature graph, Real-time feature fusion.
4. A smart city decision-making method based on edge computing as claimed in claim 3, characterized in that: The gradient size satisfies the relationship: ; in, The first Zhang feature map, and The target road images are The recognition results and labels of The target road image The cross entropy loss is For the The gradient size of the feature map.
5. The smart city decision-making method based on edge computing as claimed in claim 1, characterized in that: The training methods of the image encoder and classifier in the target edge computing unit include: Input the road image into an image encoder to obtain a feature map of each road image, obtain a fusion result of all feature maps, and input the fusion result of all feature maps into a classifier to output a recognition result of the target monitoring device; A loss function value is obtained according to the cross entropy of the recognition result and the label, and the image encoder and the classifier are updated using the gradient descent method; the image encoder and the classifier are iteratively trained, and the training is completed in response to the loss function value being less than a set value.
6. A smart city decision-making method based on edge computing as claimed in claim 5, characterized in that: The loss function value satisfies the relationship: ; in, The target road image The cross entropy loss is The target road image Tags, The target road image The recognition result of For road images No. Zhang feature map, For road images No. Zhang feature map, Representation calculation and The Hadamard is the number of feature maps corresponding to a road image, is the number of road images in a set of training data, Represents the target road image The cross entropy loss has an effect on The gradient of the feature map, is the loss function value.
7. A smart city decision-making method based on edge computing as claimed in claim 5, characterized in that: The obtaining of the fusion results of all feature maps includes: Setting an initial weight vector, wherein the initial weight vector includes an initial weight of each feature map, and the initial weight of each feature map is the same; The fusion results of all feature maps are calculated based on the initial weighted vector, and the fusion results satisfy the relationship: ; in, For the Zhang feature map, is the number of road images in a set of training data, is the number of feature maps corresponding to a road image, represents the number of all feature maps in a set of training data, is the initial weight, is the fusion result of all feature maps.
8. A smart city decision-making method based on edge computing as claimed in claim 7, characterized in that: The step of obtaining the fusion results of all feature maps further includes: Perform weighted fusion on each feature map according to the weighted vector of the current training to obtain the fusion result of all feature maps, input the fusion result into the classifier in the target edge computing unit, and obtain the recognition result of the target road image in the current training; Calculating the cross entropy loss of the target road image in the current training, and calculating the gradient size of each feature map based on the cross entropy loss; The gradient sizes of all feature maps are normalized to obtain a weight vector for the next training; wherein, in response to the current training being the first training, the weight vector for the current training is an initial weight vector.
9. A smart city decision-making system based on edge computing, characterized in that: It comprises a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a smart city decision-making method based on edge computing according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Vision-based illegal operation detection method and system
CN114155217A
Urban management illegal behavior recognition system and method based on large model
CN119091504A