Smart city management event prediction methods, devices, equipment and storage media
By employing a pose detection model based on global semantic information completion and attention collaboration mechanisms, the problem of misassociation in target overlap detection in smart city management is solved, achieving high-precision prediction of target objects and abnormal behaviors, and ensuring timely response to urban management events.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2026-03-13
AI Technical Summary
In existing smart city management, manual detection is uncertain and it is difficult to effectively distinguish overlapping targets, resulting in a high rate of false association of key target points, low detection accuracy, and inability to respond to abnormal behavior in a timely manner.
A pre-trained pose detection model combining global semantic information completion and attention collaboration mechanism is adopted. Through the target detection, global semantic information completion and attention collaboration mechanism modules, a key point heatmap of the target object is generated, and the prediction results of the target object and abnormal behavior are output. Event prediction is performed by combining it with the urban management matching model.
It improves the detection accuracy of target objects and abnormal behaviors, enables accurate prediction of urban management events, reduces the false association rate of key points, and improves the accuracy of event prediction.
Smart Images

Figure CN120580637B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart city management technology, and in particular to a method, apparatus, equipment, medium and computer program product for predicting events in smart city management. Background Technology
[0002] Currently, in the construction of smart cities, the management of urban sites mainly relies on the detection of target areas or on-site patrols by urban management law enforcement personnel.
[0003] Current technologies typically require manual detection of target areas. When abnormal crowd behavior is detected in a target area, manual detection is necessary before alerts are issued and appropriate management strategies are implemented. However, manual detection is inherently uncertain. If an anomaly has already occurred in the target area but relevant personnel fail to notice it, the severity of the anomaly may increase, making management more difficult. Furthermore, existing target detection models for smart city management cannot effectively distinguish overlapping targets, exhibiting a high rate of false associations of key points and low detection accuracy. Therefore, providing a posture detection-based smart city management event prediction method that effectively distinguishes overlapping targets, reduces the rate of false associations of key points, and improves the detection accuracy of target objects and abnormal behaviors is a pressing technical problem that needs to be solved. Summary of the Invention
[0004] Based on this, it is necessary to provide a smart city management event prediction method, device, equipment, medium, and computer program product that can effectively distinguish overlapping targets, reduce the false association rate of key points of targets, and improve the detection accuracy of target objects and abnormal behaviors, in order to address the above-mentioned technical problems.
[0005] Firstly, this application provides a method for predicting events in smart city management. The method includes:
[0006] Acquire real-time image data of the target city's management and monitoring area;
[0007] Based on the real-time image data, a pre-trained pose detection model combining global semantic information completion and attention coordination mechanism is used to output the urban management prediction result at the first preset time. The urban management prediction result includes the target object prediction result and the abnormal behavior prediction result of the target object.
[0008] Based on the target object prediction results and the abnormal behavior prediction results at the first preset time, the urban management event prediction results are output through the urban management matching model.
[0009] In one embodiment, the pre-trained pose detection model combining global semantic information completion and attention coordination mechanism includes a target detection module, multiple global semantic information completion modules, an attention coordination mechanism module, and an output module.
[0010] The target detection module detects multiple target object image features based on the real-time image data;
[0011] The global semantic information completion module generates a corresponding target object attention feature map based on the image features of each target object;
[0012] The attention collaboration mechanism module obtains a heatmap of key points of the target object in a specified region by completing different target object attention feature maps of different global semantic information completion modules.
[0013] The output module obtains the target object's pose result based on the target object's key point heatmap, and generates the target object prediction result and abnormal behavior prediction result based on the target object's pose result.
[0014] In one embodiment, the global semantic information completion module includes a first backbone network branch and a first global semantic information completion branch corresponding to the first backbone network branch, or a second backbone network branch and a second global semantic information completion branch corresponding to the second backbone network branch;
[0015] The first global semantic information completion branch eliminates the background response of the target object features extracted from the first backbone network branch;
[0016] The second global semantic information completion branch enhances the target object key point region corresponding to the target object features extracted by the second backbone network branch.
[0017] In one embodiment, the target object prediction result includes target object density prediction information. The step of outputting urban management event prediction results based on the target object prediction result and the abnormal behavior prediction result at a first preset time point, through an urban management matching model, includes:
[0018] Based on the target object density prediction information and the abnormal behavior prediction results at the first preset time, the first urban management event and the first urban management emergency level corresponding to the first urban management event are matched through the urban management matching model.
[0019] Output the first urban management incident and the first urban management emergency level.
[0020] In one embodiment, the method further includes:
[0021] By using a pre-trained pose detection model that combines global semantic information completion and attention coordination mechanism, the city management prediction results at the first and second preset times are output.
[0022] Based on the target object prediction results and abnormal behavior prediction results at the first preset time and the second preset time, the target object trend information and behavior trend information are output through the urban management matching model.
[0023] In one embodiment, the method further includes:
[0024] Based on the target object trend information and the behavioral trend information, the city management matching model is used to match the associated second urban management event and the second urban management emergency level corresponding to the second urban management event.
[0025] Output the second urban management event and the second urban management emergency level.
[0026] Secondly, this application also provides a smart city management event prediction device. The device includes:
[0027] The data acquisition module is used to acquire real-time image data of the target city's management and monitoring area;
[0028] The multi-person human posture detection module is used to output the urban management prediction result at the first preset time based on the real-time image data and through a pre-trained posture detection model that combines global semantic information completion and attention collaboration mechanism. The urban management prediction result includes the target object prediction result and the abnormal behavior prediction result.
[0029] The urban management matching and prediction module is used to output urban management event prediction results based on the prediction results of the target object and the prediction results of the abnormal behavior at a first preset time, through the urban management matching model.
[0030] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0031] Acquire real-time image data of the target city's management and monitoring area;
[0032] Based on the real-time image data, a pre-trained pose detection model combining global semantic information completion and attention coordination mechanism is used to output the urban management prediction result at the first preset time. The urban management prediction result includes the target object prediction result and the abnormal behavior prediction result of the target object.
[0033] Based on the target object prediction results and the abnormal behavior prediction results at the first preset time, the urban management event prediction results are output through the urban management matching model.
[0034] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0035] Acquire real-time image data of the target city's management and monitoring area;
[0036] Based on the real-time image data, a pre-trained pose detection model combining global semantic information completion and attention coordination mechanism is used to output the urban management prediction result at the first preset time. The urban management prediction result includes the target object prediction result and the abnormal behavior prediction result of the target object.
[0037] Based on the target object prediction results and the abnormal behavior prediction results at the first preset time, the urban management event prediction results are output through the urban management matching model.
[0038] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0039] Acquire real-time image data of the target city's management and monitoring area;
[0040] Based on the real-time image data, a pre-trained pose detection model combining global semantic information completion and attention coordination mechanism is used to output the urban management prediction result at the first preset time. The urban management prediction result includes the target object prediction result and the abnormal behavior prediction result of the target object.
[0041] Based on the target object prediction results and the abnormal behavior prediction results at the first preset time, the urban management event prediction results are output through the urban management matching model.
[0042] The embodiments of this application have the following beneficial effects:
[0043] The smart city management event prediction method, device, equipment, medium, and computer program product provided in this application can distinguish the key points of adjacent target objects and adjacent pedestrians by using a pre-trained posture detection model that combines global semantic information completion and attention coordination mechanism. This improves the positioning accuracy of key points of target objects or target pedestrians with small targets, thereby improving the prediction accuracy of target object prediction results and abnormal behavior prediction results, and thus accurately predicting urban management events. Attached Figure Description
[0044] Figure 1 This is a flowchart illustrating a smart city management event prediction method in one embodiment;
[0045] Figure 2 This is a schematic diagram of the target object image processing flow of the pose detection model in one embodiment;
[0046] Figure 3 This is a schematic diagram of the pose detection model that combines global semantic information completion and attention collaboration mechanism in one embodiment.
[0047] Figure 4 This is a schematic diagram of the feature extraction process for global semantic information completion in one embodiment;
[0048] Figure 5 This is a schematic diagram of a bypass branch structure for global semantic information completion in one embodiment;
[0049] Figure 6 This is a structural block diagram of a smart city management event prediction device in one embodiment. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0051] The smart city management event prediction method provided in this application can be applied to terminals or servers. The terminal communicates with the server via a network. A data storage system can store the data that the server needs to process. The data storage system can be integrated onto the server or located in the cloud or on other network servers. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart vehicle devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. The server can be implemented using a standalone server or a server cluster consisting of multiple servers.
[0052] Example 1
[0053] In one embodiment, such as Figure 1 As shown, a method for predicting events in smart city management is provided, including:
[0054] S1. Acquire real-time image data of the target city's management and monitoring area;
[0055] S2. Based on real-time image data, through a pre-trained pose detection model that combines global semantic information completion and attention coordination mechanism, output the urban management prediction results at the first preset time. The urban management prediction results include the target object prediction results and the abnormal behavior prediction results of the target object.
[0056] S3. Based on the target object prediction results and abnormal behavior prediction results at the first preset time, output the urban management event prediction results through the urban management matching model.
[0057] Specifically, in real-time image data of the target urban management monitoring area, there may be backgrounds (such as railings or trees) with similar shapes or textures to the target object (e.g., pedestrians). Traditional models relying on local features are easily misled, leading to incorrect keypoint localization. Furthermore, in crowded scenes, body parts of different individuals and objects may overlap or intersect. By combining global semantic information for completion, keypoints of adjacent objects and pedestrians can be distinguished. A pose detection model incorporating an attention-based collaborative mechanism can obtain the final feature map, which is then used to calculate a heatmap of the target object / pedestrian keypoints. The location of the target object / pedestrian keypoint in the network output space is obtained by shifting the position from the maximum response location to the second-largest response location in the heatmap by a certain degree (e.g., one-quarter). This location is then mapped from the network output space to the input image space, and the keypoints in the input image space are connected according to the target human body structure to obtain the target object's pose. The attention-based collaborative mechanism can improve the localization accuracy of small target keypoints. The first preset time can be the current time, a historical time, or a predicted future time. Based on the target object prediction results and abnormal behavior prediction results at the first preset time point, an urban management matching model can be combined to output urban management event prediction results. By adopting this technical solution, a pre-trained pose detection model that combines global semantic information completion and attention coordination mechanisms can distinguish the key points of adjacent target objects and adjacent pedestrians, improving the localization accuracy of key points of small target objects or target pedestrians. This, in turn, improves the prediction accuracy of target object prediction results and abnormal behavior prediction results, thereby accurately predicting urban management events.
[0058] In one embodiment, the pre-trained pose detection model combining global semantic information completion and attention collaboration mechanism includes an object detection module, multiple global semantic information completion modules, an attention collaboration mechanism module, and an output module. Specifically, the object detection module detects multiple target object image features based on real-time image data; the global semantic information completion module generates a corresponding target object attention feature map based on each target object image feature; the attention collaboration mechanism module obtains a target object keypoint heatmap for a specified region based on the different target object attention feature maps from different global semantic information completion modules; and the output module obtains the target object pose result based on the target object keypoint heatmap, and generates a target object prediction result and an abnormal behavior prediction result based on the target object pose result.
[0059] For example, the dataset used by the pose detection model can be the COCO dataset, which contains complex and varied human behaviors. The evaluation metric for the pose detection model is based on Object Keypoint Similarity (OKS), and the specific formula is as follows:
[0060]
[0061] In the formula, It is the Euclidean distance between the ground truth annotation of the i-th human body keypoint and the location of the detected i-th human body keypoint. This is the visibility annotation of the truth value of this human body keypoint; After one standard deviation The Gaussian expression is given by , where s is the square root of the area occupied by the target pedestrian. It is a constant that controls the decay of the i-th human body key point. function, when When, the function value is 1; when When the value is zero, the function value is 0. This function represents human keypoints that are detected but not labeled, and does not affect the similarity of the target keypoint. Each labeled human keypoint will generate a keypoint similarity score between 0 and 1. These keypoint similarities are averaged over the total number of labeled keypoints on the human body. The optimal prediction will be... For all key points on the human body, the deviation exceeds the standard deviation. According to the prediction, OKS tends to be 0.
[0062] For example, refer to Figure 2Pose detection models can include ViTPose and HRNet-W48. Both require the bounding box of the target object as additional input. ViTPose uses Vision Transformer (ViT) for image feature extraction, while HRNet-W48 uses HRNet (High-Resolution Network). Both use convolution operations to convert the features into Gaussian heatmaps, which are then decoded using distribution-aware coordinates to obtain the coordinates of 17 human keypoints in COCO format. These are the nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, and right ankle. For each target pedestrian, the structure of the human keypoint annotation is {x1, y1, v1, ..., xk, yk, vk}, where x and y are the positions of the human keypoints, and v is a visibility label. v=0 indicates that the human body keypoint is not labeled, v=1 indicates that the human body keypoint is labeled but not visible, and v=2 indicates that the human body keypoint is labeled and visible. Each labeled target also has a dimension s, which is the square root of the area occupied by the target pedestrian in the image. The bounding box coordinates of the target pedestrian are also provided. The test set only provides image data, without ground truth annotations, and does not provide the location of pedestrians in the images. The pose estimation model needs to simultaneously detect target pedestrians and locate their human body keypoints. The generated prediction results are uploaded to the evaluation server for evaluation of the pose estimation model.
[0063] For example, refer to Figure 3 First, a Faster R-CNN (Faster Region-based Convolutional Neural Network) object detection model detects pedestrians and uses bounding boxes to determine their locations. Next, pedestrian images are cropped from the original images and scaled to a fixed size before being input into a pose estimation network to extract visual features. A human global semantic information bypass branch eliminates background responses to features and emphasizes key human regions in the feature map. Then, an attention-based collaborative mechanism is used to obtain the final feature map, which is used to calculate a heatmap of human key points. The locations of the human key points in the network output space are obtained by shifting the heatmap from the location of the maximum response to the location of the second-largest response by one-quarter. These locations are then mapped from the network output space to their positions in the input image space, and the key points in the input image space are connected according to human structure to obtain the pose result.
[0064] For example, the attention collaboration mechanism module can take four feature maps of different levels as input, first perform element-wise addition to obtain feature map X. First, a 1×1 convolution is used to fuse the features of the four different levels along the channels, then two 3×3 convolutions are used to capture range dependencies. After each convolution operation, a ReLU (Rectified Linear Unit) activation function is used, and finally, a Sigmoid activation function is applied to obtain the spatial attention map Satt, as shown in the following formula:
[0065] Satt=
[0066] In the formula, W1, W3, and W3' represent the weights of the 1×1 convolution and two 3×3 convolutions in the information collaboration module, respectively, and * indicates the convolution operation. In the information collaboration module, a full average pooling operation is also performed to obtain the channel descriptor Z, where each channel Zc is calculated from the entire spatial features of the c-th channel of the feature map X, as shown in the following formula:
[0067]
[0068] In the formula, H and W represent the height and width of feature map X, and Xc refers to the entire spatial feature on the c-th channel of feature map X.
[0069] The channel attention map Catt is obtained after two fully connected layers and a sigmoid activation function, and is calculated as follows:
[0070] Catt=
[0071] In the formula, and This represents the weights of the two fully connected layers in the information collaboration module. Finally, the obtained attention maps Satt and Catt are multiplied by the Hadamard product with the feature map X, and then element-wise added to the feature map X to obtain the final feature map used for human keypoint localization.
[0072] Table 1. Comparison of experimental results (%) of different models on the COCO dataset
[0073] Model backbone network Input dimensions AP AP.50 AP.75 APM APL AR Mask-RCNN ResNet-FPN - 63.1 87.3 68.7 57.8 71.4 - G-RMI ResNet-101 353×257 64.9 85.5 71.3 62.3 70.0 69.7 CFN - - 72.6 86.1 69.7 78.3 64.1 - Simple Base line ResNet-152 384×288 73.7 91.9 81.1 70.3 80.0 79.0 Global semantic information completion model ResNet-152 384×288 74.3 91.7 81.7 70.7 80.5 79.5
[0074] For example, the distance between the detected target keypoint similarity results and the ground truth labels can be used to evaluate the performance of a pose estimation model. Detection results in an image are ranked by confidence scores, and ground truth labels with the highest target keypoint similarity are assigned to the detected human keypoints. After matching, matches with low confidence scores are discarded. After matching, target keypoint similarities greater than a set threshold are considered true positives (TP), and those less than the threshold are considered false positives (FP). Targets without a match are considered false negatives (FN). Precision and recall can be calculated from these values. Average Precision (AP) is calculated based on the similarity threshold of 10 target keypoints. Above, calculate 11 recall values. The corresponding average accuracy. For example... When the threshold The AP obtained at that time. The AP is obtained when the target pedestrian pixel area size is between 322 and 962. This is calculated when the pixel area of the target pedestrian is greater than 962. AR is the maximum recall of a fixed number of detected objects in each image, based on a similarity threshold of 10 target keypoints. The average value is calculated. These evaluation metrics will be used to analyze the object detection experimental results. Referring to Table 1, which shows the performance comparison results of the global semantic information completion model and other methods on the COCO dataset. Mask-RCNN uses predicted bounding boxes to crop features from the feature map, which is not ideal for locating human keypoints. G-RM uses heatmaps and offset values to address the FP problem caused by cluttered backgrounds, without considering the influence of human information. CFN outputs estimation results uniformly from multiple detectors of different levels; if the detectors based on low-level features have large biases, it will adversely affect the final result. In contrast, the human global semantic information model supplements the network backbone with global human semantic information, suppressing background responses similar to human regions in the feature map, alleviating the cluttered background problem, and significantly improving model performance.
[0075] By adopting this technical solution, the background response of features can be eliminated by global semantic information completion and the key points of human body in the feature map can be emphasized, and the key points of adjacent target objects and adjacent pedestrians can be distinguished. The localization accuracy of key points of target objects or target pedestrians with small targets can be improved by attention collaboration mechanism. Furthermore, the prediction accuracy of target object prediction results and abnormal behavior prediction results can be improved by combining the pose model of global semantic information completion and attention collaboration mechanism, thereby accurately predicting urban management events.
[0076] In one embodiment, the global semantic information completion module includes a first backbone network branch and a first global semantic information completion branch corresponding to the first backbone network branch, or a second backbone network branch and a second global semantic information completion branch corresponding to the second backbone network branch; the first global semantic information completion branch eliminates the background response of the target object features extracted by the corresponding first backbone network branch; the second global semantic information completion branch enhances the target object key point region of the target object features extracted by the corresponding second backbone network branch.
[0077] Specifically, refer to Figure 4 Attention maps can be computed in the shallow layers of the network's global semantic information bypass branches to eliminate background responses to features extracted from the corresponding main branches. In the deeper layers of the network, the global semantic information bypass branches emphasize key human point regions in the feature maps.
[0078] For example, refer to Figure 5 The human global semantic information bypass branch, given an input feature map X, obtains a descriptor F(X) through a bottom-up and top-down structure. Then, it generates an attention map Φ(X) through a sigmoid activation function, as shown in the following formula:
[0079]
[0080] Then, Φ(X) is used to enhance the feature map f(X) extracted from the corresponding backbone branch of the network, as shown in the following formula:
[0081]
[0082] The operator refers to performing the Hadamard product along the dimensions of the attention map and the feature map. This represents a feature map supplemented and guided by a bypass branch of global semantic information of the human body. The resolution is equal to that of f(X). Because the attention map's value range is [0,1], when the Hadamard product is calculated multiple times, the deep feature map will obtain very small values, resulting in a decrease in expressive power. Therefore, the enhanced feature map... Perform an element-wise addition operation with the feature map f(X) extracted from the corresponding main branch, and the output is... This will be input into the next main branch of the network and its corresponding bypass branch for global human semantic information, as shown in the following formula:
[0083]
[0084] For example, the first global semantic information completion branch can be understood as a background suppression branch, generating a spatial attention mask to reduce the weight of background regions, thus eliminating the background response of the target object features extracted by the first backbone network branch. The second global semantic information completion branch can be understood as a keypoint enhancement branch, dynamically enhancing the feature response of keypoint regions through heatmap prediction, thereby enhancing the target object keypoint regions of the target object features extracted by the second backbone network branch. The outputs of the two branches can be weighted and fused with the original feature map to generate a globally semantically enhanced feature map.
[0085] For example, for the first global semantic information completion branch, global average pooling can be performed on the input feature map to generate channel description vectors. Dependencies between channels are learned through fully connected layers, and channel attention weights are output. These channel attention weights are multiplied channel-by-channel with the original feature map to generate a preliminary attention map. Dilated convolutions are used to expand the receptive field, capturing long-range contextual relationships. The attention map is then normalized to a 0, 1 mask matrix using the sigmoid function. Finally, the mask matrix is multiplied point-by-point with the original feature map to suppress background weight regions.
[0086] For example, for the second global semantic information completion branch, a lightweight convolutional head (such as a 1×1 convolution + ReLU) can be superimposed on the feature map of the corresponding second backbone network branch to output a preliminary keypoint heatmap. A Gaussian kernel function is used to smooth the heatmap, generating a probability distribution map. The heatmap is multiplied channel-by-channel with the original feature map to amplify the feature responses of the keypoint regions, and residual connections are introduced to preserve the integrity of the original features. Finally, feature maps at different levels are upsampled and aligned, and then fused through channel concatenation. The SE (Squeeze-and-Excitation) module is used to adaptively adjust the channel weights.
[0087] By adopting this technical solution, the background response of the target object features extracted from the first backbone network branch can be eliminated by the first global semantic information completion branch; the key point region of the target object features extracted from the second backbone network branch can be enhanced by the second global semantic information completion branch, thereby generating a feature map with global semantic enhancement, which can then accurately distinguish the key points of adjacent target objects and adjacent pedestrians.
[0088] In one embodiment, the target object prediction result includes target object density prediction information, based on which S3 includes:
[0089] S31. Based on the target object density prediction information and abnormal behavior prediction results at the first preset time, match the associated first urban management event and the first urban management emergency level corresponding to the first urban management event through the urban management matching model.
[0090] S32, Output the first urban management incident and the first urban management emergency level.
[0091] Specifically, the target object density prediction information includes pedestrian density data and the density level of the target object, which can be specific density data or a density level representation. The density level includes high-density, medium-density, or low-density states. Abnormal behavior prediction results can include detected target object behaviors, detected abnormal behaviors of the target object, and the degree of matching between the detected target object behaviors and one or more abnormal behaviors. Abnormal behaviors can include watching, squatting, abnormal standing, knocking, setting up stalls, dumping, gathering, stampeding, fighting, etc. The first urban management event and its corresponding emergency level can be matched using an urban management matching model. For example, if the target object density prediction information describes a high density level, and the abnormal behavior prediction results describe abnormal target object behaviors including watching or gathering, performing similarity matching based on the urban management matching model might result in multiple people watching or abnormal gatherings. Further confirmation can then be made regarding whether the gathering is caused by a traffic accident, a stall-based crowd, street vending, or illegal gatherings, etc. Based on the pre-trained urban management matching model, and using the target object density prediction information and abnormal behavior prediction results at a first preset time, the model outputs similarity analyses for different scenarios and comprehensively provides the corresponding first urban management emergency level, outputting the first urban management event and its corresponding emergency level for timely handling by relevant personnel. By employing this technical solution, the model can accurately output the target object density prediction information and abnormal behavior prediction results at the first preset time based on a pose model that combines global semantic information completion and attention collaboration mechanisms. Further analysis yields the matched first urban management event and its corresponding emergency level, allowing relevant personnel to quickly and accurately understand the urban management event and its corresponding emergency situation at the first preset time, facilitating efficient subsequent processing.
[0092] In one embodiment, the method further includes:
[0093] By using a pre-trained pose detection model that combines global semantic information completion and attention coordination mechanism, the city management prediction results at the first and second preset times are output.
[0094] Based on the target object prediction results and abnormal behavior prediction results at the first and second preset times, the target object trend information and behavior trend information are output through the urban management matching model.
[0095] Specifically, the first preset time can be the current time, and the second preset time can be the next time after the current time; the first preset time can also be a historical time, and the second preset time can be the current time; or the first preset time can be a historical time, and the second preset time can be a future time. The selection and analysis of the time can be tailored to different scenarios, and are not limited here. Based on the urban management prediction results at different times, such as the first and second preset times, these can be used as temporal features. Through an urban management matching model with deep learning on temporal features, target object trend information and behavioral trend information can be obtained. Specifically, the target object trend information can describe the population flow trend between the first and second preset times, or after the second preset time; the behavioral trend information can describe the behavioral trend of the target object between the first and second preset times, or after the second preset time. By adopting this technical solution, target object trend information and behavioral trend information can be output based on temporal features, enabling relevant personnel to understand the urban event trends in the current urban management scenario and facilitate the implementation of corresponding measures.
[0096] In one embodiment, the method further includes:
[0097] Based on the target object's trend information and behavioral trend information, the city management matching model is used to match the associated second urban management events and the second urban management emergency level corresponding to the second urban management events.
[0098] Output the second urban management incident and its emergency level.
[0099] Specifically, the urban management matching model can also be pre-trained for different types of crowd trends and behavioral trends. For example, target object trend information can include crowd flow trends and the direction and location of those flows. Crowd flow trends can include denser or sparser crowd flow; the direction and location can include crowds moving from one location to another. Behavioral trend information can include changes in crowd behavior from static to stampede, from static to fighting, or from stampede back to static. For example, if the target object trend information is denser crowd flow and the behavioral trend is a change from static to stampede, the matched second urban management incident might be a stampede, and the second urban management emergency level might be red. If the target object trend information is sparser crowd flow and the behavioral trend is a change from static to fighting, the second urban management emergency level might be yellow. If the target object trend information is sparser crowd flow and the behavioral trend is a change from stampede to static, the second urban management emergency level might be green. It should be noted that the above is merely an illustrative example and is not intended to limit the actual situation. By adopting this technical solution, it is possible to update and learn based on different population trends and behavioral trends, and match them to obtain the second urban management incident and the second urban management emergency level. This allows relevant personnel to obtain the comprehensive population and behavioral trends of urban management within a certain time range based on the second urban management incident and the second urban management emergency level, so as to take corresponding measures for the corresponding trend events and levels.
[0100] In this embodiment, a pre-trained pose detection model that combines global semantic information completion and attention coordination mechanism can distinguish key points of adjacent target objects and adjacent pedestrians, improve the localization accuracy of key points of target objects or target pedestrians with small targets, and thus improve the prediction accuracy of target object prediction results and abnormal behavior prediction results, thereby accurately predicting urban management events.
[0101] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0102] Example 2
[0103] Based on the same inventive concept, this application also provides a smart city management event prediction device for implementing the smart city management event prediction method described above. The solution provided by this device is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more smart city management event prediction device embodiments provided below can be found in the limitations of the smart city management event prediction method described above, and will not be repeated here.
[0104] In one embodiment, such as Figure 6 As shown, a smart city management event prediction device is provided, comprising: a data acquisition module for acquiring real-time image data of a target city management monitoring area; a multi-person human posture detection module for outputting a city management prediction result at a first preset time based on the real-time image data using a pre-trained posture detection model combining global semantic information completion and attention collaboration mechanism, wherein the city management prediction result includes a target object prediction result and an abnormal behavior prediction result; and a city management matching prediction module for outputting a city management event prediction result based on the target object prediction result and the abnormal behavior prediction result at the first preset time using a city management matching model.
[0105] Furthermore, the pre-trained pose detection model combining global semantic information completion and attention collaboration mechanism includes a target detection module, multiple global semantic information completion modules, an attention collaboration mechanism module, and an output module. The target detection module detects multiple target object image features based on the real-time image data. The global semantic information completion module generates a corresponding target object attention feature map based on each target object image feature. The attention collaboration mechanism module obtains a target object key point heatmap for a specified region based on different target object attention feature maps from different global semantic information completion modules. The output module obtains the target object pose result based on the target object key point heatmap and generates a target object prediction result and an abnormal behavior prediction result based on the target object pose result.
[0106] Furthermore, the global semantic information completion module includes a first backbone network branch and a first global semantic information completion branch corresponding to the first backbone network branch, or a second backbone network branch and a second global semantic information completion branch corresponding to the second backbone network branch; the first global semantic information completion branch eliminates the background response of the target object features extracted by the first backbone network branch; the second global semantic information completion branch enhances the key point region of the target object corresponding to the target object features extracted by the second backbone network branch.
[0107] Furthermore, the target object prediction result includes target object density prediction information. Based on this, the urban management matching prediction module is also used to match the associated first urban management event and the first urban management emergency level corresponding to the first urban management event through the urban management matching model according to the target object density prediction information and the abnormal behavior prediction result at the first preset time. It is also used to output the first urban management event and the first urban management emergency level.
[0108] Furthermore, the urban management matching prediction module is also used to output urban management prediction results at the first preset time and the second preset time by using a pre-trained pose detection model that combines global semantic information completion and attention coordination mechanism; and to output target object trend information and behavior trend information by using the urban management matching model based on the target object prediction results and abnormal behavior prediction results at the first preset time and the second preset time.
[0109] Furthermore, the urban management matching prediction module is also used to match the associated second urban management event and the second urban management emergency level corresponding to the second urban management event through the urban management matching model based on the target object trend information and the behavioral trend information; and to output the second urban management event and the second urban management emergency level.
[0110] The modules in the aforementioned smart city management event prediction device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0111] Example 3
[0112] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0113] Acquire real-time image data of the target city's management and monitoring area;
[0114] Based on the real-time image data, a pre-trained pose detection model combining global semantic information completion and attention coordination mechanism is used to output the urban management prediction result at the first preset time. The urban management prediction result includes the target object prediction result and the abnormal behavior prediction result of the target object.
[0115] Based on the target object prediction results and the abnormal behavior prediction results at the first preset time, the urban management event prediction results are output through the urban management matching model.
[0116] In one embodiment, the pre-trained pose detection model combining global semantic information completion and attention coordination mechanism includes a target detection module, multiple global semantic information completion modules, an attention coordination mechanism module, and an output module.
[0117] The target detection module detects multiple target object image features based on the real-time image data;
[0118] The global semantic information completion module generates a corresponding target object attention feature map based on the image features of each target object;
[0119] The attention collaboration mechanism module obtains a heatmap of key points of the target object in a specified region by completing different target object attention feature maps of different global semantic information completion modules.
[0120] The output module obtains the target object's pose result based on the target object's key point heatmap, and generates the target object prediction result and abnormal behavior prediction result based on the target object's pose result.
[0121] In one embodiment, the global semantic information completion module includes a first backbone network branch and a first global semantic information completion branch corresponding to the first backbone network branch, or a second backbone network branch and a second global semantic information completion branch corresponding to the second backbone network branch.
[0122] The first global semantic information completion branch eliminates the background response of the target object features extracted from the first backbone network branch;
[0123] The second global semantic information completion branch enhances the target object key point region corresponding to the target object features extracted by the second backbone network branch.
[0124] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0125] Based on the target object density prediction information and the abnormal behavior prediction results at the first preset time, the first urban management event and the first urban management emergency level corresponding to the first urban management event are matched through the urban management matching model.
[0126] Output the first urban management incident and the first urban management emergency level.
[0127] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0128] By using a pre-trained pose detection model that combines global semantic information completion and attention coordination mechanism, the city management prediction results at the first and second preset times are output.
[0129] Based on the target object prediction results and abnormal behavior prediction results at the first preset time and the second preset time, the target object trend information and behavior trend information are output through the urban management matching model.
[0130] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0131] Based on the target object trend information and the behavioral trend information, the city management matching model is used to match the associated second urban management event and the second urban management emergency level corresponding to the second urban management event.
[0132] Output the second urban management event and the second urban management emergency level.
[0133] Example 4
[0134] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0135] Acquire real-time image data of the target city's management and monitoring area;
[0136] Based on the real-time image data, a pre-trained pose detection model combining global semantic information completion and attention coordination mechanism is used to output the urban management prediction result at the first preset time. The urban management prediction result includes the target object prediction result and the abnormal behavior prediction result of the target object.
[0137] Based on the target object prediction results and the abnormal behavior prediction results at the first preset time, the urban management event prediction results are output through the urban management matching model.
[0138] In one embodiment, the pre-trained pose detection model combining global semantic information completion and attention coordination mechanism includes a target detection module, multiple global semantic information completion modules, an attention coordination mechanism module, and an output module.
[0139] The target detection module detects multiple target object image features based on the real-time image data;
[0140] The global semantic information completion module generates a corresponding target object attention feature map based on the image features of each target object;
[0141] The attention collaboration mechanism module obtains a heatmap of key points of the target object in a specified region by completing different target object attention feature maps of different global semantic information completion modules.
[0142] The output module obtains the target object's pose result based on the target object's key point heatmap, and generates the target object prediction result and abnormal behavior prediction result based on the target object's pose result.
[0143] In one embodiment, the global semantic information completion module includes a first backbone network branch and a first global semantic information completion branch corresponding to the first backbone network branch, or a second backbone network branch and a second global semantic information completion branch corresponding to the second backbone network branch.
[0144] The first global semantic information completion branch eliminates the background response of the target object features extracted from the first backbone network branch;
[0145] The second global semantic information completion branch enhances the target object key point region corresponding to the target object features extracted by the second backbone network branch.
[0146] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0147] Based on the target object density prediction information and the abnormal behavior prediction results at the first preset time, the first urban management event and the first urban management emergency level corresponding to the first urban management event are matched through the urban management matching model.
[0148] Output the first urban management incident and the first urban management emergency level.
[0149] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0150] By using a pre-trained pose detection model that combines global semantic information completion and attention coordination mechanism, the city management prediction results at the first and second preset times are output.
[0151] Based on the target object prediction results and abnormal behavior prediction results at the first preset time and the second preset time, the target object trend information and behavior trend information are output through the urban management matching model.
[0152] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0153] Based on the target object trend information and the behavioral trend information, the city management matching model is used to match the associated second urban management event and the second urban management emergency level corresponding to the second urban management event.
[0154] Output the second urban management event and the second urban management emergency level.
[0155] Example 5
[0156] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:
[0157] Acquire real-time image data of the target city's management and monitoring area;
[0158] Based on the real-time image data, a pre-trained pose detection model combining global semantic information completion and attention coordination mechanism is used to output the urban management prediction result at the first preset time. The urban management prediction result includes the target object prediction result and the abnormal behavior prediction result of the target object.
[0159] Based on the target object prediction results and the abnormal behavior prediction results at the first preset time, the urban management event prediction results are output through the urban management matching model.
[0160] In one embodiment, the pre-trained pose detection model combining global semantic information completion and attention coordination mechanism includes a target detection module, multiple global semantic information completion modules, an attention coordination mechanism module, and an output module.
[0161] The target detection module detects multiple target object image features based on the real-time image data;
[0162] The global semantic information completion module generates a corresponding target object attention feature map based on the image features of each target object;
[0163] The attention collaboration mechanism module obtains a heatmap of key points of the target object in a specified region by completing different target object attention feature maps of different global semantic information completion modules.
[0164] The output module obtains the target object's pose result based on the target object's key point heatmap, and generates the target object prediction result and abnormal behavior prediction result based on the target object's pose result.
[0165] In one embodiment, the global semantic information completion module includes a first backbone network branch and a first global semantic information completion branch corresponding to the first backbone network branch, or a second backbone network branch and a second global semantic information completion branch corresponding to the second backbone network branch.
[0166] The first global semantic information completion branch eliminates the background response of the target object features extracted from the first backbone network branch;
[0167] The second global semantic information completion branch enhances the target object key point region corresponding to the target object features extracted by the second backbone network branch.
[0168] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0169] Based on the target object density prediction information and the abnormal behavior prediction results at the first preset time, the first urban management event and the first urban management emergency level corresponding to the first urban management event are matched through the urban management matching model.
[0170] Output the first urban management incident and the first urban management emergency level.
[0171] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0172] By using a pre-trained pose detection model that combines global semantic information completion and attention coordination mechanism, the city management prediction results at the first and second preset times are output.
[0173] Based on the target object prediction results and abnormal behavior prediction results at the first preset time and the second preset time, the target object trend information and behavior trend information are output through the urban management matching model.
[0174] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0175] Based on the target object trend information and the behavioral trend information, the city management matching model is used to match the associated second urban management event and the second urban management emergency level corresponding to the second urban management event.
[0176] Output the second urban management event and the second urban management emergency level.
[0177] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0178] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0179] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0180] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for predicting events in smart city management, characterized in that, The method includes: Acquire real-time image data of the target city's management and monitoring area; Based on the real-time image data, a pre-trained pose detection model combining global semantic information completion and attention coordination mechanism is used to output the urban management prediction result at the first preset time. The urban management prediction result includes the target object prediction result and the abnormal behavior prediction result of the target object. Based on the target object prediction results and abnormal behavior prediction results at the first preset time, the urban management event prediction results are output through the urban management matching model. The pre-trained pose detection model that combines global semantic information completion and attention collaboration mechanism includes a target detection module, multiple global semantic information completion modules, an attention collaboration mechanism module, and an output module. The target detection module detects multiple target object image features based on the real-time image data; The global semantic information completion module generates a corresponding target object attention feature map based on the image features of each target object; The attention collaboration mechanism module obtains a heatmap of key points of the target object in a specified region by completing different target object attention feature maps of different global semantic information completion modules. The output module obtains the target object's pose result based on the target object's key point heatmap, and generates the target object prediction result and abnormal behavior prediction result based on the target object's pose result.
2. The method according to claim 1, characterized in that, The global semantic information completion module includes a first backbone network branch and a first global semantic information completion branch corresponding to the first backbone network branch, or a second backbone network branch and a second global semantic information completion branch corresponding to the second backbone network branch; The first global semantic information completion branch eliminates the background response of the target object features extracted from the first backbone network branch; The second global semantic information completion branch enhances the target object key point region corresponding to the target object features extracted by the second backbone network branch.
3. The method according to claim 1, characterized in that, The target object prediction result includes target object density prediction information. The step of outputting urban management event prediction results based on the target object prediction result and the abnormal behavior prediction result at a first preset time point, through an urban management matching model, includes: Based on the target object density prediction information and the abnormal behavior prediction results at the first preset time, the first urban management event and the first urban management emergency level corresponding to the first urban management event are matched through the urban management matching model. Output the first urban management incident and the first urban management emergency level.
4. The method according to claim 1, characterized in that, The method further includes: By using a pre-trained pose detection model that combines global semantic information completion and attention coordination mechanism, the city management prediction results at the first and second preset times are output. Based on the target object prediction results and abnormal behavior prediction results at the first preset time and the second preset time, the target object trend information and behavior trend information are output through the urban management matching model.
5. The method according to claim 4, characterized in that, The method further includes: Based on the target object trend information and the behavioral trend information, the city management matching model is used to match the associated second urban management event and the second urban management emergency level corresponding to the second urban management event. Output the second urban management event and the second urban management emergency level.
6. A smart city management event prediction device, characterized in that, The device includes: The data acquisition module is used to acquire real-time image data of the target city's management and monitoring area; The multi-person human posture detection module is used to output the urban management prediction result at the first preset time based on the real-time image data and through a pre-trained posture detection model that combines global semantic information completion and attention collaboration mechanism. The urban management prediction result includes the target object prediction result and the abnormal behavior prediction result. The urban management matching and prediction module is used to output urban management event prediction results based on the prediction results of the target object and the prediction results of the abnormal behavior at a first preset time, through the urban management matching model. The pre-trained pose detection model that combines global semantic information completion and attention collaboration mechanism includes a target detection module, multiple global semantic information completion modules, an attention collaboration mechanism module, and an output module. The target detection module detects multiple target object image features based on the real-time image data; The global semantic information completion module generates a corresponding target object attention feature map based on the image features of each target object; The attention collaboration mechanism module obtains a heatmap of key points of the target object in a specified region by completing different target object attention feature maps of different global semantic information completion modules. The output module obtains the target object's pose result based on the target object's key point heatmap, and generates the target object prediction result and abnormal behavior prediction result based on the target object's pose result.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Multi-target detection model and method for city appearance event management
CN117576569A
Three-dimensional semantic scene completion method and device, storage medium and computer equipment
CN119379562A