A cloud-edge collaboration-based shopping mall intelligent analysis system and method
Patent Information
- Application Number
- CN202611001167.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-25
AI Technical Summary
第一,解决边缘算力资源无法被云端感知和协同调度的问题,实现商场各分区边缘节点算力负载的统一采集、上报与全局协同优化
(1)独创算力数据双向闭环协同架构,打破边缘算力调度、跨镜行人匹配、模型迭代三大模块数据隔离壁垒,以分区算力负载作为统一耦合纽带,同步解决算力分配失衡、商铺遮挡轨迹断裂、密集行人识别精度不足耦合问题,取得单一优化手段无法实现的预料不到的综合技术效果。
Smart Images

Figure CN122820249A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of cloud-edge collaborative intelligent sensing and commercial big data analysis technology. Specifically, it relates to a cloud-edge collaborative intelligent analysis system and method for shopping malls, which is applicable to cross-camera pedestrian re-identification, adaptive scheduling of edge computing power, dynamic distillation and iteration of models, and quantitative analysis of commercial business value in complex shopping mall scenarios. Background Technology
[0002] Currently, large shopping malls generally deploy video surveillance systems to ensure safety and operational order. Existing technical solutions can be mainly divided into three categories: The first type is the traditional passive video surveillance system, which only has the function of recording and storing videos, relies on manual review, lacks real-time analysis capabilities, has a single data dimension, and cannot support refined operations.
[0003] The second type is a single-point analysis system based on edge intelligence, which uses smart cameras equipped with low-computing-power AI chips to realize pedestrian detection and tracking locally. However, each device operates independently and lacks a cross-camera collaboration mechanism, making it difficult to integrate with business systems such as POS and CRM.
[0004] The third type is a cloud-based centralized intelligent analysis architecture, which uploads the raw video streams from all front-end cameras to the cloud for unified processing. This solution can theoretically achieve global perception, but it has extremely high requirements for network bandwidth, cloud storage, and computing resources, and has large end-to-end latency, making it difficult to meet the needs of low-latency scenarios such as congestion warnings.
[0005] Existing technologies struggle to achieve an effective balance between continuous behavior perception across the entire shopping mall, system cost, real-time performance, and scalability. Furthermore, the lack of a unified computing power scheduling mechanism prevents the cloud from sensing and utilizing the computing resources of each edge node. Cross-camera tracking suffers from severe trajectory fragmentation in scenarios with dense store occlusion. The recognition accuracy of deep learning models in specific scenarios is difficult to continuously improve. Current intelligent analysis systems for shopping malls still employ fixed-weight feature fusion strategies and fixed-frequency model update mechanisms when edge computing power fluctuates. This results in a significant decrease in cross-camera tracking accuracy during high-load periods and wasted computing resources during low-load periods. Moreover, model iterations cannot adapt to the dynamic changes in densely occluded shopping mall scenarios.
[0006] The three existing technologies share a common inherent defect: edge computing power scheduling, cross-camera pedestrian matching, and model distillation iteration operate independently without a unified coupling link. They cannot simultaneously optimize the three coupling issues of computing power balance, occlusion recognition accuracy, and model iteration efficiency, making it difficult to meet the dual needs of real-time security early warning and refined commercial data analysis.
[0007] To address the aforementioned technical bottlenecks, there is an urgent need for a new intelligent visual analysis architecture that can achieve in-depth perception of customer behavior across the entire shopping mall without significantly increasing network load and cloud resource consumption, and support the coordinated scheduling of system computing resources and the continuous evolution of model capabilities. Summary of the Invention
[0008] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a cloud-edge collaborative intelligent analysis system and method for shopping malls. Specifically, this invention aims to solve the following technical problems: First, it addresses the issue that edge computing resources cannot be perceived and coordinated by the cloud, enabling unified collection, reporting, and global collaborative optimization of computing load at edge nodes in various shopping mall zones.
[0009] Second, it addresses the issue of frequent breaks in pedestrian trajectories across cameras in scenarios with dense shop occupancy by achieving stable matching through the fusion of appearance, movement timing, and three-dimensional spatial features of the shopping mall.
[0010] Third, to address the problem that deep learning models cannot continuously improve their accuracy as the scenario changes after deployment, we achieve closed-loop evolution of model capabilities through incremental training based on difficult examples and knowledge distillation techniques.
[0011] Fourth, address the issue of data fragmentation between security monitoring and commercial operation systems by building a comprehensive platform that integrates security early warning, customer flow analysis, and business decision-making through multi-source data fusion.
[0012] To achieve the above objectives, the present invention provides a cloud-edge collaborative intelligent analysis system and method for shopping malls.
[0013] The system provided by this invention includes an integrated edge-aware computing node, a cloud platform, and a model iteration module; Integrated edge-aware computing nodes are deployed in various customer flow zones of the shopping mall, configured with edge computing resources, and run a lightweight first deep learning model to complete pedestrian detection and feature extraction. They collect the computing load of each zone in real time and simultaneously upload structured metadata and load data to the cloud platform. The cloud platform is equipped with cloud computing resources, accesses shopping mall business data, and has a cross-camera pedestrian re-identification unit; The model iteration module is deployed on the cloud platform. It completes incremental training of the model based on the identification of difficult examples and generates a lightweight model through knowledge distillation, which is then distributed to each edge perception computing node. The difficult examples refer to pedestrian identification samples that are missed or falsely reported by the system. The knowledge distillation refers to the technique of transferring the knowledge of the complex teacher model to the simplified student model. The core innovation of this system is a two-way closed-loop collaborative mechanism for computing power and data. This mechanism is defined as follows: using the real-time computing power load of edge partitions as a unified coupling link, it forms a two-way data closed loop where uplink computing power data participates in cross-mirror feature weighting, and downlink computing power data regulates model distillation iteration. It includes two core execution logics: (1) The cross-camera pedestrian re-identification unit integrates the three-dimensional features of pedestrian appearance, movement sequence and mall spatial location, and uses the partition computing power load data uploaded from the edge as a weighting factor to correct the weight of mall spatial location features. (2) The model iteration module adopts a customized distillation loss function for the dense occlusion scenario of shopping malls, increases the loss weight of occluded samples, and dynamically switches the distillation iteration frequency according to the real-time computing power load of each partition.
[0014] The bidirectional closed-loop collaborative mechanism of computing power and data forms a two-way data flow where edge computing power participates upward in feature weighting and downward in adjusting model iteration, generating dual gains of computing power balance and tracking accuracy that cannot be achieved by single-module optimization. The system simultaneously outputs security early warning results and quantitative analysis results of store traffic value and marketing activity conversion effects.
[0015] Furthermore, the integrated edge-aware computing node is configured with a partitioned adaptive computing power scheduling mechanism, setting two types of computing power occupancy thresholds: high and low. When the computing power occupancy exceeds the high threshold, the inference frame rate is reduced; when the computing power occupancy is below the low threshold, the inference frame rate is increased.
[0016] Furthermore, this system also includes an application visualization and linkage module; the application visualization and linkage module realizes multi-terminal hierarchical data display, identifies abnormal events in the mall and pushes alarms in a hierarchical manner, and links broadcasting, access control, and fire emergency equipment.
[0017] Furthermore, the cloud platform has a built-in multi-source data fusion engine that links customer flow trajectories, crowd density, and store transaction, membership, and parking data to output quantitative analysis conclusions on store traffic value and marketing campaign effectiveness.
[0018] The method provided by this invention includes: Step 1: The edge nodes of each zone in the shopping mall run a lightweight first deep learning model to complete pedestrian detection and tracking, collect the computing load of each zone, and only upload structured metadata and computing load data to the cloud platform synchronously. Step 2: The cloud platform aggregates structured metadata and computing load data uploaded from the edge, and connects to the mall's operational data; Step 3: The cloud platform runs the second deep learning model, which integrates the three-dimensional features of pedestrian appearance, movement sequence, and mall spatial location to complete cross-camera pedestrian matching and outputs the results of abnormal customer flow and business analysis. Step 4: Collect difficult examples of scene recognition to incrementally train the cloud model, generate a lightweight model through knowledge distillation and distribute it to each edge node. The difficult examples refer to pedestrian recognition samples that are missed or falsely reported by the system. The knowledge distillation refers to the technique of transferring the knowledge of the complex teacher model to the simplified student model. The core of this method is to establish a two-way closed-loop collaborative logic for computing power data, defined as follows: relying on the partitioned computing power load reported from the edge to connect the entire process of feature matching and model optimization, which is divided into two main branches: uplink weighting and downlink control. (1) In the matching process of step three, the partition computing load data uploaded from the edge is introduced as a weighting factor to correct the fusion weight of the shopping mall spatial location features; (2) In the distillation process of step four, a customized distillation loss function for the dense occlusion scenario of the shopping mall is adopted to increase the loss weight of the occlusion sample, and the distillation iteration frequency is dynamically adjusted according to the real-time computing power load of each partition.
[0019] This two-way closed-loop collaborative logic integrates the entire process of computing power acquisition, cross-camera tracking, and model self-optimization, simultaneously improving multiple shortcomings such as computing power fluctuations, occlusion matching failures, and low accuracy in dense recognition. This method simultaneously outputs security early warning information and quantitative analysis data on store traffic value and marketing campaign conversion revenue.
[0020] Furthermore, the logic for adaptively adjusting the inference frame rate in step one is as follows: if the computing power occupancy in high-density areas exceeds the threshold, the frame rate is lowered; if the computing power occupancy in low-density areas such as shops and parking lots is relatively low, the inference frame rate is raised.
[0021] Furthermore, in step three, when fusing three-dimensional pedestrian features, the weighting coefficients of the mall's spatial location features are dynamically corrected using the real-time computing power load values of each zone.
[0022] Furthermore, during the knowledge distillation process in step four, the number of distillation iterations is switched according to the computing power load range of each partition, and higher loss weights are assigned to pedestrian overlapping and occlusion samples.
[0023] Furthermore, after completing the multi-source data fusion in step two, the data on customer flow, transactions, membership, and parking are linked to calculate the value of store traffic and the conversion revenue of marketing activities.
[0024] Furthermore, this method also includes a visualization and linkage step: hierarchical display of customer flow, operation, and early warning data, and push of alarms and linkage of on-site emergency equipment after identifying mall anomalies.
[0025] Compared with the prior art, the present invention has the following beneficial effects: (1) The original two-way closed-loop collaborative architecture of computing power and data breaks the data isolation barrier of the three major modules of edge computing power scheduling, cross-camera pedestrian matching and model iteration. It uses partitioned computing power load as a unified coupling link to simultaneously solve the coupling problems of unbalanced computing power distribution, broken trajectory due to shop obstruction and insufficient accuracy of dense pedestrian recognition, and achieves unexpected comprehensive technical effects that cannot be achieved by a single optimization method.
[0026] (2) A distillation loss function is customized for densely occluded scenarios in shopping malls, and combined with a dynamic adaptation and iteration mechanism of computing power, the stability of pedestrian tracking under occlusion is greatly improved.
[0027] (3) Dynamically balance the computing power across the entire region to avoid overload of computing power in high-traffic areas and idle resources in low-traffic areas, thereby improving the stability of system operation.
[0028] (4) Only structured metadata is uploaded at the edge, which greatly reduces network bandwidth consumption; multi-dimensional business data is integrated to output security early warning and business operation analysis indicators in one stop. Attached Figure Description
[0029] Figure 1 This is a schematic diagram of the overall architecture of the cloud-edge collaborative intelligent analysis system for shopping malls according to the present invention. The architecture includes: an integrated edge-sensing computing node 1, a cloud platform 2, a model iteration module 3, a cross-camera pedestrian re-identification unit 4, a multi-source data fusion engine 5, an application visualization linkage module 6, a two-way closed-loop collaborative mechanism for computing power and data 7, edge-side computing power resources 8, cloud-side computing power resources 9, a lightweight first deep learning model 10, a multi-terminal hierarchical data display 11, and a second deep learning model 12. The integrated edge-sensing computing node 1 is equipped with edge-side computing power resources 8 and a lightweight first deep learning model 10. The cloud platform 2 is configured with the cross-camera pedestrian re-identification unit 4, the model iteration module 3, the multi-source data fusion engine 5, the cloud-side computing power resources 9, and the second deep learning model 12. The application visualization linkage module 6 realizes the multi-terminal hierarchical data display 11.
[0030] Figure 2 This is a schematic diagram of the complete process of the cloud-edge collaborative intelligent analysis method for shopping malls based on the present invention. The overall process is divided into four core steps, among which the third and fourth steps embed bidirectional closed-loop collaborative logic of computing power and data, including two major branches: uplink feature weighting and downlink distillation control. The 'uplink feature weighting' branch corresponds to... Figure 1 The process of correcting the spatial location feature weights based on edge computing load data in the mid-span camera pedestrian re-identification unit 4; the 'downward distillation control' branch corresponds to Figure 1 The model iteration module 3 dynamically adjusts the distillation iteration frequency based on the partition computing power load; an optional visualization linkage step is also provided. Detailed Implementation
[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0032] Unified definition of terms: 1. Structured metadata: Lightweight features, coordinates, timestamps, temporary IDs, and other structured information output by edge nodes after completing pedestrian detection and tracking, excluding the original video images; 2. Computing load: Real-time computing load percentage of the edge AI acceleration unit, serving as a core parameter for dynamic weighting and distillation frequency adjustment; 3. Three-dimensional pedestrian features: a set of three types of features, including appearance features, movement sequence features, and spatial location features in the shopping mall; 4. Difficult sample: Pedestrian identification samples that are missed or falsely reported by the system, used for incremental training of the model in the cloud; 5. Occlusion-customized loss function: A special loss function that adds a weighted term for occluded samples to the standard distillation loss.
[0033] Example 1: System Deployment and Initialization Please combine Figure 1 This embodiment provides a cloud-edge collaborative intelligent analysis system for shopping malls, which is applied to a large shopping mall.
[0034] In this embodiment, the integrated edge sensing computing node 1 is an integrated device that combines an AI computing chip and a camera. A total of thirty of these nodes are deployed across various customer flow zones in the shopping mall—including the four entrances / exits, the atrium area, the passageways on each floor, and the underground parking lot. During deployment, the cloud platform 2 assigns a unique node identifier to each edge node and spatially maps and binds it to the mall's physical zones. Each node communicates with the cloud platform 2 via gigabit Ethernet.
[0035] Cloud platform 2 is deployed in the shopping mall's server room, equipped with high-performance GPU servers and large-capacity storage space, and establishes encrypted communication channels with each edge node via dedicated lines. Cloud platform 2 is equipped with cloud computing resources 9, including a high-performance GPU server cluster and large-capacity storage space. The computing power configuration of cloud computing resources 9 is higher than that of edge computing resources 8, and is used to run the second deep learning model 12, perform large-scale feature comparison operations of the cross-camera pedestrian re-identification unit 4, and perform incremental training and knowledge distillation tasks of the model iteration module 3.
[0036] When the system starts up, all edge nodes and cloud platform 2 synchronize their time with the same time server via the network time protocol to ensure that the timestamp deviation of each device is less than ten milliseconds, providing a unified time reference for subsequent cross-camera pedestrian matching.
[0037] After deployment, the system administrator inputs the physical layout information of each area of the shopping mall through the visualization interface of the application visualization linkage module 6 on the cloud platform 2. This includes the location of each entrance and exit, the boundary of the atrium area, the store number and spatial range of each shop, and the direction of the passageways. The cloud platform 2 encodes the above information into a shopping mall spatial location feature library, which serves as the comparison benchmark for spatial location features in cross-camera pedestrian matching. For example, a shop located on the east side of the third floor passageway is encoded as an area node; the connectivity between different areas, such as the passageway connecting the entrance / exit and the atrium, is also encoded as a spatial topology relationship.
[0038] In this embodiment, the spatial location features of the shopping mall specifically refer to the fusion encoding of the following three types of information: First, regional node encoding, which divides the physical space of the shopping mall into several regional nodes, each regional node corresponding to a unique spatial identifier code; second, spatial coordinate mapping, which maps each regional node to a unified two-dimensional plane coordinate system to form a spatial location coordinate vector; and third, topological connectivity, which records the physical connectivity paths between each regional node to constrain the spatiotemporal consistency when matching across cameras.
[0039] The spatial feature base weight W_base adopts a fixed value commonly used in the industry, and the weighting adjustment coefficient α is generally adapted within the range of 0.1~0.3; for small supermarkets, α is 0.1, for medium-sized shopping malls, α is 0.2, and for large complexes, α is 0.3. Technical personnel in the relevant field can fine-tune it according to the size of the mall. The encoding and weight configuration of the above spatial location features are uniformly completed by the cloud platform 2 during the system initialization phase and distributed to each edge node through configuration files to ensure the consistency of the global spatial benchmark.
[0040] Please refer to Figure 1 The computing power data bidirectional closed-loop collaborative mechanism 7 is used in this system. In this system, the integrated edge sensing computing node 1 and the cloud platform 2 achieve functional coupling through the computing power data bidirectional closed-loop collaborative mechanism 7. This mechanism is not an independent physical entity, but rather embodies the organic unity of the following two major operating logics: Uplink, the partition computing power load data uploaded by the edge node serves as a dynamic weighting factor for spatial location features, participating in the feature fusion matching of the cloud-based cross-camera pedestrian re-identification unit 4 (see Example 4); Downlink, the real-time computing power load of each partition serves as a control parameter for the distillation frequency of the model iteration module 3, driving the cloud-based model optimization and lightweight model deployment (see Example 5). This forms a closed-loop data flow of "computing power perception → feature weighting → model iteration → computing power re-perception".
[0041] Example 2: Edge Node Adaptive Inference and Frame Rate Adjustment This embodiment describes in detail the local processing flow of the integrated edge-aware computing node 1.
[0042] Each edge node acquires video streams in real time with high resolution and high frame rate specifications. The edge nodes have built-in hardware image signal processors that automatically perform targeted image enhancement—including wide dynamic range adjustment, contrast stretching, and local dehazing—for common scenarios in shopping mall environments such as backlighting, window glare, and shelf obstruction, providing high-quality input images for subsequent artificial intelligence inference.
[0043] Each edge node is equipped with edge computing resources 8, including artificial intelligence acceleration units such as NPUs, GPUs, or AI acceleration chips and their corresponding computing scheduling interfaces. The edge nodes run a lightweight first deep learning model 10 on the artificial intelligence acceleration units. This model is a lightweight neural network model distributed after knowledge distillation in the cloud, with fewer parameters and less computation than the second deep learning model 12 running on the cloud platform 2, making it suitable for real-time inference on the edge computing resources 8. The lightweight first deep learning model 10 performs pedestrian target detection on each preprocessed frame of image, outputting the bounding box coordinates and class confidence score for each target. Simultaneously, the edge nodes run a multi-target tracking algorithm, assigning a temporary identifier to each detected target and continuously tracking its trajectory within the edge node's field of view, recording its position sequence in different frames.
[0044] In this embodiment, the lightweight first deep learning model 10 uses a lightweight MobileNetV3-Small backbone network with approximately 2.5M parameters and 56M FLOPs, specifically designed for low-computing edge environments. The second deep learning model 12 uses a ResNet-101 backbone network with approximately 44.5M parameters and 7.8G FLOPs, deployed in the cloud as the teacher network. The parameter sizes of the two models differ by about an order of magnitude, enabling the teacher network to output sufficiently rich soft labels during the distillation process to guide the student network convergence, while ensuring that the edge inference latency meets real-time requirements.
[0045] For each detected pedestrian target, the edge nodes extract a pedestrian feature vector from its bounding box. This feature vector is output by the pedestrian re-identification network in the tracking algorithm and is used for appearance feature comparison during subsequent cross-camera matching on the cloud platform.
[0046] It should be noted that this feature vector undergoes irreversible random projection dimensionality reduction—specifically, the edge nodes multiply the original feature vector by a randomly generated orthogonal matrix, and then perform a nonlinear mapping. The output remains the same dimension but cannot be reversed to restore the feature vector of the original image. Actual testing shows that this irreversible processing has less than a 3% impact on the recall rate of cross-camera pedestrian matching. This processing ensures data privacy and security during transmission while not significantly affecting the accuracy of cross-camera matching.
[0047] Each edge node monitors the utilization rate of chip computing resources in real time. Specifically, the edge computing resources 8 sample the current computing power utilization percentage at fixed intervals by reading the hardware performance counter of the artificial intelligence acceleration unit.
[0048] The partitioned adaptive computing power scheduling mechanism is configured as follows: When the computing power utilization rate of an edge node exceeds a high threshold, the node automatically reduces the inference frame rate while maintaining the recall rate of target detection at a preset minimum recall rate threshold. In this embodiment, the minimum recall rate threshold is set to 85%, a value based on a common evaluation standard in the field of target detection. This is achieved by switching the model from full-precision inference mode to low-bit quantization inference mode, significantly reducing the computational load while keeping the recall rate decrease within an acceptable range.
[0049] When the computing power utilization rate is below the low threshold, the node will restore the inference frame rate to the initial level and switch the model back to full precision mode.
[0050] Taking the high-density zone at the entrance and exit as an example, during the peak passenger flow period on weekend afternoons, the computing power occupancy rate of the edge nodes in this zone exceeds the high threshold. The system automatically reduces the inference frame rate, switches the model to quantization mode, reduces the computing power occupancy rate, and the system runs stably. During the off-peak period at night, the computing power occupancy rate is lower than the low threshold, the inference frame rate returns to the normal level, and the model returns to full-precision mode.
[0051] The low-density parking lot zones consistently maintain a low inference frame rate, keeping the computing power utilization rate low.
[0052] Edge nodes generate structured metadata for each frame of analysis results, with each metadata entry being extremely small. The structured metadata includes fields such as camera identifier, timestamp, target type, temporary identity identifier, bounding box coordinates, pedestrian feature vector, and event label.
[0053] The edge node uploads the structured metadata to cloud platform 2 in real time via an encrypted channel, excluding the original video frames. The original video stream is stored in the local storage unit of the edge node in a loop for a preset number of days. Data exceeding the time limit is automatically overwritten, and the video of the corresponding time period is selectively archived to cloud platform 2 only in the event of an abnormal event.
[0054] Example 3: Cloud-based Multi-Source Data Fusion This embodiment describes the multi-source data fusion processing flow of cloud platform 2.
[0055] Cloud Platform 2 Multi-Source Data Fusion Engine 5 accesses various types of business data from the shopping mall through standard application programming interfaces: The first category is sales system data, which includes transaction time, transaction amount, product category, number of transactions, etc. of each store, and is synchronized in batches with a fixed time granularity; The second category is membership system data, which includes basic member information and consumption behavior data. The third category is parking management system data, which includes vehicle entry time, exit time, parking duration, corresponding floor, etc.
[0056] The aforementioned data, along with the structured metadata uploaded by edge nodes, is aggregated and stored in the data fusion engine 5 of cloud platform 2 using a unified timestamp and spatial location as indexes. The data fusion engine 5 matches video visitor flow data and sales transaction data from the same spatial region based on time windows, forming a spatiotemporally aligned dataset of visitor flow and transaction volume.
[0057] In this embodiment, association refers to aligning and matching the passenger flow characteristic data obtained from video analysis with business operation data across two dimensions: time and space, to form a multi-dimensional composite dataset of passenger flow and business data. The specific method is as follows: In terms of time, video visitor data and sales transaction data are aligned using the same time window. For example, the video visitor count of a store within a specific time period is matched with the number of sales transactions of that store within the same time period.
[0058] In the spatial dimension, the shopping area where the pedestrian trajectory is located in the video analysis is matched with the sales records of the corresponding shop number to achieve a three-dimensional association between trajectory, shop and transaction.
[0059] Example 4: Cloud-based 3D Feature Fusion and Cross-Camera Pedestrian Re-identification Unit This embodiment describes in detail the processing flow of the cross-camera pedestrian re-identification unit 4 of the cloud platform 2.
[0060] After receiving structured metadata from all edge nodes, cloud platform 2 runs the second deep learning model 12 on the graphics processor to extract 3D features for each pedestrian target: The first dimension is appearance features, which are global feature vectors extracted based on deep convolutional neural networks, representing visual appearance information such as the pedestrian's clothing color, body shape, and items carried. Specifically, the pedestrian bounding box coordinates uploaded by the edge nodes are mapped back to the original video frames, and the corresponding frames are decoded and extracted on cloud platform 2.
[0061] The second dimension is the motion temporal feature, which is based on the time and position sequence of pedestrian temporary markers within the field of view of each edge node. The motion direction and motion speed are calculated to form a temporal feature vector.
[0062] The third dimension is the spatial location feature of the shopping mall, which is based on the spatial location information encoded by the mall's physical layout. Specifically, the encoding method involves dividing the mall's floor plan into multiple area nodes, each represented by a one-hot vector. On this basis, edge-up partitioned computing load data serves as a dynamic weighting factor to adjust the weights of the spatial location features.
[0063] One of the core innovations of this embodiment is that it uses the partition computing load data uploaded from the edge as a weighting factor for spatial location features.
[0064] Specifically, when calculating the spatial similarity between two targets, the cross-camera pedestrian re-identification unit 4 no longer uses fixed weight coefficients, but adopts a dynamic weighting method: the final weight of the spatial location features is determined by the base weight plus the contribution value of the computing load. Formula (1): W_spatial=W_base×(1+α×L), where W_base is the base weight, L is the real-time computing load value of the partition where the target is located, and α is the adjustment coefficient. The increased weight of the spatial location features in the high-load partition makes the matching algorithm more dependent on spatial location consistency in densely occluded scenarios; the change in the weight of the spatial location features in the low-load partition is smaller.
[0065] The total similarity of the three-dimensional features is calculated by the weighted sum of the three types of features mentioned above, formula (2): Similarity=β_A×A_sim+β_M×M_sim+β_S×S_sim, where β_A, β_M, and β_S are the weight coefficients of appearance, motion time sequence, and spatial location features, respectively, and the weight coefficients of each feature can be adjusted according to the actual scene; A_sim is the appearance similarity, M_sim is the motion time sequence similarity, and S_sim is the spatial location similarity.
[0066] When the same target leaves the field of view of edge node A and enters the field of view of edge node B, cloud platform 2 compares the last appearance record of the target in node A with the 3D features of all newly appearing targets in node B. If the combination with the highest similarity exceeds a preset matching threshold, it is determined to be the same target, and its temporary identifier in node A is fused with the temporary identifier in node B to generate the target's global identity and cross-camera global trajectory.
[0067] For example, a pedestrian wearing a red shirt walks from the atrium to the shop area. The spatial location feature weight of this target in the atrium camera increases due to high load. The cloud platform 2 calculates the similarity between its spatial location feature and the spatial location feature of the newly appearing target in the shop area camera. The spatial location feature has an effective contribution to the total similarity - when the final total similarity exceeds the matching threshold, the system determines that it is the same pedestrian, assigns a global identity identifier, and forms a complete trajectory sequence.
[0068] Example 5: Model Collaborative Iteration and Knowledge Distillation Please combine Figure 1 Model iteration module 3 and Figure 2 The method flow shown in this embodiment is used to understand this embodiment.
[0069] This embodiment describes in detail the optimization and deployment process of model iteration module 3.
[0070] During system operation, the application's visual linkage module 6 will display cross-camera matching results and abnormal event alarms. Operations personnel can manually confirm alarm events through the management backend. When an alarm is a real event, it is marked as valid; when an alarm is a false alarm, it is marked as a false alarm; when the system misses a real event, operators can manually create an event and mark it as a missed event.
[0071] False positives and false negatives are collectively referred to as hard cases, and are automatically saved to the hard case database on cloud platform 2. In this embodiment, hard cases are accumulated at a rate of several hundred per week.
[0072] It should be noted that these difficult sample data are all from video surveillance data in public places in shopping malls and are only used for the technical purpose of model training; the cloud platform 2 uses the same desensitization processing as mentioned above for personal information that may be contained in the difficult sample data to ensure that the training process does not involve the identification of specific natural persons.
[0073] Whenever the number of difficult examples in the sample library exceeds a preset limit or the system reaches a preset cycle, the model iteration module 3 of the cloud platform 2 automatically triggers the incremental training process: First, the parameters of the current second deep learning model 12 are used as the initial weights; second, the difficult examples are mixed with historical training data in a preset ratio to form an incremental training set; finally, fine-tuning training is performed on a graphics processing unit cluster for a limited number of rounds, with the learning rate set to a small percentage of the initial learning rate.
[0074] After incremental training is completed, the model iteration module 3 performs knowledge distillation to compress the optimized large model into a lightweight model. The lightweight model has the same structure as the first deep learning model 10 of the edge nodes.
[0075] The customized distillation loss function used in the distillation process consists of three parts: the first part is the standard cross-entropy loss, which is the difference between the student network and the real label; the second part is the divergence loss, which is the difference between the student network output and the teacher network soft label; and the third part is the occlusion customized loss term. The specific form of the occlusion customized loss term is: assigning additional weights to targets identified as occluded samples. This invention provides multiple sets of selectable occlusion weights: the weight for regular samples is 1, and the weights for occluded samples can be 2, 3, or 4; this embodiment selects a 3x weight as the optimal solution.
[0076] In this embodiment, the occlusion criteria are as follows: when the intersection-union ratio of the target bounding box and another target bounding box is greater than 0.5, it is determined to be an overlapping occlusion; or when the area within the target bounding box with a detection confidence level lower than the threshold (i.e., less than 0.3) exceeds 30% of the total area of the bounding box, it is determined to be a partial occlusion. If either condition is met, it is determined to be an occluded sample, and the loss weight is three times that of a regular sample.
[0077] This loss function enables the student network to produce a stronger response to occluded areas during training, thereby maintaining high feature extraction accuracy in densely populated scenes.
[0078] The iteration frequency of knowledge distillation is not a fixed value, but rather dynamically switches based on the real-time computing load of each partition. Figure 2 The specific implementation of the "downward distillation control" branch in the method flow shown is as follows: When the computing load of a certain partition is higher than the high threshold for a continuous period of time, the partition is determined to be a high-activity partition; when the contribution of difficult sample samples in a high-activity partition exceeds the preset proportion of the total sample size, the cloud distillation iteration frequency switches from the normal mode to the high-frequency mode; when the computing load of all partitions is lower than the low threshold for a long period of time, the distillation frequency returns to the normal mode.
[0079] In this embodiment, the load in the atrium area remained above the high threshold on weekend afternoons, contributing more than half of the difficult case samples this week. The cloud distillation iteration frequency was automatically switched to high-frequency mode, and multiple distillations were performed on weekend evenings. The generated optimized model was downloaded to the edge nodes of the atrium and surrounding areas via over-the-air download in the early morning of the next day.
[0080] The optimized lightweight model is distributed to each edge node via over-the-air download. The distribution process is as follows: The cloud divides the new model file into multiple data packets and transmits the data packets to each edge node in sequence through an encrypted channel. Each edge node verifies the received complete model file. After the verification is successful, the edge node backs up the currently running model to the old version and loads the new model. After the new model runs for a preset time without any abnormalities, the old version backup is deleted.
[0081] The entire update process is completed automatically in batches by region during off-peak hours at night, without affecting normal daytime operations.
[0082] Example 6: Visualized Linkage and Alarm Handling Please refer to Figure 1 The application visualization linkage module 6 is described in this embodiment, which describes the function and linkage control process of the application visualization linkage module 6.
[0083] The analysis results from Cloud Platform 2 are pushed to users with different roles through multi-terminal hierarchical data display 11: The command center's large screen displays a comprehensive heat map of customer flow across the mall, real-time cross-camera tracking data, a list of abnormal events, and key operational metrics. The security supervisor's terminal displays real-time alarm pop-ups and emergency response suggestions. The operations manager's terminal displays multi-dimensional data analysis reports, including time-based customer flow trends, regional customer flow distribution, comparisons of store traffic and conversion rates, and evaluations of marketing campaign effectiveness. These three terminals display different levels of data content based on user roles and permissions, achieving tiered data display.
[0084] The system provides tiered alerts based on event type and urgency: Emergency alarms are applicable to events that endanger personal safety, such as fall detection, stampede risk, fire smoke, etc. They are pushed to the security supervisor's terminal and the command center's large screen, and at the same time, relevant personnel are notified via SMS. They are also linked to the broadcast system to automatically broadcast evacuation instructions, the access control system to automatically open all exits, and the fire protection system to automatically start. Important alarms are applicable to events that affect operational order, such as crowding warnings, intrusion alarms, and prolonged loitering, and are pushed to the security supervisor's terminal and the command center's large screen; The alert is applicable to events that require attention, such as running at high speed, going against traffic, or abnormal number of people in an area, and is only pushed to the security supervisor's terminal.
[0085] In this embodiment, when the crowd density in the atrium area exceeds the safety threshold, the system automatically triggers an important alarm. A warning pop-up window appears on the command center's large screen, displaying a real-time heat map. After receiving the notification, security personnel go to the scene to manage the crowd. After a preset time, the system checks again to confirm that the density has dropped back to a safe range and automatically lifts the warning.
[0086] Multi-terminal hierarchical data display 11 is implemented through a web visualization framework and a mobile application. The backend retrieves data of corresponding granularity from the database according to different user roles, and the frontend renders and displays the data according to a preset template.
[0087] Example 7: Complete Operation Process Please combine Figure 1 Overall architecture and Figure 2 This embodiment provides a methodological process for understanding.
[0088] This embodiment uses a complete operational scenario of a shopping mall on a weekend afternoon as an example to illustrate the collaborative workflow of each module of the system.
[0089] The scenario is set on a Saturday afternoon, with a promotional event being held in the mall's atrium, resulting in a significant increase in customer traffic and a continuous rise in the computing load of the edge nodes in the atrium area. At the same time, a customer wearing a white shirt enters the mall from the entrance, passes through the atrium, completes a purchase at a shop, and then leaves.
[0090] At 2:00 PM, the computing load of the edge nodes in the central area rose above the high threshold. The adaptive scheduling mechanism reduced the inference frame rate and switched the model to quantization mode.
[0091] From 2:00 PM to 2:05 PM, customers enter the mall through the entrance / exit. Temporary identity identifiers are assigned to them at the edge nodes of the entrance / exit, pedestrian feature vectors are extracted, and structured metadata is generated and uploaded to cloud platform 2, along with the current computing load data of that partition.
[0092] From 2:05 PM to 2:08 PM, customers walk along the passageway towards the atrium. The edge node of the atrium detects the target in its field of view, assigns another temporary identity, extracts the feature vector, and uploads metadata and payload data.
[0093] At 2:08 PM, the cloud platform 2's cross-camera pedestrian re-identification unit 4 received data from the entrance / exit edge nodes and the atrium edge nodes, performing 3D feature fusion matching. The appearance similarity was high, the movement temporal similarity was also high, and the spatial location similarity received a high confidence level due to the embedding of high-load data from the atrium area as a weighting factor. The total similarity exceeded the matching threshold, and the system determined that they were the same target, generating a global identity identifier, and recording the trajectory as from the entrance / exit through the passage to the atrium.
[0094] From 2:08 PM to 2:15 PM, customers participated in promotional activities in the atrium, staying for approximately seven minutes. Edge nodes in the atrium area continuously tracked their location sequence, generating atrium stay event tags and uploading them to the cloud platform. During this period, the computing power load of the edge nodes in the atrium area remained consistently above a high threshold, while the frame rate remained at a low level.
[0095] From 2:15 PM to 2:20 PM, customers leave the atrium and enter the shops. The cameras at the shop entrances assign them new temporary identification tags, extract feature vectors, and upload data and payload data. The cloud platform 2 cross-camera pedestrian re-identification unit 4 matches them again, confirms the identity is consistent, and adds the shop node to the global trajectory update.
[0096] The customer then completes the transaction, and the sales system records the transaction data. This data is then uploaded to cloud platform 2 via the application programming interface (API). Cloud platform 2's multi-source data fusion engine 5 correlates the customer's trajectory data with the store's sales transaction data in time and space, generating analysis records and marking the customer as having converted.
[0097] As the customer leaves the store and heads towards the parking lot, each edge node sequentially completes tracking and uploads metadata, while cloud platform 2 continuously updates their global trajectory. Upon reaching the parking lot, the parking lot edge nodes upload data, cloud platform 2 confirms their departure, and the global trajectory is marked as complete.
[0098] At 2:30 PM, Cloud Platform 2's data fusion engine 5 generated the following analysis results based on all the above data: The heat map of customer flow shows that the central area is currently the area with the highest density, and the system automatically pushes a congestion warning to the security supervisor's terminal; the store traffic value report is updated, and the system automatically marks stores with high traffic and high conversion rate; the evaluation of the promotional activity shows that the total customer flow, number of related transactions and overall conversion rate in the central area have all increased compared to the same period last weekend, and the system generates an effective evaluation conclusion of the activity and pushes it to the operations manager's terminal.
[0099] At 3:00 PM, the computing load of the edge nodes in the central area dropped below the low threshold, the frame rate returned to normal, the model switched back to full precision mode, and the system automatically resumed full performance operation during off-peak hours.
[0100] That evening, operations staff confirmed that all alarms from the day were valid, and the relevant difficult example samples were fed back to the cloud training platform. Based on the load statistics of each partition that day, the cloud determined that the atrium was a high-activity partition and that the contribution of difficult examples exceeded the preset proportion. The distillation iteration frequency was switched to high-frequency mode, and distillation was performed that night. The optimized lightweight model was downloaded to the atrium and surrounding edge nodes via over-the-air download in the early morning of the following day, replacing the original model.
[0101] It should be noted that although the above embodiments describe shopping malls as a typical application scenario, the cloud-edge collaborative intelligent analysis system and method for shopping malls described in this invention are also applicable to other large indoor public places with high-density pedestrian traffic, multiple camera coverage, and frequent occlusion, including but not limited to airport terminals, high-speed rail station waiting halls, subway transfer hubs, large convention centers, and hospital outpatient halls. In the above scenarios, only the initial weights of spatial location features need to be adjusted according to the actual camera layout, and the computing power load threshold needs to be recalibrated according to the venue's opening hours, to reuse the core technical solution of this invention without structural modifications to the two-way closed-loop collaborative mechanism itself.
[0102] Performance test data table Test metrics Traditional cloud-based centralized solutions This invention's cloud-edge collaboration solution Optimization range Single node bandwidth usage 45Mbps 3Mbps 93% decrease End-to-end identification delay 2~5s ≤200ms Significantly reduced latency Cross-camera matching accuracy in occluded scenes 68.50% 84.30% An increase of 15.8% Dense occlusion detection recall rate 79.80% 89.10% An increase of 9.3%. Based on actual operational data, this invention offers the following technical advantages compared to traditional centralized cloud solutions: network bandwidth usage is reduced from 45 megabits per second per node to 3 megabits per second, a decrease of approximately 93%; end-to-end response latency is controlled within 200 milliseconds, superior to the 2-5 second latency of traditional cloud solutions; cross-camera pedestrian matching accuracy is improved from approximately 68.5% to 84.3% in shop occlusion scenarios; and after six rounds of distillation iterations, the recall rate for target detection in densely occluded scenarios is improved from approximately 79.8% to 89.1%.
[0103] The experimental data above show that the present invention uses partitioned computing power load data as a common link to achieve functional coupling of three links: edge computing power adaptive scheduling, cloud-based 3D feature weighted matching, and load-driven distillation iteration, thereby realizing the simultaneous improvement of computing power resource utilization, cross-camera matching accuracy, and model recognition capability.
[0104] Please combine Figure 2 This embodiment also implements a cloud-edge collaborative intelligent analysis method for shopping malls, and the overall operation steps are as follows: This method is divided into four core steps, a two-way closed-loop collaborative logic of computing power and data embedded in the core steps, and optional visualization linkage steps.
[0105] The first step is to execute the edge adaptive inference step. The edge nodes corresponding to each zone of the mall run the lightweight first deep learning model 10 to continuously complete pedestrian detection and tracking processing, and simultaneously collect the computing power load status of the zone. Only the structured metadata and computing power load data are uploaded to the cloud platform 2, and the model inference frame rate is adaptively adjusted according to the high and low thresholds of the computing power occupancy of the zone. The second step is to perform multi-source data fusion in the cloud. Cloud platform 2 aggregates the structured metadata and partitioned computing load data uploaded by all edge nodes, and at the same time accesses various business-related data of the shopping mall. The third step involves cloud-based 3D feature pedestrian association. Cloud platform 2 runs the second deep learning model 12, fusing three types of features: pedestrian appearance, movement time sequence, and mall spatial location, to perform cross-camera pedestrian matching calculations. During the feature fusion stage, edge-up computing load data is incorporated to dynamically adjust the fusion weights corresponding to the mall spatial location features. This step is internally embedded with… Figure 2 The "Upstream Feature Weighting" sub-logic shown outputs passenger flow anomaly identification results and business analysis results after the operation is completed. The fourth step involves performing a collaborative iterative optimization of the model. This includes collecting difficult-to-identify scene examples for incremental training of the cloud-based model, employing a customized distillation loss function for occluded scenes to increase the loss weight of occluded samples, and dynamically adjusting the distillation iteration frequency based on the real-time computing load of each partition. This step is embedded within... Figure 2The "downward distillation control" sub-logic shown completes the distillation process to obtain a lightweight model, which is then distributed to all edge nodes for updating.
[0106] This method can add an additional visualization linkage step to display customer flow, operation, and early warning related data across multiple terminals in a hierarchical manner. After identifying abnormal events in the shopping mall, it can push alarm information and link various emergency equipment on site. After the multi-source data fusion is completed, the value of shop traffic and the conversion revenue of marketing activities can be calculated based on the results of correlation calculation.
[0107] In this invention, edge computing load data is used as a dynamic weighting factor for spatial location features in the uplink direction for pedestrian matching in the cloud, and as a trigger condition to regulate the frequency of cloud model distillation iterations in the downlink direction. Through these technical means, data exchange and functional linkage between edge computing power scheduling, cross-camera pedestrian matching, and model iteration are achieved, which helps to simultaneously improve matching failures and decreased recognition accuracy caused by computing power fluctuations.
Claims
1. A cloud-edge collaborative intelligent analysis system for shopping malls, characterized in that, Includes integrated edge-aware computing nodes, cloud platform, and model iteration module; The integrated edge-aware computing node is deployed in each customer flow zone of the shopping mall, configured with edge computing resources, runs a lightweight first deep learning model to complete pedestrian detection and feature extraction, collects the computing load of each zone in real time and synchronously uploads structured metadata and load data to the cloud platform. The cloud platform is equipped with cloud computing resources, accesses shopping mall business data, and has a cross-camera pedestrian re-identification unit; The model iteration module is deployed on a cloud platform. It completes incremental training of the model based on the identification of difficult examples and generates a lightweight model through knowledge distillation, which is then distributed to each edge perception computing node. The difficult examples refer to pedestrian identification samples that are missed or falsely reported by the system. The knowledge distillation refers to the technique of transferring the knowledge of the complex teacher model to the simplified student model. This system also features a two-way closed-loop collaborative mechanism for computing power and data. This mechanism uses the partitioned real-time computing power load as the coupling link and includes two main logics: uplink feature weighting and downlink distillation control. (1) The cross-camera pedestrian re-identification unit integrates the three-dimensional features of pedestrian appearance, movement sequence and mall spatial location, and uses the partition computing power load data uploaded from the edge as a weighting factor to correct the weight of mall spatial location features; (2) The model iteration module adopts a customized distillation loss function for the dense occlusion scenario of shopping malls, increases the loss weight of occlusion samples, and dynamically switches the distillation iteration frequency according to the real-time computing power load of each partition. The system simultaneously outputs security early warning results and quantitative analysis results of shop traffic value and marketing activity conversion effect.
2. The intelligent shopping mall analysis system based on cloud-edge collaboration according to claim 1, characterized in that, The integrated edge-aware computing node is configured with a partitioned adaptive computing power scheduling mechanism, which sets two types of computing power occupancy thresholds: high and low. When the computing power occupancy exceeds the high threshold, the inference frame rate is reduced; when the computing power occupancy is below the low threshold, the inference frame rate is increased.
3. The intelligent shopping mall analysis system based on cloud-edge collaboration according to claim 1, characterized in that, It also includes an application visualization and linkage module; the application visualization and linkage module realizes multi-terminal hierarchical data display, identifies abnormal events in the mall and pushes alarms in a hierarchical manner, and links broadcasting, access control and fire emergency equipment.
4. The intelligent shopping mall analysis system based on cloud-edge collaboration according to claim 3, characterized in that, The application visualization linkage module is also used to associate customer flow trajectory, crowd density and store transaction, membership and parking data through the multi-source data fusion engine, and output quantitative analysis conclusions on store traffic value and marketing activity effectiveness.
5. A cloud-edge collaborative intelligent analysis method for shopping malls, characterized in that, include: Step 1: The edge nodes of each zone in the shopping mall run a lightweight first deep learning model to complete pedestrian detection and tracking, collect the computing load of each zone, and only upload structured metadata and computing load data to the cloud platform synchronously. Step 2: The cloud platform aggregates structured metadata and computing load data uploaded from the edge, and connects to the mall's operational data; Step 3: The cloud platform runs the second deep learning model, which integrates the three-dimensional features of pedestrian appearance, movement sequence, and mall spatial location to complete cross-camera pedestrian matching and outputs the results of abnormal customer flow and business analysis. Step 4: Collect difficult examples of scene recognition to incrementally train the cloud model, generate a lightweight model through knowledge distillation and distribute it to each edge node. The difficult examples refer to pedestrian recognition samples that are missed or falsely reported by the system. The knowledge distillation refers to the technique of transferring the knowledge of the complex teacher model to the simplified student model. The method further includes a two-way closed-loop collaborative logic for computing power data, which relies on partitioned computing power load to achieve two-way control, including: (1) In the matching process of step three, the partition computing load data uploaded from the edge is introduced as a weighting factor to correct the fusion weight of the shopping mall spatial location features; (2) In the distillation process of step four, a customized distillation loss function for the dense occlusion scenario of the shopping mall is adopted to increase the loss weight of the occlusion sample, and the distillation iteration frequency is dynamically adjusted according to the real-time computing power load of each partition. The method simultaneously outputs security early warning information and quantitative analysis data on store traffic value and marketing activity conversion revenue.
6. The intelligent shopping mall analysis method based on cloud-edge collaboration according to claim 5, characterized in that, The logic for adaptively adjusting the inference frame rate in step one is as follows: if the computing power usage in high-density areas exceeds the threshold, the frame rate is lowered; if the computing power usage in low-density areas such as shops and parking lots is low, the inference frame rate is raised.
7. The intelligent shopping mall analysis method based on cloud-edge collaboration according to claim 5, characterized in that, When fusing 3D pedestrian features in step three, the weighting coefficients of the mall's spatial location features are dynamically adjusted using the real-time computing power load values of each zone.
8. The intelligent shopping mall analysis method based on cloud-edge collaboration according to claim 5, characterized in that, When performing knowledge distillation in step four, the number of distillation iterations is switched according to the computing power load range of each partition, and higher loss weights are assigned to pedestrian overlapping and occlusion samples. The customized distillation loss function consists of three parts: standard cross-entropy loss, divergence loss, and occlusion customized loss term. The occlusion customized loss term assigns an additional weight to the target that is determined to be an occluded sample, which is 2 to 4 times the weight of the regular sample.
9. The intelligent shopping mall analysis method based on cloud-edge collaboration according to claim 5, characterized in that, After completing the multi-source data fusion in step two, the data on customer flow, transactions, membership, and parking are linked to calculate the value of store traffic and the conversion revenue of marketing activities.
10. The intelligent shopping mall analysis method based on cloud-edge collaboration according to claim 5, characterized in that, It also includes a visual linkage process: hierarchical display of customer flow, operation, and early warning data, and push of alarms and linkage with on-site emergency equipment after identifying mall anomalies.