Logistics operation data processing method and device, electronic equipment and storage medium

By deploying video screening and operation classification models on edge servers and combining them with cloud verification, the accuracy problem of identifying logistics violations in logistics operation scenarios has been solved, and adaptive optimization and accuracy improvement of the model have been achieved.

CN121963029APending Publication Date: 2026-05-01SF TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SF TECH CO LTD
Filing Date
2025-12-29
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In different logistics operation scenarios, existing technologies are unable to accurately identify logistics violations, resulting in untimely warnings.

Method used

By deploying video screening and operation classification models on edge servers, logistics operation videos are coarsely screened and finely classified. Video features are extracted by combining edge multimodal large models, and then verified on cloud servers to update edge model parameters to adapt to scene changes.

Benefits of technology

It improves the accuracy of identifying logistics violations in logistics operation scenarios and ensures the model's continuous self-optimization and accuracy in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963029A_ABST
    Figure CN121963029A_ABST
Patent Text Reader

Abstract

The invention provides a logistics operation data processing method and device, electronic equipment and a storage medium, and the method is applied to an edge server, and comprises the steps: carrying out the illegal segment screening of a logistics operation video collected in a target logistics operation scene based on a video screening model, and obtaining a video screening result, carrying out violation operation classification on the violation abnormal segment based on an operation classification model corresponding to the target logistics operation scene to obtain a corresponding violation operation classification result; sending the corresponding logistics operation video to a cloud server for re-checking, and receiving an illegal operation re-checking result returned by the cloud server; when the re-checking is not passed, obtaining a reference violation operation category and a reference operation video feature generated by the cloud server; and performing parameter updating on an operation classification model corresponding to the target logistics operation scene according to the reference operation video features and the reference illegal operation categories. According to the invention, the accuracy of identifying the logistics violation operation in different logistics operation scenes can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Logistics operation data processing methods and devices, electronic equipment and storage media Technical Field

[0001] This application relates to the fields of computer vision and intelligent security technology, and in particular to a logistics operation data processing method and apparatus, electronic device and storage medium. Background Technology

[0002] Logistics transit sites are key hubs in the logistics system. They encompass various logistics operation scenarios, and the efficient and compliant operation of personnel and equipment is crucial for ensuring the normal functioning of these sites.

[0003] In actual operations, non-compliant logistics operations occur frequently, such as staff not wearing safety protective equipment, entering restricted areas, vehicles speeding, or vehicles remaining in dangerous areas. To ensure the safety of personnel on site and the standardized operation of equipment, it is necessary to detect and promptly issue warnings for non-compliant logistics operations.

[0004] However, different logistics operation scenarios (such as unloading, sorting, transportation, and loading) differ in terms of goods, facilities (conveyor belts and stacking machines), operational processes, and operating environments, leading to variations in potential logistics violations. Related technologies have low accuracy in identifying logistics violations across different operation scenarios, resulting in delayed warnings. Summary of the Invention

[0005] The main objective of this application is to provide a logistics operation data processing method, apparatus, electronic device, and storage medium, which aims to improve the accuracy of identifying logistics violations in different logistics operation scenarios.

[0006] To achieve the above objectives, a first aspect of this application proposes a logistics operation data processing method, applied to an edge server corresponding to a target logistics site. The target logistics site includes multiple target logistics operation scenarios. The edge server is equipped with a trained video filtering model and an operation classification model corresponding to different logistics violations in each target logistics operation scenario. The method includes: filtering violation segments from logistics operation videos collected in the target logistics operation scenarios based on the video filtering model to obtain video filtering results; and classifying the violation and abnormal segments indicated by the video filtering results based on the operation classification model corresponding to the target logistics operation scenario to obtain corresponding violation operation classifications. The system retrieves the violation operation classification results; sends the video screening results and the corresponding logistics operation videos to the cloud server for violation operation review, and receives the violation operation review results returned by the cloud server; when the violation operation review results indicate that at least one of the video screening results and the violation operation classification results fails the review, it obtains the reference violation operation category generated by the cloud server based on the logistics operation videos that failed the review, and the reference operation video features extracted from the logistics operation videos that failed the review; and updates the parameters of the operation classification model corresponding to the target logistics operation scenario according to the reference operation video features and the reference violation operation category.

[0007] To achieve the above objectives, a second aspect of this application proposes a logistics operation data processing device applied to an edge server corresponding to a target logistics site. The target logistics site includes multiple target logistics operation scenarios. The edge server is equipped with a trained video filtering model and an operation classification model corresponding to different logistics violations in each target logistics operation scenario. The device includes: a classification unit, configured to filter violation segments from logistics operation videos collected in the target logistics operation scenarios based on the video filtering model to obtain video filtering results, and to classify the violation and abnormal segments indicated by the video filtering results into violation operation categories based on the operation classification model corresponding to the target logistics operation scenario, thereby obtaining corresponding violation operation categories. The system includes a receiving unit, configured to send the video screening results and the logistics operation videos corresponding to the violation operation classification results to a cloud server for violation operation review, and receive the violation operation review results returned by the cloud server; an acquisition unit, configured to acquire, when the violation operation review results indicate that at least one of the video screening results and the violation operation classification results has failed review, a reference violation operation category generated by the cloud server based on the logistics operation videos that have failed review, and reference operation video features extracted from the logistics operation videos that have failed review; and an update unit, configured to update the parameters of the operation classification model corresponding to the target logistics operation scenario based on the reference operation video features and the reference violation operation category.

[0008] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.

[0009] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.

[0010] The logistics operation data processing method, apparatus, electronic device, and storage medium proposed in this application perform coarse screening of large-scale logistics operation videos using a video screening model on an edge server, and fine classification of violation and abnormal segments using an operation classification model corresponding to the logistics operation scenario. The video screening results and violation operation classification results are then uploaded to a cloud server for violation operation verification to further ensure the accuracy of the classification results. When the verification fails, the edge server can obtain the features of the reference operation video and the reference violation operation category generated by the cloud server to update the parameters of the operation classification model corresponding to the target logistics operation scenario. This allows the edge model to learn online and correct its own cognitive biases, continuously adapting to changes in specific logistics operation scenarios and achieving continuous self-optimization of model performance. This improves the accuracy of the edge model in identifying logistics violation operations in different logistics operation scenarios. Attached Figure Description

[0011] Figure 1 is a flowchart of the logistics operation data processing method provided in the embodiment of this application; Figure 2 is a flowchart of step S102 in Figure 1 provided in the embodiment of this application; Figure 3 is a flowchart of the cloud server generating the review result of the violation operation provided in the embodiment of this application; Figure 4 is a flowchart of the model adaptive fine-tuning provided in the embodiment of this application; Figure 5 is a flowchart of the cloud server performing anomaly detection provided in the embodiment of this application; Figure 6 is a schematic diagram of the logistics operation data processing system provided in the embodiment of this application; Figure 7 is a structural schematic diagram of the logistics operation data processing device provided in the embodiment of this application; Figure 8 is a hardware structure schematic diagram of the electronic device provided in the embodiment of this application. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0013] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0014] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0015] Logistics transit sites are key hubs in the logistics system. They encompass various logistics operation scenarios, and the efficient and compliant operation of personnel and equipment is crucial for ensuring the normal functioning of these sites.

[0016] In actual operations, non-compliant logistics operations occur frequently, such as staff not wearing safety protective equipment, entering restricted areas, vehicles speeding, or vehicles remaining in dangerous areas. To ensure the safety of personnel on site and the standardized operation of equipment, it is necessary to detect and promptly issue warnings for non-compliant logistics operations.

[0017] However, different logistics operation scenarios (such as unloading, sorting, transportation and loading) have differences in goods, facilities (conveyor belts and stacking machines), operation processes and operation environments, which leads to differences in the logistics violations that may occur in different logistics operation scenarios.

[0018] The relevant technologies identify logistics violations in the following ways: (1) Relying on manual duty and fixed security rules, specifically, relying on security patrols and central control room duty, supplemented by video recording and simple electronic fences (such as virtual warning lines and area intrusion alarms). This method is simple to deploy and has low cost. However, video recording is sensitive to the effects of lighting, occlusion, and image shaking. Moreover, the alarm for logistics violations depends on fixed thresholds, which is prone to false alarms. The video recording cannot understand complex logistics violations, such as forklifts being too close to pedestrians and continuously following them. (2) Rule engines based on visual algorithms. This method combines background subtraction, motion detection, color thresholds or shape thresholds, wireframe triggers or area triggers, etc., with time thresholds to form warning rules. This method is highly interpretable and easy to implement. However, it is not robust to changes in logistics scene environment (rain, snow, shadows, and night, etc.) and is sensitive to the shaking of the viewpoint and camera. It is difficult to identify wearing norms and complex interactive behaviors.

[0019] Therefore, the relevant technologies are greatly affected by changes in different logistics operation scenarios, and the accuracy of identifying logistics violations in different logistics operation scenarios is low, resulting in untimely early warnings.

[0020] Based on this, embodiments of this application provide a logistics operation data processing method and apparatus, electronic device and storage medium, aiming to improve the accuracy of identifying logistics violations in different logistics operation scenarios.

[0021] The logistics operation data processing method, apparatus, electronic device, and storage medium provided in this application are specifically described through the following embodiments. First, the logistics operation data processing method in this application embodiment is described.

[0022] The logistics operation data processing method provided in this application relates to the fields of computer vision and intelligent security technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the logistics operation data processing method, but is not limited to the above forms.

[0023] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0024] Figure 1 is an optional flowchart of a logistics operation data processing method provided in an embodiment of this application. The method in Figure 1 may include, but is not limited to, steps S101 to S104.

[0025] Step S101: Based on the video filtering model, filter the logistics operation videos collected in the target logistics operation scenario to obtain the video filtering results. Then, based on the operation classification model corresponding to the target logistics operation scenario, classify the abnormal segments indicated by the video filtering results into violations, and obtain the corresponding violation classification results. Step S102: Send the video filtering results and the logistics operation videos corresponding to the violation classification results to the cloud server for violation review, and receive the violation review results returned by the cloud server. Step S103: When the violation review results indicate that at least one of the video filtering results and the violation classification results fails the review, obtain the reference violation category generated by the cloud server based on the logistics operation videos that fail the review, and the reference operation video features extracted from the logistics operation videos that fail the review. Step S104: Update the parameters of the operation classification model corresponding to the target logistics operation scenario according to the reference operation video features and the reference violation category.

[0026] Steps S101 to S104, as illustrated in this embodiment, involve using an edge server's video screening model to coarsely screen a large number of logistics operation videos, and then using an operation classification model corresponding to the logistics operation scenario to finely classify the violation and abnormal segments. The video screening results and violation classification results are then uploaded to a cloud server for violation verification to further ensure the accuracy of the classification results. If the verification fails, the edge server can obtain the features of a reference operation video and a reference violation category generated by the cloud server to update the parameters of the operation classification model corresponding to the target logistics operation scenario. This allows the edge model to learn online and correct its own cognitive biases, continuously adapting to changes in specific logistics operation scenarios and achieving continuous self-optimization of model performance. This improves the accuracy of the edge model in identifying logistics violations in different logistics operation scenarios.

[0027] The logistics operation data processing method provided in this application is applied to an edge server corresponding to a target logistics site. The edge server is a computer device deployed locally at the logistics site, used to process large-scale logistics operation video data locally, avoiding the need to transmit all raw video data back to the cloud and solving bandwidth and latency issues. The target logistics site may include a logistics transfer area or a logistics sorting center, and includes numerous elements such as personnel, vehicles (e.g., forklifts, trucks, or automated guided vehicles). Multiple elements, including vehicles (AGVs), equipment (such as conveyor belts and stacker cranes), and various goods, operate in parallel. A target logistics site can include multiple target logistics operation scenarios, each corresponding to a different functional area within the site. For example, a logistics sorting area may include an unloading area, a sorting area, a cargo transportation area, and a loading area. The elements, operational standards, and violation definitions differ across these scenarios, leading to variations in potential logistics violations. For instance, in an unloading scenario, the main elements include trucks, unloading ports, telescopic conveyor belts, and unloading personnel. Possible violations include trucks not aligning with the unloading port and unloading personnel not wearing reflective vests as required inside the truck. Similarly, in a sorting scenario, the main elements include sorting machines, sorters, and packages. Possible violations include sorters throwing goods or climbing over operating conveyor belts.

[0028] The edge server is equipped with a video screening model, which is used to perform coarse screening of large-scale logistics operation videos. It can quickly identify video segments containing specific subjects or specific logistics operations from large-scale logistics operation videos through lightweight object detection, attribute detection and object tracking. This filters out video segments that may contain logistics violations and removes meaningless video footage, thereby reducing the bandwidth consumption of logistics operation data being uploaded to the cloud.

[0029] The edge server also deploys operation classification models corresponding to different logistics violations in each target logistics operation scenario. These models can be used to refine video segments potentially containing logistics violations, as identified by the video filtering model, further determining the video segments actually containing violations and identifying the corresponding violations. The operation classification model is a binary classification model trained based on sample videos of potential logistics violations within the target logistics operation scenario. Therefore, this embodiment does not employ general, fixed rules for determining logistics violations, but rather uses operation classification models customized for specific logistics operation scenarios to identify violations. Multiple operation classification models can be set up for a single target logistics operation scenario. For example, in a cargo transportation operation scenario, operation classification models can be set up for violations such as speeding by transport equipment or close contact with transport equipment; while in a sorting operation scenario, operation classification models can be set up for violations such as illegally crossing conveyor belts or throwing goods.

[0030] In step S101 of some embodiments, the video screening results can be obtained by first screening the illegal segments of the logistics operation video collected in the target logistics operation scenario based on the video screening model. The video screening results can include multiple video segments that are suspected of having illegal logistics operations. The video duration of each video segment can be within a preset duration range, for example, a video segment of 6s-10s.

[0031] In some embodiments, the video filtering model includes a target detection module, a target attribute recognition module, and a target tracking module. Based on the video filtering model, violation segments are filtered from logistics operation videos collected in a target logistics operation scenario to obtain video filtering results. This includes the following steps: calling the target detection module to perform target detection on the logistics operation videos collected in the target logistics operation scenario, obtaining target detection results; calling the target attribute recognition module to perform attribute recognition on the target detection results, obtaining target attribute recognition results; calling the target tracking module to perform video tracking on the detected objects contained in the target detection results, obtaining target object tracking videos; and filtering violation segments based on the target object tracking videos to obtain video filtering results.

[0032] In this embodiment, the video filtering model is pre-trained based on training sample data. Depending on the task type performed by the video filtering model, a corresponding loss function can be used for parameter updates. Specifically, the video filtering model includes an object detection module, an object attribute recognition module, and an object tracking module. For detection tasks in the object detection module, binary cross-entropy (BCE) loss and intersection over union (IoU) loss can be used.

[0033] First, the object detection module can be called to perform object detection on the logistics operation videos collected in the target logistics operation scene, obtaining the object detection results. Then, the object attribute recognition module can be called to perform attribute recognition on the object detection results, obtaining the object attribute recognition results. Specifically, the object detection module can be implemented using a lightweight deep learning network (such as the YOLO model). The object detection module is used for localization and classification, that is, it selects the positions of detection targets such as personnel, forklifts, and goods in each video frame of the logistics operation video (e.g., detection boxes) and labels the category corresponding to each detection target, obtaining the object detection results. The object attribute recognition module is used to further identify the target attributes corresponding to the detection targets based on the object detection module, obtaining the object attribute recognition results. For example, whether the personnel are wearing safety helmets, whether they are wearing reflective vests, or whether they are holding objects such as mobile phones or cigarettes, or whether the personnel are wearing hats, and what color the hats are.

[0034] Then, the target tracking module can be invoked to perform video tracking on the detected targets contained in the target detection results, obtaining the target object tracking video. Based on the target object tracking video, violation segments can be filtered to obtain the video filtering results. Specifically, the target tracking module can be implemented using multi-target tracking algorithms (such as DeepSORT or ByteTrack). This module performs the object tracking task and is pre-trained before being deployed to the edge server. During training, triplet loss and re-ID loss can be used for gradient updates. The target tracking module associates discrete detection boxes by matching visual features between consecutive frames in the logistics operation video, obtaining a continuous motion trajectory, i.e., the target object tracking video, thereby calculating the target's speed, dwell time, or direction of motion. The target object tracking video includes video segments corresponding to the complete motion process of the detected target. Then, based on the preset rule engine (such as area intrusion logic, speed threshold logic), combined with the target attribute recognition results and target object tracking videos obtained in the previous steps, video segments that meet the violation judgment rules set in the rule engine are filtered out, and video filtering results are obtained, thereby filtering out video segments of normal operations.

[0035] This application provides an efficient edge-end video coarse-screening mechanism. By combining target detection, attribute recognition, and video tracking, it can perceive not only static features (such as whether workers are wearing uniforms) but also dynamic behaviors (such as speeding or driving in the wrong direction), increasing the detection dimensions of illegal logistics operations and improving the accuracy of illegal logistics operation identification. Furthermore, this application can compress invalid data at the edge, sending only the video screening results that may indicate illegal logistics operations to subsequent processes (such as uploading to the cloud), reducing network transmission bandwidth consumption.

[0036] After generating video screening results through the aforementioned steps, the violation and abnormal segments indicated by the video screening results can be classified into violation operations based on the operation classification model corresponding to the target logistics operation scenario. The operation classification model is a lightweight classification network (such as a fully connected layer or a simple classifier). It has a simple structure, fast inference speed, and can adapt to the corresponding target logistics operation scenario (such as unloading area, sorting area, or transportation channel). The operation classification model is pre-trained using training sample data. The training sample data includes sample logistics operation videos containing violation operations collected from the corresponding target logistics operation scenario, or it can be obtained from relevant databases. The training sample data also includes the sample violation operation category corresponding to each sample logistics operation video. The operation classification model takes the sample logistics operation video as input and updates its parameters with the corresponding sample violation operation category as the output label. Then, the offline-trained operation classification model is deployed to the corresponding edge server. Thus, the operation classification model provided in this embodiment can adapt to the corresponding logistics operation scenario, such as adapting to a logistics operation scenario with low lighting or a specific color of work clothing, and finally output the specific violation operation category.

[0037] Then, the operation classification model can be used to determine the specific violation type of the video segments indicating abnormal violations in the video screening results. The video segments indicating abnormal violations are those that may involve logistics violations. The operation classification model can analyze the screened video segments to obtain violation classification results, which include specific violation types, such as not wearing a safety helmet, smoking in violation of regulations, people entering restricted areas, and close interaction between forklifts and pedestrians.

[0038] In some embodiments, the edge server is further deployed with an edge multimodal large model, which performs violation operation classification on the violation and abnormal segments indicated by the video screening results based on the operation classification model corresponding to the target logistics operation scenario, and obtains the corresponding violation operation classification results. The steps include: extracting features from the violation and abnormal segments indicated by the video screening results based on the edge multimodal large model to obtain violation and abnormal video features; and performing violation operation classification on the violation and abnormal video features based on the operation classification model corresponding to the target logistics operation scenario to obtain the corresponding violation operation classification results.

[0039] In this embodiment, an edge multimodal large model is also deployed in the edge server. Both the edge multimodal large model and the operation classification model belong to the adaptive classification module in the edge server. The edge multimodal large model refers to a neural network model capable of jointly understanding multiple modalities such as video and text. This model is a lightweight, large model with fewer parameters. Before deploying the edge multimodal large model to the edge server, it can be trained. Specifically, video-text pairs can be obtained from the logistics vertical domain, and then filtered and cleaned. This includes removing low-quality videos (such as videos with too low resolution), cleaning noise in the text (such as HTML tags and garbled characters), and ensuring the semantic relevance of the video-text pairs. During training, the edge multimodal model takes video as input and corresponding text descriptions as output. Parameters are updated using contrastive loss. Specifically, for N pairs of images and text in a batch, the cosine similarity between all image embeddings and text embeddings is calculated. Elements on the diagonal are correct pairs, and the rest are negative samples. Symmetric cross-entropy is used for model optimization, thus making semantically similar videos (such as different people performing the same throwing action) appear closer in the feature space. In the use of the edge multimodal model, the input is a video segment, and the output is the corresponding embedding vector.

[0040] The edge multimodal large model can extract features from violation and anomalous segments indicated by video screening results. Specifically, it first acquires the violation and anomalous segments screened in previous steps, then performs preprocessing operations such as frame extraction, size adjustment, and normalization on these segments before inputting them into the edge multimodal large model. The edge multimodal large model transforms the unstructured violation and anomalous segments into high-dimensional numerical vectors (i.e., embedding vectors), obtaining violation and anomalous video features. These features represent the semantic information within the violation and anomalous segments. The edge multimodal large model does not directly output classification results; instead, it outputs a fixed-dimensional embedding vector, which is the violation and anomalous video feature.

[0041] Then, based on the operation classification model corresponding to the target logistics operation scenario, the features of the violation and abnormal video can be classified into violation operation categories to obtain the corresponding violation operation classification results. The operation classification model can map the input features of the violation and abnormal video to specific violation behavior category labels. As mentioned above, the operation classification model is a binary classification model, that is, each operation classification model focuses on determining whether a certain logistics violation operation exists, such as whether it is a violent sorting behavior.

[0042] The operation classification model processes the abnormal video features using weight matrix operations and binary activation functions (such as Sigmoid or Softmax for binary classification). It calculates the probability that the abnormal video features belong to the target abnormal operation category and also to the normal operation category or other operation categories (non-target abnormal operation category). Then, by comparing these two probability values ​​or determining whether the probability of belonging to the target abnormal operation category exceeds a preset probability threshold (e.g., 0.5), the operation classification result is determined based on the result with the highest probability. For example, in a sorting operation scenario, the operation classification model determines whether the input abnormal video features meet the definition of violent sorting; if so, it outputs the violation classification result; if not, it is determined to be a non-violation category.

[0043] It should be noted that the operation classification model can classify violations based on the violation and abnormal video features output by the edge multimodal large model, or it can perform a simple task of classifying violations based on images alone. The edge multimodal large model can directly classify violations and abnormal segments in the video screening results, that is, perform a more complex violation classification task. Alternatively, it can generate corresponding violation and abnormal video features based on violation and abnormal segments, and then input the generated violation and abnormal video features into the operation classification model for specific classification of logistics violations.

[0044] This application's embodiments achieve efficient and flexible edge computing by decoupling feature extraction and operation classification. First, semantic features of the video are extracted using a large-scale multimodal edge model. Then, a scenario-based operation classification model is used to classify violations. Since the operation classification model uses feature vectors instead of raw video pixels as input, the computational load at the edge is reduced, improving inference efficiency. By deploying corresponding operation classification models for different operational scenarios, the features of the violation / abnormal video are mapped to corresponding semantic results. This solves the problem in related technologies where a single, cross-scenario general model is used to uniformly identify logistics violations in multiple logistics operation scenarios, resulting in low accuracy and high false alarm rates. Therefore, this approach improves the accuracy of identifying logistics violations in different logistics operation scenarios.

[0045] In step S102 of some embodiments, the edge server sends the video screening results and the logistics operation videos corresponding to the violation operation classification results to the cloud server for violation operation verification, and receives the violation operation verification results returned by the cloud server. The edge server uploads the screened and classified video data to the cloud server, realizing edge-cloud data interaction. The cloud server possesses more powerful computing resources and can perform secondary verification on the video screening results and the logistics operation videos corresponding to the violation operation classification results uploaded by the edge server. That is, it performs violation operation verification on the logistics operation videos uploaded by the edge server and corrects any false positives or false negatives that the model in the edge server might generate. For example, the model in the edge server might misjudge normal cargo handling operations as illegal cargo throwing operations due to line-of-sight obstruction or blurred vision in the logistics operation scenario. Or, the model deployed in the edge server might encounter a logistics violation operation it has never seen before, generating uncertain samples. In this case, the cloud server can use its powerful generalization ability to identify the logistics operations in the uncertain samples and obtain the accurate category of the violation operation. Then, the cloud server can return the generated violation operation verification results to the edge server, thereby realizing a closed-loop decision-making process for logistics violation operation identification.

[0046] In some embodiments, referring to Figure 2, sending the candidate logistics operation videos corresponding to the video screening results and the violation operation classification results to the cloud server for violation operation review includes the following steps S201 to S202: Step S201, obtaining the model classification performance parameters corresponding to the violation operation classification results within a preset time interval; Step S202, when at least one parameter in the model classification performance parameters indicates a decrease in the classification performance of the operation classification model, sending the logistics operation videos corresponding to the violation operation classification results within the preset time interval to the cloud server for violation operation review, and sending the logistics operation videos corresponding to the video screening results to the cloud server for violation operation review.

[0047] The edge model is a lightweight model, and the environment of the target logistics site (such as lighting, background, and camera angle) is dynamically changing. Therefore, the edge model may experience concept drift during long-term operation, leading to performance degradation. To ensure the long-term stable operation of the logistics security system, the edge server needs to monitor the health status and performance parameters of the edge model in real time and take appropriate measures when performance degradation is detected.

[0048] In step S201 of some embodiments, the edge server can obtain model classification performance parameters corresponding to the classification results of violations within a preset time interval. The preset time interval is a sliding time window (e.g., the past hour or the past week). The model classification performance parameters are used to analyze the short-term performance of the edge-side model and can reflect the confidence or stability of the edge-side model's judgment results on the current input data. Specifically, the edge server can detect data distribution drift or concept drift in the edge-side model. This is because, over time, lighting conditions (day and night alternation), background environment (changes in cargo stacking), camera angle (minor displacement), or the manifestation of violations (e.g., new work clothes) in the target logistics operation scenario may change, causing the data distribution during edge-side model training to be inconsistent with the current actual inference data distribution, thereby causing the model to fail. Data distribution drift or concept drift can be detected specifically through the confidence distribution, information entropy of the output probability distribution, or distribution distance of the feature vectors of the edge-side model. The confidence distribution is used to measure the accuracy of the results generated by the edge-side model, and the information entropy is used to measure the uncertainty of the edge-side model regarding the generated results. The higher the entropy value, the higher the uncertainty. Furthermore, the health of the edge-side model can be calculated based on the above data. Specifically, corresponding weight coefficients can be set for confidence and information entropy, and the corresponding parameters can be weighted and summed according to the weight coefficients to obtain the model health score. If the health score is higher than the preset health threshold, it indicates that the model is adapted to the current environment and is running well; if the health score is lower than the preset health threshold, it indicates that the model is not performing well and intervention is required. For example, if the edge-side model fails to meet the preset confidence threshold for classifying multiple logistics operation videos within a certain period of time, it indicates that the reliability of the violation operation categories generated by the edge-side model is not high, or the information entropy value corresponding to the output result of the edge-side model is high, indicating that the model has a high degree of uncertainty regarding the current output result.

[0049] It should be noted that the video input to the edge server can also be observed manually. If the operator observes that there are logistics violations in the logistics operation video, but they are not identified by the edge model, then this segment of the logistics operation video that was not identified can be extracted separately and labeled with hard sample and corresponding violation category.

[0050] In step S202 of some embodiments, when the edge server detects a decline in the health index of the model, or detects that difficult-to-handle samples cause fluctuations in accuracy, the edge server can automatically send the inaccurately identified logistics operation videos to the cloud server for review of violations. Alternatively, when the edge-side model's confidence level is detected to be lower than a preset confidence threshold, and the edge-side model cannot recover its accuracy in the short term, service degradation or a rollback strategy can be triggered. Specifically, the edge-side model can send more logistics operation videos to the cloud server, and the more capable cloud server can identify and review the violation categories of the logistics operation videos, reducing the need for the edge-side model to perform violation classification.

[0051] As mentioned above, operators can also proactively identify unreported or falsely reported logistics operation videos. They can then add corresponding tags to the detected unreported or falsely reported logistics operation videos to indicate that the videos need to be uploaded to the cloud for review of violations. Finally, the logistics operation videos and corresponding tags are sent to the cloud server.

[0052] Specifically, when at least one parameter in the model classification performance parameters indicates a decline in the classification performance of the operation classification model, the edge server can send the logistics operation videos corresponding to the violation operation classification results within a preset time interval to the cloud server for violation operation review, and also send the logistics operation videos corresponding to the video screening results to the cloud server for violation operation review. The edge server uploads not only videos already identified as violations by the operation classification model (logistics operation videos corresponding to violation operation classification results), but also videos that passed the initial screening but may not have been confirmed as violations (logistics operation videos corresponding to video screening results). This is because when the performance of the edge-side model declines, there is a high possibility of missed reports (i.e., the initial screening detects anomalies, but the classification model does not recognize them). Therefore, it is necessary to package and upload both types of videos together to the cloud server for violation operation review.

[0053] In this embodiment, when a decline in the performance parameters of the edge model is detected, the decision-making authority for classifying violations can be transferred to the cloud server, ensuring the accuracy of violation identification when the edge model performance deteriorates. When performance deteriorates, both the coarse and fine screening videos of logistics operations are uploaded to the cloud server. The cloud server can accurately capture logistics operation videos that are difficult to judge or are incorrectly judged, leading to a decline in edge model performance, thereby improving the accuracy of identifying logistics violations.

[0054] In some embodiments, referring to Figure 3, a cloud-based multimodal large model is deployed in the cloud server, and the edge multimodal large model is obtained by knowledge distillation from the cloud-based multimodal large model. The process of the cloud server generating the violation operation review result includes the following steps S301 to S304: Step S301, based on the cloud-based multimodal large model, feature extraction is performed on the logistics operation videos corresponding to the video screening results and violation operation classification results to obtain reference operation video features; Step S302, based on the cloud-based multimodal large model, a reference violation operation category corresponding to each logistics operation video is generated using the logistics operation videos corresponding to the video screening results and violation operation classification results; Step S303, the reference operation video features are matched with the violation abnormal video features to obtain a first matching result, and the reference violation operation category is matched with the violation operation classification result to obtain a second matching result; Step S304, when at least one of the first matching result and the second matching result indicates a mismatch, a violation operation review result indicating that at least one of the video screening results and the violation operation classification result has failed the review is generated.

[0055] As described above, the cloud server deploys a cloud-based multimodal large-scale model, and the aforementioned edge multimodal large-scale model is obtained through knowledge distillation from the cloud-based model. The cloud-based model acts as the teacher model, possessing a large parameter scale, a deep network structure, and high-precision feature representation capabilities trained on large-scale general data and logistics vertical domain data. The edge multimodal model, on the other hand, acts as the student model, with a relatively streamlined structure and fewer parameters, making it suitable for the limited storage and computing resources of the edge server. During training, the edge multimodal model not only learns the ground truth but also approximates the cloud-based model's thought process by minimizing the difference between its output distribution and that of the cloud-based model (such as KL divergence). In this way, the edge multimodal model can retain a very low parameter count while maximizing the inheritance of the cloud-based model's generalization understanding and feature extraction capabilities for complex logistics operation scenarios.

[0056] The parameter scale and generalization ability of cloud-based multimodal large-scale models far exceed those of edge-side models. Due to their superior generalization and feature extraction capabilities, cloud-based multimodal large-scale models can perform secondary verification on logistics operation videos corresponding to the video screening results and violation classification results uploaded from edge servers. This involves reviewing the logistics operation videos uploaded to edge servers for violations and correcting any false positives or false negatives that might occur in the models on the edge servers. For example, the model on the edge server might misclassify normal goods handling operations as illegal throwing operations due to visual obstruction or blurred vision in the logistics operation scenario. Alternatively, the model deployed on the edge server might encounter a logistics violation operation it has never seen before, generating uncertain samples. In this case, the cloud-based multimodal large-scale model can use its powerful generalization ability to identify the logistics operations in the uncertain samples and obtain the accurate category of the violation. The cloud server can then return the generated violation verification results to the edge server, thus achieving a closed-loop decision-making process for identifying logistics violations.

[0057] In step S301 of some embodiments, features are extracted from the logistics operation videos corresponding to the video screening results and violation operation classification results based on the cloud-based multimodal large model to obtain reference operation video features. The reference operation video features refer to high-dimensional vectors (Embeddings) generated by the cloud-based multimodal large model that represent the standard semantics of the logistics operation videos, and are used to measure the accuracy of the violation and abnormal video features extracted by the edge multimodal large model.

[0058] In step S302 of some embodiments, a reference violation category can be generated for each logistics operation video based on the logistics operation video corresponding to the video screening results and violation operation classification results using a cloud-based multimodal large model. Since the cloud-based multimodal large model is pre-trained using large-scale video-text pairs, it can generate corresponding feature vectors and violation operation categories based on the input video segments. Therefore, the cloud-based multimodal large model can perform multimodal analysis on the received logistics operation videos and regenerate the corresponding reference violation operation categories.

[0059] In step S303 of some embodiments, the reference operation video features and the violation / abnormal video features can be matched for similarity to obtain a first matching result. The first matching result is used to measure the semantic space consistency between the violation / abnormal video features extracted by the edge multimodal large model and the reference operation video features extracted by the cloud multimodal large model. Specifically, the similarity between the two feature vectors can be determined by calculating the cosine similarity or Euclidean distance. Taking the calculation of the cosine similarity between two features as an example, if the similarity is higher than a preset similarity threshold (e.g., 0.9), the first matching result indicates that the violation / abnormal video features and the reference operation video features have a high similarity and are matched. If the similarity does not exceed the preset similarity threshold, the first matching result indicates that the violation / abnormal video features and the reference operation video features have a low similarity and are not matched. Similarly, a second matching result can be obtained by matching the reference violation category with the violation classification result. This second matching result is obtained by comparing the text labels of the violation category output by the edge-side operation classification model with the ground truth labels generated by the cloud-based multimodal large model. If the violation classification result generated by the edge side is the same as the reference violation category generated by the cloud-based multimodal large model (e.g., the edge side determines the violation category of the logistics operation video to be category A, and the cloud side determines the violation category of the uploaded logistics operation video to be category A), then the second matching result is a match. If the violation classification result generated by the edge side is different from the reference violation category generated by the cloud-based multimodal large model (e.g., the edge side determines the violation category of the logistics operation video to be category A, and the cloud side determines the violation category of the uploaded logistics operation video to be category B), then the second matching result is a mismatch. Thus, this embodiment of the application achieves dual verification at both the feature representation level and the final result level.

[0060] In step S304 of some embodiments, when at least one of the first matching result and the second matching result indicates a mismatch, a violation operation review result is generated indicating that at least one of the video screening result and the violation operation classification result has failed the review. If the reference operation video features generated by the cloud multimodal large model are not similar to the violation and abnormal video features generated by the edge side, or if the violation operation classification result generated by the cloud multimodal large model is different from the violation operation classification result generated by the edge side, the corresponding logistics operation video is marked as failing the review. Specifically, the first matching result (the matching result between the features generated by the edge side and the features generated by the cloud) is mainly used to verify the validity of the video screening result and the video features extracted based on the video screening result, that is, it can be used to verify the accuracy of the video screening model in screening violation and abnormal segments and the accuracy of the edge multimodal large model in generating violation and abnormal video features. The second matching result (violation operation category matching) is mainly used to verify the accuracy of the violation operation classification result, that is, it can be used to verify the accuracy of the video screening model in screening violation and abnormal segments, the accuracy of the edge multimodal large model in generating violation and abnormal video features, and the accuracy of the operation classification model in generating violation operation classification results.

[0061] If the first matching result is a mismatch, it indicates that the violation and abnormal video features extracted by the edge end and the reference operation video features extracted by the cloud differ significantly in semantic space. This confirms that the video screening result review failed or the violation and abnormal video feature review generated by the edge multimodal large model failed. In this case, even if the edge end ultimately outputs the correct violation operation classification result, the reasoning process is unreliable because the violation operation classification result is generated based on incorrect features.

[0062] If the second matching result is a mismatch, it means that the text label corresponding to the violation operation category output by the edge side is inconsistent with the ground truth label generated by the cloud. For example, if the classification result output by the edge side is "violation throwing" and the classification result output by the cloud side is "normal sorting", it indicates that the operation classification model has made a qualitative error in the logistics operation. This may be due to the inaccuracy of the violation and abnormal video features generated by the edge multimodal large model, the inaccuracy of the violation operation classification result generated by the operation classification model itself, or the inaccuracy of the video screening result obtained by the video screening model.

[0063] Therefore, by combining the first and second matching results, it can be determined that the model on the edge side has an inaccurate problem in identifying logistics violations and cannot pass the review of the cloud-based multimodal large model. In this case, the violation review result output by the cloud-based multimodal large model indicates that at least one of the video screening result and the violation classification result has failed the review. If both the first and second matching results pass the review, then the violation review result output by the cloud-based multimodal large model indicates that the review has passed.

[0064] It should be noted that when the confidence level of the reference violation operation category generated by the cloud-based multimodal large model is low or the uncertainty is high (for example, the probability of each logistics violation operation is not much different, and the cloud-based multimodal large model cannot determine the specific violation operation category contained in the logistics operation video), the video screening results and the logistics operation videos corresponding to the violation operation classification results can be sent to the abnormal sample library for manual review.

[0065] This application's embodiments introduce reference operation video features and reference violation operation categories generated by a large cloud-based model as dual references. This not only detects explicit errors in edge-side judgment results (i.e., label inconsistency) but also uncovers implicit errors in edge-side judgment results based on incorrect features (i.e., feature inconsistency). This ensures the accuracy of cloud-based review results, thereby accurately locating the bottleneck data (i.e., difficult samples, uncertain samples, or long-tailed samples) that cause insufficient processing capabilities of the edge-side model from large-scale logistics operation videos. This provides high-quality supervision samples for subsequent online adaptive parameter fine-tuning of the edge-side model, solving the problems of non-targeted model iteration, low training efficiency, and inability to adapt to specific logistics operation scenarios in related technologies.

[0066] In some embodiments, referring to Figure 4, after the cloud server generates the review result of the violation operation, the cloud server is further configured to perform the following steps S401 to S403: Step S401, the reference operation video features and the corresponding reference violation operation categories are identified as difficult-to-handle sample data, and the difficult-to-handle sample data is stored in a preset cloud sample database; Step S402, the difficult-to-handle sample data is sent to the edge server, so that the edge server updates the parameters of the edge multimodal large model and the operation classification model corresponding to the target logistics operation scenario based on the difficult-to-handle sample data; Step S403, when the difficult-to-handle sample data in the cloud sample database reaches the preset sample quantity threshold, the parameters of the cloud multimodal large model are updated according to the difficult-to-handle sample data in the cloud sample database.

[0067] The violation review results obtained through the aforementioned steps actually reveal the cognitive blind spots or capability shortcomings of the edge-side model in the target logistics operation scenario. To prevent similar errors from recurring, the logistics operation videos that incorrectly identify violations can be converted into training sample data. By constructing training sample data, targeted parameter updates can be performed on the edge-side model, and the parameters of the cloud-based multimodal large model can also be updated using the constructed training sample data, thus achieving a decision-making closed loop for the cloud-based review results.

[0068] In step S401 of some embodiments, the reference operation video features and the corresponding reference violation operation categories can be identified as difficult-to-handle sample data, and the difficult-to-handle sample data can be stored in a preset cloud sample database. The difficult-to-handle sample data can include uncertain samples, long-tail samples, and negative samples. Uncertain samples represent sample data for which the model cannot determine the violation operation category. Long-tail samples represent sample data corresponding to violation operation categories that the model has rarely seen or has never seen in the target logistics operation scenario. Negative samples represent sample data for which the model incorrectly identifies the violation operation category, such as when the aforementioned reference operation video features are inconsistent with the violation and abnormal video features, or when the reference violation operation category is inconsistent with the violation operation classification result. Specifically, the reference operation video features, reference violation operation categories, and corresponding logistics operation videos generated by the cloud multimodal large model can be associated, and metadata tags corresponding to the difficult-to-handle sample data can be added. Scene identifiers for the corresponding target logistics operation scenario can also be added. It should be noted that, as described above, manual review can be used, and the results of manual review can also generate corresponding difficult-to-handle sample data. Furthermore, the data for difficult-to-handle sample data can be stored in a pre-defined cloud sample database. This cloud sample database can include training sample data previously used to train the cloud-based multimodal large model. Difficult-to-handle sample data is generated based on the violation review results generated by the cloud, and then the cloud sample database is incrementally updated using this data. For example, when the cloud confirms that the edge detection device mistakenly identifies an employee holding a black barcode scanner as smoking illegally, the correct features and category label corresponding to that logistics operation video can be stored as a difficult-to-handle sample record in the cloud sample database. It should be noted that, as mentioned above, the logistics operation videos corresponding to the video screening results generated by the video screening model are also uploaded to the cloud server. In addition to logistics operation videos that fail the review, corresponding training sample data can be generated based on the logistics operation videos corresponding to the video screening results and stored in the cloud sample database, thus expanding the cloud sample database. This training sample data can be identified as general training sample data and is not specific to any particular logistics operation scenario.

[0069] In step S402 of some embodiments, the heavy-duty sample data can then be sent to the edge server so that the edge server can update the parameters of the edge multimodal large model and the operation classification model corresponding to the target logistics operation scenario based on the heavy-duty sample data. Specifically, the heavy-duty sample data can be sent to the edge server in the corresponding target logistics site in real time so that the edge server can fine-tune the parameters of the edge multimodal large model and the operation classification model corresponding to the target logistics operation scenario based on the heavy-duty sample data. Specifically, parameter fine-tuning can be performed through a lightweight incremental layer, which can be achieved through an adapter module, low-rank adaptation (LoRA), and batch normalization. This is achieved through methods such as Normalization (BN) and Test-Time Augmentation (TTN). The adapter module is a lightweight neural network module inserted into the pre-trained model for task-specific fine-tuning, enabling transfer learning without updating most parameters of the original model. Low-rank adaptation is an efficient parameter fine-tuning method that reduces computational and storage overhead by training only a small number of additional parameters through low-rank decomposition of the weight matrix. Batch normalization is a technique to accelerate neural network training by standardizing each batch of data, mitigating the vanishing gradient problem and improving training stability. Test-time augmentation involves performing various augmentations on the input data during the model inference phase (such as flipping, scaling, and cropping), and then fusing multiple prediction results to improve the model's robustness and generalization ability. For large-scale edge multimodal models, updating the integrated lightweight adapter allows for better differentiation of confused targets in target logistics operation scenarios during feature extraction. For operation classification models, which are typically small-scale classifiers, parameters can be directly updated using hard-to-bear sample data delivered from the cloud to adjust their classification decision logic.

[0070] In step S403 of some embodiments, when the number of difficult-to-bear sample data in the cloud sample database reaches a preset sample number threshold, the parameters of the cloud-based multimodal large model can be updated based on the difficult-to-bear sample data in the cloud sample database. Since the cloud-based multimodal large model is a base model with a huge parameter scale, updating it once takes a long time, and the purpose of training the cloud-based multimodal large model is to improve its generalization ability. Therefore, the cloud-based multimodal large model can be trained only when the number of difficult-to-bear sample data in the cloud sample database increases to a certain amount. Specifically, when the number of difficult-to-bear sample data in the cloud sample database reaches a preset sample number threshold, the parameters of the cloud-based multimodal large model are updated based on the difficult-to-bear sample data in the cloud sample database. For example, if the preset sample number threshold is 10,000 difficult-to-bear sample data, then when the number of difficult-to-bear sample data in the cloud sample database reaches 10,000, the cloud-based multimodal large model is trained uniformly. Furthermore, an update frequency can be set to periodically update the parameters of the cloud-based multimodal large model, for example, updating the parameters of the cloud-based multimodal large model every month using training sample data or incremental sample data from the cloud sample database. This ensures the stability of model training while improving the utilization of computing resources.

[0071] This application embodiment realizes the decision-making closed loop of cloud-based violation operation review results and adaptive updates of the edge-side model. By feeding back the difficult-to-handle sample data to the edge side and training the edge-side model with a small number of incremental samples, the edge-side model is adaptively updated for the target logistics operation scenario. This enables the edge-side model to flexibly cope with environmental and element differences in different logistics operation scenarios, solving the problem of low accuracy of single-model cross-scenario identification of logistics violations in related technologies. Furthermore, through the large-scale accumulation and training of cloud-based training sample data, the entire system gradually improves the accuracy of identifying logistics violations over time.

[0072] In some embodiments, referring to Figure 5, the cloud server is further configured to perform the following steps S501 to S503: Step S501, acquiring meteorological data and violation events at the logistics site, and generating risk violation operation categories corresponding to the relevant logistics site based on the meteorological data and violation events at the logistics site; Step S502, when the relevant logistics site includes the target logistics site, generating a violation operation warning strategy corresponding to the target logistics site based on the risk violation operation category; Step S503, distributing the violation operation warning strategy to the edge server in the target logistics site, so that the edge server adjusts the logistics violation operation detection strategy for the target logistics site based on the violation operation warning strategy.

[0073] Operational safety at target logistics sites is often greatly affected by external environmental factors (such as severe weather) and sudden safety incidents. In order to further improve the sensitivity of identifying logistics violations and the ability to issue timely warnings of logistics violations, this application embodiment can actively sense environmental changes and safety incidents occurring at the target logistics site through a cloud server, thereby actively pushing warning instructions to the edge server so that it can adjust the detection focus in advance to adapt to the complex and ever-changing logistics site environment.

[0074] In step S501 of some embodiments, the cloud server can acquire meteorological data and violation events at the logistics site, and generate risk violation operation categories corresponding to the relevant logistics site based on the meteorological data and violation events. Meteorological data refers to natural environmental parameters affecting the operational safety of the logistics site, including rainfall, wind speed, and visibility. The cloud server can obtain real-time weather conditions for the target logistics site's location by calling the API interface of a third-party meteorological service provider. Violation events at the logistics site can include sudden safety violations obtained through instant messaging (IM), work order systems, or safety broadcasts, such as a cargo collapse accident at a nearby logistics site. The cloud server can receive violation event notifications by subscribing to messages from the enterprise's internal safety management system. Then, the cloud server can perform correlation analysis based on the meteorological data and violation events at the logistics site to deduce possible logistics violations under the current environment, thus obtaining risk violation operation categories. For example, meteorological data from a logistics site indicates that the area will experience heavy rain. The cloud server infers that the heavy rain may cause slippery ground, increasing the risk of falls for personnel, reducing visibility, increasing vehicle speeds and accident risks, and increasing the probability of personnel rushing to seek shelter and illegally crossing work areas. Then, personnel falls, vehicle skidding, and personnel illegally crossing work areas can be categorized as risky violations, and corresponding abnormal monitoring and early warning strategies can be generated based on these categories. The cloud server can also identify relevant logistics sites affected by the above factors based on the meteorological data and violations. It can determine the affected logistics sites based on the geographical location covered by the meteorological data. For example, when the cloud server receives a red rainstorm warning for a certain district in a city, it can determine that all logistics transit points in that area will be affected by the rainstorm. Furthermore, logistics site violations often target sites of the same type with certain attributes. For instance, if the cloud server receives a safety notification stating that a fire safety inspection of a cold chain warehouse is required, this instruction applies to all logistics sites handling cold chain goods.

[0075] In step S502 of some embodiments, after the cloud server determines the relevant logistics sites affected by meteorological data or violations at the logistics sites through the aforementioned steps, it then determines whether the target logistics site belongs to the affected relevant logistics sites. When the relevant logistics sites include the target logistics site, a violation warning strategy corresponding to the target logistics site can be generated based on the risk violation operation category. Specifically, the violation warning strategy is used to send warning information to the edge server and guide the edge-side devices and models to adjust their strategies for identifying physical violations. Specifically, the recognition focus of the edge-side model can be adjusted, or the video acquisition strategy of the edge-side video acquisition device can be adjusted. For example, if the risk violation operation category is high vehicle speed, the generated violation warning strategy may include increasing the priority of the operation classification model corresponding to speeding and reducing the trigger threshold for speeding behavior. Or, if the risk violation operation category is reduced visibility, the generated violation warning strategy may include increasing the sampling frequency of cameras in the relevant area and strengthening the detection weight of areas prone to water accumulation and slipping, with cameras automatically turning on wipers or adjusting parameters to improve image quality.

[0076] In step S503 of some embodiments, the violation operation warning strategy is sent to the edge server in the target logistics area, so that the edge server adjusts its logistics violation operation detection strategy for the target logistics area based on the violation operation warning strategy. As exemplified above, the edge server can increase the sampling frequency of video acquisition devices in slippery areas such as unloading areas where there is a risk of people falling.

[0077] This application embodiment proactively senses abnormal environmental changes through a cloud server, improving the initiative and adaptability of logistics safety detection. This embodiment can dynamically adjust the detection strategy for logistics violations at the edge based on weather changes or safety incidents, solving the problems of high false alarm rates, high missed alarm rates, and untimely warnings in related technologies' logistics safety systems under harsh environments. This improves the logistics safety system's ability to perceive abnormal environments, enabling intelligent adaptation to diverse logistics operation scenarios and environmental changes, thereby enhancing the safety control level of logistics sites.

[0078] In step S103 of some embodiments, when the violation operation review result indicates that at least one of the video screening result and the violation operation classification result fails the review, the edge server can obtain the reference violation operation category generated by the cloud server based on the logistics operation video that failed the review, as well as the reference operation video features extracted from the logistics operation video that failed the review. Specifically, as described above, the cloud server can generate hard-to-bear sample data based on the reference operation video features and the reference violation operation category, and send the hard-to-bear sample data to the edge server. In this way, the embodiments of this application can avoid the problem of high bandwidth consumption and slow transmission speed caused by sending a large amount of training sample data to the edge server, thereby enabling incremental updates of the model on the edge side with a small amount of training sample data.

[0079] In step S104 of some embodiments, after receiving the reference operation video features and reference violation operation categories sent by the cloud server, the edge server can update the parameters of the operation classification model corresponding to the target logistics operation scenario based on the reference operation video features and reference violation operation categories. Specifically, the edge server can perform online fine-tuning or incremental learning of the operation classification model corresponding to the target logistics operation scenario based on the reference operation video features and reference violation operation categories. In this way, the operation classification model on the edge side can be trained specifically for the target logistics operation scenario, and can also be trained using historical sample data in the cloud sample database, quickly adapting to changes in the logistics operation scenario without forgetting historical knowledge. This update is continuous, and as the running time progresses, the operation classification model on the edge side can become more adapted to the target logistics operation scenario, solving the problems of poor generalization and long model iteration cycles across logistics operation scenarios in related technologies, thereby achieving online adaptive adaptation for multiple logistics operation scenarios.

[0080] In some embodiments, please refer to Figure 6, which illustrates the system architecture and data flow diagram of an edge-cloud collaborative logistics security system. The logistics security system includes video acquisition equipment, an edge server, a cloud server, and a cloud sample database. The edge server is deployed in the target logistics site and contains adaptive classification units. These adaptive classification units include an edge multimodal large model and operation classification models corresponding to different logistics violations in each target logistics operation scenario. The edge server also deploys a video filtering model, which includes a target detection module, a target attribute recognition module, and a target tracking module. The cloud server deploys a cloud multimodal large model and a cloud sample database. Specifically, the video acquisition equipment is used to acquire real-time logistics operation videos in the target logistics operation scenario. The acquired logistics operation videos can be transmitted to the edge server for violation classification processing or sent to the cloud sample database for storage, so as to generate corresponding training sample data based on the logistics operation videos later.

[0081] After the collected logistics operation videos are transmitted to the edge server, the following two data processing streams can be executed in parallel: The edge server uses a video filtering model to filter the logistics operation videos, obtaining the filtering results. Specifically, a lightweight object detection module identifies whether there are targets such as people, vehicles, and objects in each video frame; then, the object attribute recognition module further extracts the relevant attribute information of the detected targets; finally, the object tracking module analyzes the movement trajectory of the detected targets. It should be noted that the video filtering model can use only the object detection module. Some business scenarios may not require attribute detection and object tracking, and these two steps can be omitted. However, the core of the video filtering model is to obtain illegal and abnormal segments through coarse screening, improve the recall rate of illegal and abnormal segments, and reduce bandwidth costs. In this way, by using the video filtering model to perform initial screening of logistics operation videos, meaningless background images can be removed, and video segments containing potentially abnormal behaviors can be filtered out.

[0082] The video screening results generated in the aforementioned steps enter the adaptive classification unit, or can be directly uploaded to the cloud server. Uploading large-scale logistics operation videos to the cloud server after screening significantly saves transmission bandwidth. The adaptive classification unit can extract features from the logistics operation videos in the video screening results using a lightweight edge multimodal model. Based on the extracted features, a lightweight operation classification model performs specific violation operation classification, generating violation operation classification results. The generated violation and abnormal video features and the corresponding violation operation classification results are sent to the cloud sample database for storage. The edge server uploads the processed video screening results and the logistics operation videos corresponding to the violation operation classification results to the cloud server, while simultaneously sending the generated violation and abnormal video features and the corresponding violation operation classification results to the cloud server for further processing.

[0083] The cloud server can use a cloud-based multimodal large model to review logistics operation videos uploaded by the edge server for violation checks. If the cloud server determines that the edge server's violation check is correct, it confirms the violation and issues a logistics violation alert to the edge server. If the cloud server determines that the edge server's check is incorrect (e.g., a false alarm), it generates difficult-to-handle sample data and sends it to the cloud sample database for storage. The cloud server can also periodically retrieve training sample data from the cloud sample database and update the parameters of the cloud-based multimodal large model based on the training sample data. The cloud server can also distribute difficult-to-handle sample data to the edge server. After receiving the difficult-to-handle sample data, the edge server fine-tunes the parameters of the model in the adaptive classification unit, enabling online adaptive updates of the edge-side model to the target logistics operation scenario.

[0084] The cloud server can also proactively detect anomalies. Specifically, it can generate early warning strategies for violations by combining external meteorological data or security incidents, and then distribute these strategies to edge servers. This allows the edge servers to adjust their logistics violation detection strategies accordingly. The cloud server can also manage video capture devices, specifically by detecting their online status and image quality, and verifying whether the device ID matches the corresponding logistics operation scenario ID, thereby improving the accuracy of logistics operation video capture.

[0085] Steps S101 to S104, as illustrated in this embodiment, involve using an edge server's video screening model to coarsely screen a large number of logistics operation videos, and then using an operation classification model corresponding to the logistics operation scenario to finely classify the violation and abnormal segments. The video screening results and violation classification results are then uploaded to a cloud server for violation verification to further ensure the accuracy of the classification results. If the verification fails, the edge server can obtain the features of a reference operation video and a reference violation category generated by the cloud server to update the parameters of the operation classification model corresponding to the target logistics operation scenario. This allows the edge model to learn online and correct its own cognitive biases, continuously adapting to changes in specific logistics operation scenarios and achieving continuous self-optimization of model performance. This improves the accuracy of the edge model in identifying logistics violations in different logistics operation scenarios.

[0086] Referring to Figure 7, this application embodiment also provides a logistics operation data processing device 700, applied to an edge server corresponding to a target logistics site. The target logistics site includes multiple target logistics operation scenarios. The edge server is equipped with a trained video filtering model and an operation classification model corresponding to different logistics violations in each target logistics operation scenario. This enables the implementation of the aforementioned logistics operation data processing method. The device includes: a classification unit 710, used to filter violation segments from logistics operation videos collected in the target logistics operation scenarios based on the video filtering model to obtain video filtering results, and to classify the violation and abnormal segments indicated by the video filtering results based on the operation classification model corresponding to the target logistics operation scenario, thereby obtaining a classification of violations. The system includes: a receiving unit 720, which sends the video screening results and the logistics operation videos corresponding to the violation operation classification results to the cloud server for violation operation review, and receives the violation operation review results returned by the cloud server; an obtaining unit 730, which obtains the reference violation operation category generated by the cloud server based on the logistics operation videos that failed the review, and the reference operation video features extracted from the logistics operation videos that failed the review, when the violation operation review results indicate that at least one of the video screening results and the violation operation classification results has failed the review; and an updating unit 740, which updates the parameters of the operation classification model corresponding to the target logistics operation scenario based on the reference operation video features and the reference violation operation category.

[0087] The specific implementation of this logistics operation data processing device is basically the same as the specific embodiment of the logistics operation data processing method described above, and will not be repeated here.

[0088] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described logistics operation data processing method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0089] Please refer to Figure 8, which illustrates the hardware structure of an electronic device according to another embodiment. The electronic device includes: a processor 801, which can be implemented using a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, for executing related programs to implement the technical solutions provided in the embodiments of this application; and a memory 802, which can be implemented using a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM), etc. The memory 802 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 802 and is called and executed by the processor 801 to execute the logistics operation data processing method of the embodiments of this application. The input / output interface 803 is used to realize information input and output. The communication interface 804 is used to realize communication interaction between this device and other devices. Communication can be realized by wired means (such as USB, network cable, etc.) or by wireless means (such as mobile network, WIFI, Bluetooth, etc.). The bus 805 transmits information between the various components of the device (such as processor 801, memory 802, input / output interface 803 and communication interface 804). The processor 801, memory 802, input / output interface 803 and communication interface 804 realize communication connection between each other within the device through the bus 805.

[0090] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described logistics operation data processing method.

[0091] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0092] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0093] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0094] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0095] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0096] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0097] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0098] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0099] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0100] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0101] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0102] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method for processing logistics operation data, characterized in that, An edge server corresponding to a target logistics site, the target logistics site including multiple target logistics operation scenarios, is deployed on the edge server. The edge server is equipped with a trained video filtering model and an operation classification model corresponding to different logistics violations in each target logistics operation scenario. The method includes: filtering violation segments from logistics operation videos collected in the target logistics operation scenarios based on the video filtering model to obtain video filtering results; classifying the violation anomaly segments indicated by the video filtering results into violation operation categories based on the operation classification model corresponding to the target logistics operation scenario to obtain corresponding violation operation classification results; sending the video filtering results and the logistics operation videos corresponding to the violation operation classification results to a cloud server for violation operation review, and receiving the violation operation review results returned by the cloud server; when the violation operation review results indicate that at least one of the video filtering results and the violation operation classification results fails the review, obtaining a reference violation operation category generated by the cloud server based on the logistics operation videos that failed the review, and reference operation video features extracted from the logistics operation videos that failed the review; and updating the parameters of the operation classification model corresponding to the target logistics operation scenario based on the reference operation video features and the reference violation operation category.

2. The method according to claim 1, characterized in that, The edge server is also deployed with an edge multimodal large model. The step of classifying the illegal and abnormal segments indicated by the video screening results based on the operation classification model corresponding to the target logistics operation scenario to obtain the corresponding illegal operation classification results includes: extracting features from the illegal and abnormal segments indicated by the video screening results based on the edge multimodal large model to obtain illegal and abnormal video features; and classifying the illegal and abnormal video features based on the operation classification model corresponding to the target logistics operation scenario to obtain the corresponding illegal operation classification results.

3. The method according to claim 2, characterized in that, The cloud server is equipped with a cloud-based multimodal large model. The edge multimodal large model is obtained by knowledge distillation from the cloud-based multimodal large model. The process of the cloud server generating the review result of the violation operation includes the following steps: based on the cloud-based multimodal large model, feature extraction is performed on the logistics operation videos corresponding to the video screening results and the violation operation classification results to obtain reference operation video features. Based on the cloud-based multimodal big model, the reference violation category corresponding to each logistics operation video is generated using the video filtering results and the violation operation classification results. The reference operation video features are matched with the violation and abnormal video features to obtain a first matching result, and the reference violation operation category is matched with the violation operation classification result to obtain a second matching result; When at least one of the first matching result and the second matching result indicates a mismatch, a violation review result is generated indicating that at least one of the video filtering result and the violation classification result has failed the review.

4. The method according to claim 3, characterized in that, After the cloud server generates the review result of the violation operation, the cloud server is further configured to perform the following steps: identify the reference operation video features and the corresponding reference violation operation category as difficult-to-handle sample data, and store the difficult-to-handle sample data in a preset cloud sample database; send the difficult-to-handle sample data to the edge server, so that the edge server updates the parameters of the edge multimodal large model and the operation classification model corresponding to the target logistics operation scenario based on the difficult-to-handle sample data; when the difficult-to-handle sample data in the cloud sample database reaches a preset sample quantity threshold, update the parameters of the cloud multimodal large model according to the difficult-to-handle sample data in the cloud sample database.

5. The method according to claim 2, characterized in that, The cloud server is also used to perform the following steps: acquiring meteorological data and violation events at logistics sites, and generating risk violation operation categories corresponding to relevant logistics sites based on the meteorological data and violation events at the logistics sites; when the relevant logistics sites include a target logistics site, generating a violation operation early warning strategy corresponding to the target logistics site based on the risk violation operation category; and distributing the violation operation early warning strategy to the edge server in the target logistics site, so that the edge server adjusts its logistics violation operation detection strategy for the target logistics site based on the violation operation early warning strategy.

6. The method according to claim 1, characterized in that, The step of sending the candidate logistics operation videos corresponding to the video screening results and the violation operation classification results to the cloud server for violation operation review includes: obtaining the model classification performance parameters corresponding to the violation operation classification results within a preset time interval; when at least one of the model classification performance parameters indicates that the classification performance of the operation classification model has decreased, sending the logistics operation videos corresponding to the violation operation classification results within the preset time interval to the cloud server for violation operation review, and sending the logistics operation videos corresponding to the video screening results to the cloud server for violation operation review.

7. The method according to claim 1, characterized in that, The video filtering model includes a target detection module, a target attribute recognition module, and a target tracking module. The step of filtering violation segments from logistics operation videos collected in a target logistics operation scenario based on the video filtering model to obtain video filtering results includes: calling the target detection module to perform target detection on the logistics operation videos collected in the target logistics operation scenario to obtain target detection results; calling the target attribute recognition module to perform attribute recognition on the target detection results to obtain target attribute recognition results; calling the target tracking module to perform video tracking on the detected objects contained in the target detection results to obtain target object tracking videos; and filtering violation segments based on the target object tracking videos to obtain video filtering results.

8. A logistics operation data processing device, characterized in that, An edge server corresponding to a target logistics site, the target logistics site including multiple target logistics operation scenarios, is applied to the edge server. The edge server is equipped with a trained video filtering model and an operation classification model corresponding to different logistics violations in each target logistics operation scenario. The device includes: a classification unit, used to filter violation segments from logistics operation videos collected in the target logistics operation scenarios based on the video filtering model to obtain video filtering results, and to classify the violation and abnormal segments indicated by the video filtering results based on the operation classification model corresponding to the target logistics operation scenario to obtain corresponding violation operation classification results; and a receiving unit, used to receive the video filtering results. The system sends the results and the corresponding logistics operation videos of the violation operation classification results to the cloud server for violation operation review, and receives the violation operation review results returned by the cloud server; the acquisition unit is used to acquire, when the violation operation review result indicates that at least one of the video screening results and the violation operation classification results has failed the review, a reference violation operation category generated by the cloud server based on the logistics operation video that failed the review, and reference operation video features extracted from the logistics operation video that failed the review; the update unit is used to update the parameters of the operation classification model corresponding to the target logistics operation scenario according to the reference operation video features and the reference violation operation category.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the logistics operation data processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the logistics operation data processing method according to any one of claims 1 to 7.