A method for preventing cheating of a truck scale based on edge computing

CN122842008APending Publication Date: 2026-09-29NANJING UNIV OF SCI & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610827743.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]本发明提出了一种基于边缘计算的汽车衡智能监控与防作弊方法,旨在解决现有技术中车辆停稳判定可靠性低、复杂环境下车辆与人员检测精度不足、人员作弊行为难以精准识别、边缘节点计算资源受限难以实时部署,以及缺乏自动化闭环防作弊监管体系的技术问题

Benefits of technology

[0020](1)本发明针对汽车衡夜间低光照、车辆遮挡、小目标检测的复杂工业场景,构建了专用目标图像数据集,包括不同光照条件、车辆角度及背景,并对图像进行组合增强处理,使边缘节点部署的改进RTDETRV2模型,骨干网络替换为RepViT,颈部网络替换为Lite结构,能够有效提取车辆及人员多特征信息。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122842008A_ABST
    Figure CN122842008A_ABST
Patent Text Reader

Abstract

This invention discloses a method for preventing cheating on truck scales based on edge computing, comprising: acquiring vehicle weighing videos, constructing a dedicated image dataset and performing combined enhancements; training an improved RTDETRV2 model using RepViT and Lite and deploying it on edge nodes to achieve real-time detection of vehicles and personnel; using BoTSORT to track and assign unique IDs; calculating the offset of the bounding box center point based on the vehicle trajectory; generating a visual stabilization signal through moving average; simultaneously analyzing the weight fluctuation of the weighbridge; confirming vehicle stabilization when two conditions are met; calibrating the ROI mask of the weighing area; extracting the bottom coordinates of the personnel detection box; determining whether the person has entered the area; verifying cheating through multi-frame verification and triggering an alarm; the edge nodes realize a closed loop of detection, tracking, stabilization judgment, anti-cheating, alarm, and data recording; and pushing structured data to the central control center. Compared with existing technologies, this invention has high detection accuracy and reliable stabilization judgment in complex environments such as low light and occlusion, realizing automated and traceable management of truck scale anti-cheating.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of truck scale measurement and video monitoring technology, specifically relating to an intelligent monitoring and anti-cheating method for truck scales based on edge computing. Background Technology

[0002] Truck scales are crucial weighing devices used for weighing large vehicles, widely applied in mining, ports, and logistics. Existing truck scale monitoring systems typically employ a distributed architecture, including front-end camera equipment, edge computing nodes, and a remote monitoring center. The edge nodes, equipped with built-in NPU chips, can deploy target detection models, perform real-time inference on the video stream, and upload the processed structured data to the monitoring center, forming a distributed processing model of "front-end acquisition—edge inference—central management."

[0003] Existing technologies for vehicle stability determination and anti-cheating monitoring suffer from low efficiency and insufficient accuracy. Vehicle stability determination often relies on manual visual inspection, which is highly subjective and inefficient, or solely on weight sensor data, judging stability by detecting weight fluctuations. However, this is prone to misjudgment when the weighbridge is subjected to vehicle braking impacts or sensor noise interference. Meanwhile, anti-cheating monitoring mainly relies on manual review of surveillance footage or simple area intrusion detection, making it difficult to accurately determine whether personnel have actually entered the weighing area, nor can it distinguish continuous behavior from different individuals. Some methods acquire multiple consecutive frames of surveillance images and use convolutional neural networks, attention mechanisms, or autoencoders to reconstruct features for abnormal behavior identification. However, these methods have limitations in decoder reconstruction capabilities, continuous behavior differentiation, and collaborative analysis with weight sensor data. This leads to potential missed or false alarms in complex scenarios such as low light, occlusion, and small targets. Furthermore, the computational cost of these models is high, making it difficult to meet the requirements of high reliability, real-time performance, and automated monitoring in industrial settings. In summary, existing technologies for vehicle stability determination and anti-cheating monitoring still suffer from technical deficiencies such as reliance on single data sources, insufficient accuracy, poor real-time performance, and difficulty in achieving highly reliable automated monitoring. Therefore, how to improve the algorithm model and data processing logic on the edge inference node based on the existing hardware and network architecture to realize vehicle stability determination and automated anti-cheating monitoring in complex environments has become a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0004] This invention proposes an intelligent monitoring and anti-cheating method for truck scales based on edge computing, aiming to solve the technical problems in existing technologies such as low reliability of vehicle stability determination, insufficient vehicle and personnel detection accuracy in complex environments, difficulty in accurately identifying personnel cheating behavior, limited computing resources of edge nodes making real-time deployment difficult, and lack of an automated closed-loop anti-cheating supervision system.

[0005] The technical solution for implementing this invention is: a method for preventing cheating on truck scales based on edge computing, comprising the following steps:

[0006] Step S1: Collect video images of vehicles weighing in the weighbridge operation area, construct a dedicated target image dataset covering different lighting conditions, angles, and vehicle states, and perform combined image enhancement processing on the dedicated target image dataset to obtain a target enhancement dataset, then proceed to step S2.

[0007] Step S2: Construct an improved RTDETRV2 model to output the bounding boxes and confidence information of vehicle and personnel targets.

[0008] The improved RTDETRV2 network includes a Patch Embedding module, a RepViT backbone network, a Lite module, and an RTDETRTransformerv2 decoder.

[0009] The input image is an RGB three-channel image. The Patch Embedding module performs feature mapping on the input image, downsampling through a spatial convolution operation to generate a preliminary feature map with a preset number of channels. An activation function is then applied for a non-linear transformation to obtain a smaller, expanded preliminary feature map. This preliminary feature map is then sequentially input into the RepViT backbone network for hierarchical feature extraction, outputting multiple three-channel feature maps P with different downsampling ratios. 3e P 4e and P 5e HybridEncoder performs cross-scale feature encoding on these three feature maps to obtain the encoded three feature maps P3, P4 and P5.

[0010] The Lite module performs multi-scale fusion and enhancement processing on the encoded three feature maps P3, P4, and P5, generating corresponding enhanced feature maps. ;

[0011] Proceed to step S3.

[0012] Step S3: Train the improved RTDETRV2 network using the target augmentation dataset to obtain the improved RTDETRV2 model. Deploy the improved RTDETRV2 model at edge nodes and pull the video stream in real time through the camera. Let the input image tensor of the i-th frame be... , , ,in For batch size, For image resolution, For the total number of frames, Input the improved RTDETRV2 model It extracts multi-scale features of vehicles and people, and generates corresponding target detection results, namely the continuous target trajectory of vehicles and the continuous target trajectory of people.

[0013] Step S4: For the same vehicle target, extract the vehicle detection box sequence of consecutive frames from the vehicle detection results obtained in step S3, calculate the vehicle pixel offset based on the position change of the center point of the vehicle detection box in consecutive frames, and smooth the vehicle pixel offset. When the smoothed vehicle pixel offset is lower than the first judgment threshold in multiple consecutive frames and the duration reaches the first preset time, generate a vehicle visual stabilization judgment signal and proceed to step S5.

[0014] Step S5: After the vehicle visual stopping determination signal is triggered, the weight data of the secondary meter of the truck scale is read synchronously. The weight fluctuation amplitude is calculated within the preset stopping determination time window. When the weight fluctuation amplitude is lower than the second determination threshold and the duration reaches the second preset time, the vehicle is confirmed to have entered the stopping weighing state, and the process proceeds to step S6. If the weight fluctuation amplitude does not meet the second determination threshold, the process proceeds to step S7.

[0015] Step S6: Manually mark the weighing area of ​​the truck scale in the image and generate the corresponding pixel-level mask. After confirming that the vehicle has entered the stopped weighing state, the personnel get off the vehicle. Use the improved RTDETRv2 model from step S3 to detect the personnel in real time, obtain the personnel detection box, extract the pixel coordinates of the two endpoints of the bottom edge of the detection box, and determine whether the two endpoints of the bottom edge are located within the marked area mask. If the same personnel target has any one of the two endpoints of the bottom edge of its detection box located within the area mask in N consecutive frames, it is determined that the personnel have entered the weighing area and constitute cheating behavior, and proceed to step S8; otherwise, proceed to step S9. .

[0016] Step S7: When the vehicle visual stability determination signal is valid and the weight fluctuation amplitude is not lower than the second determination threshold, call the video image of the corresponding time period, extract the video key frame of the corresponding time period, and determine whether there is weighbridge swaying or weight sensor interference based on the key frame image. When it is determined to be interference, ignore the weight abnormality and go to step S9; when it is determined not to be interference, return to step S5 and re-determine stability.

[0017] Step S8: After determining that it is a cheating behavior, the edge node executes an abnormal alarm and saves the video clips, vehicle stopping determination results and weight data of the corresponding time period to form a traceable event evidence chain, and then proceeds to step S9.

[0018] Step S9: The edge node pushes the vehicle detection results, target tracking results, vehicle stability determination results, anti-cheating determination results, and event recording results to the monitoring center in the form of structured data.

[0019] Compared with the prior art, the significant advantages of this invention are:

[0020] (1) This invention addresses the complex industrial scenarios of low light, vehicle occlusion, and small target detection at night in the case of truck scales. It constructs a dedicated target image dataset, including different lighting conditions, vehicle angles, and backgrounds. The images are combined and enhanced, so that the improved RTDETRV2 model deployed at the edge nodes is replaced with RepViT in the backbone network and with a Lite structure in the neck network, which can effectively extract multiple feature information of vehicles and personnel.

[0021] (2) To address the issues of human reliance and cheating in vehicle stability determination, this invention proposes a visual and weight-based stability determination method: The center point offset is calculated by tracking the bounding box of the vehicle in consecutive frames using RTDETRV2, and combined with secondary table weight fluctuation analysis, to achieve accurate determination of the vehicle's stable state. Compared to relying solely on weight sensors for stability determination, this invention effectively reduces misjudgments caused by braking impacts, sensor noise, and other interferences, thus improving the reliability of stability determination.

[0022] (3) This invention addresses the issues of reliance on manual monitoring and simplistic rules in weighbridge anti-cheating methods by proposing a ROI-based method for determining personnel cheating: After a vehicle comes to a complete stop, the coordinates of the two ends of the bottom edge of the personnel detection frame are extracted. If any one of the two ends falls within the ROI for multiple consecutive frames, it is determined that the person is cheating on the weighbridge; otherwise, it is determined that the person is not on the weighbridge. Combined with video and weight data, automated alarms and event recording are achieved. Actual testing shows a high accuracy rate in identifying personnel cheating. Compared to infrared beam detection and simple area intrusion detection, the false alarm rate is significantly reduced, significantly improving the automation level of anti-cheating and avoiding human oversight.

[0023] (4) This invention is designed for complex situations where multiple vehicles and personnel are monitored simultaneously. It can independently calculate the stability and anti-cheating indicators of each vehicle and each person, and distinguish different targets by tracking IDs, ensuring accurate stability judgment and anti-cheating effects in multi-target situations. At the same time, it takes into account low-latency real-time processing to meet the automation needs of industrial sites.

[0024] (5) This invention constructs a real-time closed-loop processing flow at the edge node, including real-time detection, target tracking, fusion stability judgment, anti-cheating judgment, anomaly alarm, and event recording, and pushes the structured results to the monitoring center in real time. Compared with the lack of automated closed-loop in the prior art, this invention realizes automated, traceable, and highly reliable management of vehicle stability judgment and anti-cheating in industrial scenarios, improving the overall intelligence level and engineering feasibility of the system. Attached Figure Description

[0025] Figure 1 This is a flowchart of the method of the present invention.

[0026] Figure 2This is a schematic diagram illustrating the vehicle coming to a complete stop according to the present invention.

[0027] Figure 3 This is a real-time weight curve of the secondary table of the present invention.

[0028] Figure 4 This is a schematic diagram illustrating the anti-cheating determination of personnel in the ROI area according to the present invention.

[0029] Figure 5 This is a network structure diagram of the Lite module of the present invention. Detailed Implementation

[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0031] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indication will also change accordingly.

[0032] The technical solutions of the various embodiments of the present invention can be combined with each other, but only if they can be implemented by those skilled in the art. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.

[0033] The following section will further introduce the specific implementation method, as well as the technical difficulties and inventive points of this invention, using this design example as an example.

[0034] Combination Figure 1 The specific steps of the edge computing-based anti-cheating method for truck scales described in this invention are as follows:

[0035] Step S1: Collect vehicle weighing videos in the weighbridge operation area using cameras positioned on the four corner pillars of the weighbridge. In practice, due to the complex weighbridge environment, various reflection problems and motion blur problems caused by changes in lighting are common. Therefore, images are extracted from the video every 200ms to form a dataset of 3000 images.

[0036] To make the dataset cover a more comprehensive range of image states, the samples will undergo combined image enhancement processing. The combined image enhancement methods include adaptive brightness adjustment, color equalization, local contrast enhancement, and random geometric transformation.

[0037] The mathematical representation of image enhancement is as follows: For the original image After random rotation angle get satisfy:

[0038] ,

[0039] Scaled image for

[0040] ,

[0041] This refers to the coordinate position corresponding to the scaling. Brightness adjustment is achieved by multiplying by a coefficient. , This is the brightness multiplication factor, plus the bias. , As a brightness bias, comprehensive training samples are generated through the above transformation combination.

[0042] It should be noted that, since cheating is a low-probability event in actual operations, the number of positive samples collected is limited. To meet the requirements of deep learning models for massive training data, and to take into account the cost of data collection and annotation, various data augmentation operations are performed on the initial samples, and the original data and the augmented data are merged to form the target augmented dataset.

[0043] The images are labeled using LabelMe, with labels categorized into two types: vehicles and people. The target augmentation dataset is divided into training and testing sets, with a sample size ratio of 8:2. The model input image size is 640×640, and the output consists of L candidate boxes, where L is 300. Each candidate box corresponds to 80 COCO categories. Only the confidence scores and normalized bounding box coordinates for the vehicle and person categories are used. Proceed to step S2.

[0044] Step S2: The improved RTDETRv2 network includes a Patch Embedding module, a RepViT backbone network, a HybridEncoder module, a Lite module, and an RTDETRTransformerv2 decoder. This network is used to output the bounding boxes and confidence information of vehicle targets and personnel targets.

[0045] The input image is an RGB three-channel image. The Patch Embedding module performs feature mapping on the input image, that is, it performs a 7×7 spatial convolution operation on the input RGB image with a stride of 4 to achieve spatial downsampling and generate a preliminary feature map of 64 channels. Then, the GELU activation function is applied to the preliminary feature map of 64 channels to obtain a preliminary feature map of size 160×160 with 64 channels. The preliminary feature map is then input into the RepViT backbone network for hierarchical feature extraction, passing through four stages: Stage 1, Stage 2, Stage 3, and Stage 4. Each stage consists of several RepViTBlock modules stacked together. The output of the previous stage is used as the input of the next stage. As the network depth increases, the spatial size of the feature map gradually decreases, and the number of channels gradually increases. The shallow stages mainly preserve the edges of vehicles and people, wheels, etc. The deep layer primarily extracts target semantic information, including contour and texture information. Stage 1 contains two RepViTBlock modules with a stride of 1, outputting feature maps C3 / P3 with a downsampling factor of 8, 64 channels, and a spatial size of 160×160. Stage 2 contains three RepViTBlock modules with a stride of 2, outputting feature maps C4 / P4 with a downsampling factor of 16, 128 channels, and a spatial size of 80×80. Stage 3 contains four RepViTBlock modules with a stride of 2, outputting feature maps C5 / P5 with a downsampling factor of 32, 256 channels, and a spatial size of 40×40. Stage 4 contains two RepViTBlock modules with a stride of 2, outputting high-level features with 512 channels and a spatial size of 20×20. The three feature maps P3, P4, and P5 with downsampling factors of 8, 16, and 32 are output sequentially. 3e P 4e and P 5e HybridEncoder performs cross-scale feature encoding on these three feature maps to obtain the encoded three feature maps P3, P4 and P5;

[0046] When the stride of the RepViTBlock module is 1, the spatial information mixing module first inputs the input feature map into the RepVGGDW module for depthwise separable convolution processing. The RepVGGDW module includes three branches: depthwise separable convolution and identity mapping. The outputs of the three branches are summed element-wise and then subjected to batch normalization to enhance the local spatial feature representation. Subsequently, depending on the module configuration, a channel attention mechanism is selected to adaptively weight the importance of different channels to obtain the spatially enhanced feature map. The spatially enhanced feature map is then input into the channel feature fusion module, which includes a GELU activation function and pointwise convolution to complete channel dimensionality enhancement, nonlinear mapping, and channel recombination. Finally, the output of the channel feature fusion module is summed element-wise with the output of the Shortcut branch to obtain the output feature map of the RepViTBlock.

[0047] When the stride of RepViTBlock is 2, Token Mixer no longer uses the RepVGGDW module. Instead, it uses a 3×3 depthwise convolution with stride=2 to spatially downsample the input feature map, reducing the height and width of the feature map to half of their original size. Depending on the module configuration, it selects whether to use a channel attention mechanism for channel weighting. Then, the downsampled feature map is input into the channel feature fusion module, which performs channel feature mixing by sequentially passing through pointwise convolution and the GELU activation function. At the same time, the Shortcut branch uses pointwise convolution with a corresponding stride=2 to match the size and number of channels of the input feature map. Finally, the output of the channel feature fusion module is added to the output of the Shortcut branch to obtain the downsampled output feature map.

[0048] The Lite module, such as Figure 5 As shown, multi-scale fusion and enhancement processing is performed on the encoded three feature maps P3, P4, and P5 to generate corresponding enhanced feature maps. :

[0049] The three feature maps P3, P4, and P5 are represented as follows:

[0050] , , .

[0051] Where 32 represents the batch size, and 64, 128, and 256 represent the number of channels in the three feature maps, respectively. , , These are the height and width of the three feature map outputs, respectively.

[0052] The three feature maps are subjected to a pointwise convolution operation to uniformly map the channels, unifying the number of channels to 256, resulting in the mapped feature maps. , , :

[0053] ,

[0054] ,

[0055] .

[0056] Subsequently, the mapped feature map and Upsampling is performed to map the spatial dimensions of the feature map to the maximum scale. Alignment yields the upsampled feature map. Specifically:

[0057] ,

[0058] .

[0059] Upsampled feature map and Element-wise addition is performed to fuse the features at the three scales, resulting in a fused feature map. :

[0060] .

[0061] Subsequently, the feature maps were fused. Sequentially pass through depthwise separable convolutions to extract local spatial features. :

[0062] ,

[0063] right Pointwise convolution is performed to achieve channel information mixing and linear mapping, resulting in the pointwise convolution output. :

[0064] ,

[0065] Pointwise convolution output Batch normalization is performed, and the LeakyReLU activation function is applied to generate enhanced feature maps. :

[0066] ,

[0067] Enhanced feature maps Space dimensions and The same applies; the number of channels is uniformly 256.

[0068] Enhance feature maps and , The RTDETRTransformerv2 decoder is used for target querying and detection to achieve accurate positioning of vehicles and personnel at multiple scales and targets.

[0069] Step S3: Train the improved RTDETRV2 model using the target augmentation dataset to establish a joint detection model that simultaneously identifies vehicles and people. This model can simultaneously output the target bounding boxes and confidence information of vehicles and people during the first-stage inference.

[0070] The improved RTDETRV2 model was trained using the AdamW optimizer, with an initial learning rate set to... Using a cosine annealing strategy, the weights decay to... The batch size is set to 32, the number of worker threads is 8, the iteration is 300 rounds, and the decay rate is 0.9999.

[0071] Export the trained and improved RTDETRV2 model to ONNX format, set the number of query boxes to 300, the evaluation index to -1, use a model conversion tool to convert the ONNX model to an RKNN model, deploy the trained and improved RTDETRV2 model on edge nodes, and use a multi-threaded pipeline architecture to achieve real-time video stream detection.

[0072] In this embodiment of the invention, the real-time video feed from the Hikvision camera is read using the SDK streaming method.

[0073] Vehicle and pedestrian target detection is performed using the improved RTDETRV2 model. Let the tensor of the input image in the i-th frame be... i=1,2,… , ,in For batch size, For image resolution, The total number of frames is then fed into the improved RTDETRV2 model. Multi-scale features of vehicles and personnel are extracted to obtain the corresponding vehicle detection result set. Collection of personnel test results That is, the continuous target trajectory of vehicles and the continuous target trajectory of personnel:

[0074] ,

[0075] ,

[0076] in, The index of the vehicle detection result in the set, with a value ranging from 1 to... , This represents the total number of vehicle targets detected in the i-th frame. This represents the index of the personnel detection result in the set, with a value ranging from 1 to... , This represents the total number of human targets detected in the i-th frame. Let i be the set of vehicle detection results for the i-th frame. Let i be the set of personnel detection results for the i-th frame. For the i-th frame, the first... Category labels for individual vehicle targets For the i-th frame, the first... Category labels for individual personnel goals, For the i-th frame, the first... Detection confidence of individual vehicle targets For the i-th frame, the first... Detection confidence level of individual targets, For the i-th frame, the first... The bounding box coordinates and dimensions of each vehicle target. For the i-th frame, the first... The bounding box coordinates and dimensions of each personnel target.

[0077] Each bounding box Represented as:

[0078] ,

[0079] , For the i-th frame, the first... The coordinates of the top left corner of the vehicle target bounding box , For the i-th frame, the first... The height and width of the bounding box of each vehicle target.

[0080] Each bounding box Represented as:

[0081] ,

[0082] , For the i-th frame, the first... The coordinates of the top left corner of the vehicle target bounding box , For the i-th frame, the first... The height and width of the bounding box of each vehicle target.

[0083] Simultaneously initialize the BotSORT tracker to obtain model input and output attributes. The input size is fixed at 640x640, and the output consists of two tensors—a classification confidence tensor. and bounding box regression tensor 300x80 refers to the shape of the output tensor, 300 represents the maximum number of candidate bounding boxes detected in the current frame, 80 represents the number of categories, and 4 represents the four coordinate parameters of the bounding box. Load the category label file and initialize the data structures required for post-processing. Configure the Kalman filter parameters, IoU threshold, maximum number of lost frames, and global motion compensation parameters of the BoTSORT tracker, and create a unique tracking instance for subsequent inter-frame target association.

[0084] Read the weighbridge area file pre-annotated with LabelMe, parse out the polygon vertex coordinates, and calculate the scaling factor based on the actual size of the current video frame. , The original polygon vertices are mapped to the pixel coordinate system of the current frame to generate a ROI polygon, which is used for subsequent personnel position determination.

[0085] Let the coordinates of the original polygon's vertices be... If the current frame width is W and the height is H, then the mapped coordinates are... for:

[0086] , ,

[0087] in , To label the original dimensions of the image.

[0088] Read real-time frames from the camera, convert the original BGR frames to RGB format, and adjust the image to the model input size of 640×640 using a scaling method that maintains the aspect ratio. Fill any insufficient areas with a specified background color, and record the scaling ratio. and fill offset Let the original image size be... , Refers to the original image width. The scaling factor is the original image height. :

[0089] ,

[0090] The coordinates of the top left corner of the original region in the filled image for:

[0091] , .

[0092] The processed image data is pushed into the processing pool through a thread-safe queue, and the worker threads are notified.

[0093] Each edge node maintains an independent worker thread that continuously checks the queue. When the queue is not empty, it retrieves a replacement frame and performs the following operations: It feeds the preprocessed image data into the neural network accelerator for inference, obtaining the model's output classification confidence tensor O and bounding box regression tensor G. For each candidate box, it indexes d, class q, and class confidence. Applying the sigmoid function, we obtain the probability. :

[0094] ,

[0095] The Top-K algorithm is used to select the K highest-probability bounding boxes, where K=300, and their indices and corresponding categories are recorded. The selection criteria are:

[0096] , ,

[0097] This refers to the probability that the d-th candidate box belongs to category q. This refers to the highest probability of the d-th candidate box. Refers to the predicted category of the d-th candidate box. This refers to the confidence threshold. If so, then the candidate box will be retained.

[0098] For the retained candidate boxes, the normalized bounding box coordinates will be... Convert to absolute coordinates The conversion formula is:

[0099] , ,

[0100] , ,

[0101] , The coordinates of the normalized center point of the candidate box are given, w is the width of the candidate box, and h is the height of the candidate box. , Refers to the pixel coordinates of the top-left corner of the bounding box. , Refers to the pixel coordinates of the bottom right corner of the bounding box.

[0102] Map back to the original image size using the scaling and padding parameters from preprocessing:

[0103] , ,

[0104] , ,

[0105] This indicates the horizontal fill offset. This refers to the vertical fill offset, maintaining the aspect ratio to fill to 640×640. This refers to the scaling ratio. This refers to the absolute pixel coordinates of the top-left corner of the original image. This refers to the absolute pixel coordinates of the bottom right corner of the original image.

[0106] Non-maximum suppression is performed independently for each category to remove redundant bounding boxes with high overlap. IoU calculation formula:

[0107] ,

[0108] Where A and B are the areas of the two bounding boxes.

[0109] All detected human targets in the current frame are converted into the detection data structure required by the BoTSORT tracker, and then input into the cross-frame target association function respectively. Cross-frame target association is performed on the detection results and personnel detection results respectively. The association process introduces bounding box overlap and feature vector similarity:

[0110] ,

[0111] ,

[0112] Refers to the i-th frame. Feature representation of a vehicle target. Refers to the i-th frame. The characteristics of individual personnel goals This is the vehicle target index for the previous frame. For the personnel target index of the previous frame, For the current frame vehicle target target in the previous frame Match score, For the current frame personnel target target in the previous frame Match score, The intersection-union ratio (IU / U) of the target bounding box, used for spatial matching. Cosine similarity of feature vectors is used for target appearance matching. , These are the weighting coefficients. It is a polygon area function used to measure the degree of overlap between two bounding boxes.

[0113] Based on the scoring results, a unique tracking identifier is assigned to the same vehicle target. :

[0114] ,

[0115] It is the set of tracking IDs for all vehicle targets in the (i-1)th frame.

[0116] Assign a unique tracking identifier to the same person / target :

[0117] ,

[0118] This is the set of tracking IDs for all personnel targets in frame i-1.

[0119] Finally, the continuous target trajectory of the vehicle in consecutive frames is obtained. Continuous target trajectory of personnel :

[0120] ,

[0121] ,

[0122] The generated continuous target trajectory data is used as input for step S4 for vehicle stability determination and personnel area anti-cheating detection.

[0123] Step S4: For the same vehicle target, extract the sequence of vehicle detection bounding boxes for consecutive frames from the vehicle detection results obtained in Step S3. , This represents the center coordinates and size of the vehicle detection box in the i-th frame. Total number of frames Let be the width of the detection bounding box in the i-th frame image. Let be the height of the detection bounding box in the i-th frame image. Let x be the x-coordinate of the center point of the detection box in the i-th frame image. Let be the ordinate of the center point of the detection box in the i-th frame image.

[0124] ,

[0125] Calculate the vehicle pixel offset based on the position change of the vehicle detection box center point in consecutive frames. :

[0126] ,

[0127] Let x be the x-coordinate of the center point of the detection box in the (i-1)th frame of the image. It represents the ordinate of the center point of the detection box in the (i-1)th frame of the image.

[0128] Let the coordinates of the center point of the vehicle bounding box in frame t be... The vehicle pixel offset is calculated based on the position change of the center point of the vehicle detection box in consecutive frames. Pixel offset between adjacent frames To filter out single-frame jitter, a sliding window averaging method is used for smoothing.

[0129] ,

[0130] Where D is the size of the sliding window. This refers to the pixel offset of the center point of the vehicle detection bounding box in the i-th frame, where i is the frame index within the sliding window, and t is the current frame index, representing the frame currently being stabilized. This refers to the smoothed offset of the vehicle's center point, used for visual stability assessment. The sliding window size represents the number of consecutive frames used for averaging, and a first threshold is set. Pixels, first preset time Second.

[0131] The vehicle pixel offset is smoothed, and the smoothed vehicle pixel offset is compared with a first determination threshold. and the first preset time For comparison, the value of j ranges from frame i-N+1 to frame i, and the smooth offset of consecutive frames j... The value is below the threshold for N consecutive frames. And the duration reached Generate vehicle visual stability determination signal :

[0132] .

[0133] Proceed to step S5.

[0134] Step S5: Vehicle visual stop determination signal Upon triggering, the weight data sequence of the secondary weighbridge is read synchronously. , For time series indexes, such as Figure 3 As shown, this weight curve records the trend of real-time weighing data over time during a complete weighing process. The horizontal axis represents weighing time, and the vertical axis represents real-time weighing data. The curve shows that the weight rises rapidly and fluctuates briefly during the vehicle's entry onto the scale, then enters a relatively stable plateau. The stable reading in this plateau area represents the effective weight of the weighing. During the vehicle's exit from the scale, the weight quickly drops back to zero. Any jitter or abnormal jumps in the curve may be caused by the vehicle not coming to a complete stop, scale vibration, or cheating.

[0135] The original weight data at time t is subjected to low-pass filtering to remove noise, and the filtered weight is denoted as . For the current vehicle target, within the preset stopping determination time window... Internally calculated weight fluctuation range:

[0136] ,

[0137] This refers to the vehicle's weight at time t. This refers to the preset stopping and stabilization determination time. This refers to the vehicle's weight fluctuation range, which is the difference between the maximum and minimum weight values ​​within a time window. The second threshold value is set to 50 kg. Refers to the second preset time. This refers to the start time of the time window, corresponding to the moment when the vehicle has just detected a stop signal. In practical applications, the above parameters can be adaptively adjusted according to the camera resolution, weighbridge accuracy, and on-site environment. It can be adjusted within the range of 1-5. It can be adjusted within the range of 30-100kg;

[0138] When the weight fluctuation range Below the second judgment threshold And the duration reaches the second preset time. At that time, a vehicle weight stabilization determination signal is generated:

[0139] ,

[0140] when When the vehicle has entered a stopped weighing state, a stop signal is output, and the process proceeds to step S6; if If the vehicle is not completely stopped, proceed to step S7.

[0141] Step S6: Manually calibrate the weighing area of ​​the truck scale in the image, generate the corresponding pixel-level ROI mask Q, and save it as a JSON file. During manual calibration, establish a pixel coordinate system with the top left corner of the image as the origin, with the x-axis along the horizontal direction and the y-axis along the vertical direction. Map the coordinates of the polygon vertices in the JSON data generated by LabelMe onto each frame of the image. After confirming that the vehicle has entered the stopped weighing state, the personnel get off the vehicle. Use the improved RTDETRv2 model from step S3 to perform real-time detection of the personnel and obtain the personnel detection box. , Indicates the first For a person target in a frame, extract the pixel coordinates of the two endpoints of the bottom edge of the detection box:

[0142] ,

[0143] ,

[0144] Calculate the pixel coordinates of the left and right endpoints of the bottom edge of the detection box. and :

[0145] ,

[0146] ,

[0147] Points at both ends of the bottom edge The coordinates of the ROI region are compared with those marked in the JSON file. The cross product method is used to determine whether the two endpoints of the bottom edge are located within the marked region mask. The specific steps are as follows:

[0148] Let the four vertices of the ROI polygon be denoted as follows, in clockwise order: , , , For the point to be judged Calculate the points in sequence Relative to the cross product of each edge:

[0149] ,

[0150] ,

[0151] ,

[0152] ,

[0153] like , , , If all are greater than 0 or all are less than 0, then the point It is located inside the quadrilateral, otherwise it is located outside.

[0154] The left end of the bottom edge of the personnel detection frame and right endpoint Perform the above judgments separately, where:

[0155] , (Left end of the bottom edge of the frame)

[0156] , (Right end of the bottom edge of the frame)

[0157] The absolute pixel coordinates of the left endpoint on the original image. This refers to the absolute pixel coordinates of the right endpoint on the original image.

[0158] When the same person or target appears in N consecutive frames, if S and If at least one point is located inside the ROI, the person is determined to be in the weighbridge area in the current frame, and is considered to have entered the weighing area and committed cheating; otherwise, the person is determined not to have boarded the weighbridge.

[0159] Define the decision function:

[0160] ,

[0161] When the same person target appears in N consecutive frames, and any one of the two endpoints of the bottom edge of its detection box lies within the region mask, the condition is met. :

[0162] ,

[0163] when If the person is determined to have entered the weighing area and committed cheating, proceed to step S8; otherwise, proceed to step S9. .

[0164] Step S7: When the vehicle visually stops, a stop determination signal is received. Effective, and the fluctuation range of vehicle weight Not lower than the second judgment threshold Call the corresponding time period video frames Pixel difference analysis is performed on consecutive frame images to determine whether the weighbridge is swaying or the weight sensor is interfering. Let the consecutive frame differences be:

[0165] ,

[0166] in It is the sum of the absolute values ​​of the pixel differences between the k-th frame and the previous frame, reflecting the amplitude of image motion. For the k-th frame image, This is the image of the (k-1)th frame.

[0167] Define interference determination function :

[0168] ,

[0169] The image difference threshold is used to determine whether there is weighbridge swaying or interference when the threshold is exceeded, covering the entire time period. Aggregation determination:

[0170] ,

[0171] If the error is detected, it is determined to be interference. The weight abnormality is ignored, and the process proceeds to step S9. If no interference is detected, return to step S5 to re-evaluate stability. To prevent an infinite loop, a maximum stability evaluation waiting time is set. If the vehicle fails to stop after exceeding the limit, it will be considered cheating.

[0172] Step S8: After determining that it is a cheating behavior, the edge node executes an abnormal alarm and saves the video clips, vehicle stopping determination results and weight data of the corresponding time period to form a traceable event evidence chain, and then proceeds to step S9.

[0173] Step S9: The edge node pushes the vehicle detection results, target tracking results, vehicle stability determination results, anti-cheating determination results, and event recording results to the monitoring center in structured data format, retrieves the latest frame image for real-time display, and writes it to the video file. For example... Figure 2 As shown, the improved RTDETRV2 model displays confidence information in blue, tracker outputs confidence and ID information in red, and speed information in green, based on the video frame rate to control display speed, confidence, and ID. Figure 4 As shown, all personnel detection boxes are drawn in different colors according to the judgment result. If they are within the ROI, they are drawn in blue; otherwise, they are drawn in red. The category and confidence level are marked above the box. The tracking box is drawn in green and marked with the tracking ID and confidence level. The ROI polygonal area is filled and overlaid with semi-transparent green. After drawing, the image is converted to BGR format. The main thread checks the display queue in a loop and pushes the data to the main thread through the display queue. When the trajectory has not been updated for a long time, the relevant data of the trajectory is cleared from the history.

Claims

1. A method for preventing cheating on truck scales based on edge computing, characterized in that, The method comprises the following steps: Step S1, collecting vehicle weighing video images of the working area of the truck scale, constructing a special target image dataset covering different illuminations, angles and vehicle states, and performing combined image enhancement processing on the special target image dataset to obtain a target enhanced dataset, and proceeding to step S2; Step S2, constructing an improved RTDETRV2 model for outputting the boundary box and confidence information of the vehicle target and the personnel target; The improved RTDETRV2 network comprises a Patch Embedding module, a RepViT backbone network, a Lite module and an RTDETRTransformerv2 decoder; The input image is an RGB three-channel image. The Patch Embedding module performs convolutional downsampling on the input image to reduce its spatial size and generate a preliminary feature map. The preliminary feature map is then sequentially input into the RepViT backbone network for hierarchical feature extraction, outputting three feature maps P at different downsampling ratios. 3e P 4e and P 5e HybridEncoder performs cross-scale feature encoding on these three feature maps to obtain the encoded three feature maps P3, P4 and P5; The Lite module performs multi-scale fusion and enhancement processing on the encoded three feature maps P3, P4, and P5, generating corresponding enhanced feature maps. ; proceeding to step S3; Step S3: Train the improved RTDETRV2 network using the target augmentation dataset to obtain the improved RTDETRV2 model. Deploy the improved RTDETRV2 model at edge nodes and pull the video stream in real time through the camera. Let the input image tensor of the i-th frame be... , , ,in For batch size, For image resolution, For the total number of frames, Input the improved RTDETRV2 model Extract multi-scale features of vehicles and people, and generate corresponding target detection results, namely the continuous target trajectory of vehicles and the continuous target trajectory of people. Step S4, for the same vehicle target, extracting a vehicle detection box sequence of consecutive frames from the vehicle detection result obtained in step S3, calculating a vehicle pixel offset based on the position change of the center point of the vehicle detection box of the consecutive frames, and performing smoothing processing on the vehicle pixel offset, when the smoothed vehicle pixel offset is lower than a first determination threshold in consecutive frames and the duration reaches a first preset time, a vehicle visual stop determination signal is generated, and proceeding to step S5; Step S5, after the vehicle visual stop determination signal is triggered, synchronously reading the truck scale secondary scale weight data, calculating the weight fluctuation amplitude within a preset stop determination time window, when the weight fluctuation amplitude is lower than a second determination threshold and the duration reaches a second preset time, it is determined that the vehicle enters a stop and weighing state, and proceeding to step S6; if the weight fluctuation amplitude does not satisfy the second determination threshold, proceed to step S7; Step S6: Manually mark the weighing area of ​​the truck scale in the image and generate the corresponding pixel-level mask. After confirming that the vehicle has entered the stopped weighing state, the personnel get off the vehicle. Use the improved RTDETRv2 model from step S3 to detect the personnel in real time, obtain the personnel detection box, extract the pixel coordinates of the two endpoints of the bottom edge of the detection box, and determine whether the two endpoints of the bottom edge are located within the marked area mask. If the same personnel target has any one of the two endpoints of the bottom edge of its detection box located within the area mask in N consecutive frames, it is determined that the personnel have entered the weighing area and constitute cheating behavior, and proceed to step S8; otherwise, proceed to step S9. ; Step S7, when the vehicle visual stop determination signal is valid, and the weight fluctuation amplitude is not lower than the second determination threshold, the video images of the corresponding period are called, the video key frames of the corresponding period are extracted, and whether there is a weighbridge shaking or weight sensor interference is judged according to the key frame images, when it is determined that there is interference, the weight anomaly is ignored, and proceeding to step S9; if it is determined that there is no interference, return to step S5 and re-determine the stability; Step S8, after it is determined that it is cheating, the edge node performs abnormal alarm, and saves the video clips, vehicle stop determination result and weight data of the corresponding period, forms a traceable event evidence chain, and proceeds to step S9; Step S9, the edge node pushes the vehicle detection result, target tracking result, vehicle stop determination result, anti-cheating determination result and event record result to the monitoring center in the form of structured data.

2. The method according to claim 1, characterized in that, In step S2, the encoded three feature maps P3, P4 and P5 are represented as: , , , in, Indicates batch size. , , These represent the number of channels in the three feature maps. , , These are the height and width of the three feature map outputs, respectively. The three feature maps are each subjected to a pointwise convolution operation to perform unified channel mapping, unifying the number of channels to ch (preferably ch = 256), resulting in the mapped feature maps. , , : , , , Subsequently, the mapped feature map and Upsampling is performed to map the spatial dimensions of the feature map to the maximum scale. Alignment yields the upsampled feature map. Specifically: , , Upsampled feature map and Element-wise addition is performed to fuse the features at the three scales, resulting in a fused feature map. : , Subsequently, the feature maps were fused. Sequentially pass through depthwise separable convolutions to extract local spatial features. : , right Pointwise convolution is performed to achieve channel information mixing and linear mapping, resulting in the pointwise convolution output. : , Pointwise convolution output Batch normalization is performed, and the LeakyReLU activation function is applied to generate enhanced feature maps. : , Enhanced feature maps Space dimensions and The same applies; the number of channels is uniformly ch. Enhance feature maps and , The RTDETRTransformerv2 decoder is used for target querying and detection to achieve accurate positioning of vehicles and personnel at multiple scales and targets.

3. The method of claim 2, wherein, Step S3 is as follows: , , in, The index of the vehicle detection result in the set, with a value ranging from 1 to... , This represents the total number of vehicle targets detected in the i-th frame. This represents the index of the personnel detection result in the set, with a value ranging from 1 to... , This represents the total number of human targets detected in the i-th frame. Let i be the set of vehicle detection results for the i-th frame. Let i be the set of personnel detection results for the i-th frame. For the i-th frame Category labels for individual vehicle targets For the i-th frame Category labels for individual personnel goals, For the i-th frame, the first... Detection confidence of individual vehicle targets For the i-th frame, the first... Detection confidence level of individual targets, For the i-th frame The bounding box coordinates and dimensions of each vehicle target. For the i-th frame, the first... The bounding box coordinates and dimensions of each personnel target; Each bounding box Represented as: , For the i-th frame The coordinates of the top left corner of the vehicle target bounding box For the i-th frame The height and width of the bounding box of each vehicle target; Each bounding box Represented as: , For the i-th frame, the first... The coordinates of the top left corner of the target bounding box for each person. For the i-th frame, the first... The height and width of the target bounding box for each individual; The vehicle detection result and the personnel detection result of the continuous frames are respectively input into a cross-frame target association function Data association is performed on the same target in continuous frames, and a matching score function considers the bounding box overlap degree and feature vector similarity: , , Refers to the i-th frame. Feature representation of a vehicle target. Refers to the i-th frame. The characteristics of individual personnel goals This is the vehicle target index for the previous frame. For the personnel target index of the previous frame, For the current frame vehicle target target in the previous frame Match score, For the current frame personnel target target in the previous frame Match score, The intersection-union ratio (IU / U) of the target bounding box, used for spatial matching. Cosine similarity of feature vectors is used for target appearance matching. , These are the weighting coefficients. It is a polygon area function used to measure the degree of overlap between two bounding boxes; assigning a unique tracking identification to the same vehicle object based on the scoring result : , a set of tracking IDs for all vehicle targets in the i-1th frame; Assigning a unique tracking identification to the same person object : , a set of tracking IDs for all person targets in the i-1th frame; a continuous target trajectory of the vehicle in consecutive frames and a continuous target trajectory of the person : , , The generated continuous target trajectory data is used as the input of step S4 for vehicle stop determination and personnel area anti-cheating detection.

4. The method of claim 3, wherein, In step S4, the vehicle pixel offset is the Euclidean distance of the vehicle detection box center point of the same vehicle target in adjacent frames; A sequence of vehicle bounding boxes of consecutive frames is extracted from the continuous target trajectory of the vehicle obtained in step S3 for the same vehicle target , denote the center coordinates and size of the vehicle bounding box of the i-th frame, is the total number of frames, is the width of the i-th frame image bounding box, is the height of the i-th frame image bounding box, is the horizontal coordinate of the center point of the i-th frame image bounding box, is the vertical coordinate of the center point of the i-th frame image bounding box; , According to the position change of the center point of the vehicle detection frame of the continuous frame, the vehicle pixel offset is calculated : , xi-1is a horizontal coordinate of the center point of the detection frame for the i-1th frame image, yi-1is a vertical coordinate of the center point of the detection frame for the i-1th frame image. The vehicle pixel offset is smoothed, and the smoothed vehicle pixel offset is compared with a first determination threshold and a first preset time , j is in a range from i-N+1 to i, and when the smoothed offsets of the continuous j frames are all lower than the threshold in the continuous N frames and the duration reaches , a vehicle visual stop determination signal is generated : 。 5. The method of claim 4, wherein, The vehicle pixel offset is smoothed, that is, the vehicle pixel offset of consecutive frames is averaged in a sliding window to filter out single-frame detection jitter. Vehicle vision stop determination signal After triggering, synchronously read the weight data sequence of the truck scale secondary meter , Index the time series, and calculate the weight fluctuation amplitude within the preset stop determination time window for the current vehicle target ​ , denotes the weighing value of the vehicle at time t, denotes the preset stationary determination time, denotes the weight fluctuation range of the vehicle, i.e. the difference between the maximum and minimum values of the weight within the time window, denotes the second determination threshold, denotes the second preset time, denotes the start time of the time window, corresponding to the time when the vehicle just detects the stationary signal. When the weight fluctuation amplitude is lower than the second determination threshold and the duration reaches the second preset time , a vehicle weight stop determination signal is generated : , When the vehicle enters the stationary state, a stationary signal is outputted and the process goes to step S6; if the vehicle does not enter the stationary state, the process goes to step S7.

6. The method of claim 5, wherein, ​ In the image, the weighing area of the truck scale is manually calibrated, and the corresponding pixel-level mask Q is generated. After confirming that the vehicle enters the stable weighing state, the personnel get off the vehicle, and the improved RTDETRV2 model in step S3 is used for real-time detection of the personnel to obtain the personnel detection frame , represents the detection frame of a certain personnel target in the i-th frame, is the total number of frames, and the pixel coordinates of the two end points of the bottom edge of the detection frame are extracted :​ , , determining whether the two end points of the bottom side are located within the above-mentioned calibrated region mask, defining a decision function : whether the two end points of the bottom side are located within the above-mentioned calibrated region mask, defining a decision function : , When the same person target is in consecutive N frames, and any one of the two end points of the bottom side of the detection frame is located in the region mask, the condition is satisfied : , When , it is determined that the person enters the weighing area and constitutes a cheating behavior, and step S8 is entered, otherwise step S9 is entered, in which .

7. The method of claim 6, wherein, ​ When the vehicle vision is stopped and the determination signal is judged Effective, and the vehicle weight fluctuation range , call the corresponding period Video frame , the pixel difference analysis of the continuous frame image is carried out to judge the scale swing or weight sensor interference, and the continuous frame difference is: , wherein is the sum of the absolute values of the pixel difference between the kth frame and the previous frame, reflecting the image motion amplitude, is the kth frame image, is the k-1th frame image; Defining an interference decision function : , For image difference threshold, when exceeding the threshold, it is considered that there is a scale swing or disturbance, and the whole time period Aggregation determination: , Aggregation decision function When, the interference is determined, the weight abnormality is ignored, and the process goes to step S9; , the non-interference is determined, the process returns to step S5, and the stabilization is determined again. In order to prevent the dead loop, the maximum stabilization waiting time is set , the cheating is determined when the over-limit is still not stopped.

8. The method of claim 7, wherein, ​ ​ In step S5, the weight fluctuation range is defined as the difference between the maximum and minimum values of the weight in the stable stopping time window; The first determination threshold, the second determination threshold, the continuous frame number N, the first preset time and the second preset time in step S4 and step S5 are adaptively adjusted according to the camera resolution, the truck scale accuracy and the on-site environment.

9. The method of claim 8, wherein: The trajectory calculation, the vehicle stable stopping determination, the region mask determination and the alarm determination are independently performed on each vehicle target and each personnel target in the working area based on the unique tracking identifier of each vehicle target and each personnel target. The edge node performs the video image preprocessing, the target detection, the target tracking, the vehicle stable stopping determination, the anti-cheating determination and the result output in a multi-thread pipeline manner.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the method of any one of claims 1-9.