Electric drive parameter acquisition method based on stirring transport vehicle platform
By using a camera on the mixing transport vehicle to capture the display screen, perform data annotation and anti-shake processing, designing the excitation branch structure and cross-mixed model, the problem of low efficiency and accuracy of electric drive parameters acquisition by the mixing transport vehicle is solved, and efficient, accurate identification and real-time monitoring of electric drive parameters are achieved.
Patent Information
- Application Number
- CN202510285371.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, the electric drive parameter acquisition of mixing transport trucks has the problem of poor identification efficiency and accuracy, and the sensor is inconvenient to maintain, easy to damage, and inevitably damage the circuit of the tanker truck, resulting in inefficient losses.
The camera captures the screen of the display screen, labels and annotates it, obtains the original data set, and then performs anti-shake processing on the data, designs the excitation branch structure, uses the cross-pseudo-supervision algorithm to train the cross-mixed model to obtain the target model. Finally, the real-time display screen is collected and inputs into the target model, identifying the electric drive parameters and uploading it to the cloud platform.
It improves the efficiency and accuracy of electric drive parameter identification, reduces the complexity and damage risk of sensor maintenance, and realizes real-time monitoring of electric drive parameters and macro-control strategies of cloud platform.
Smart Images

Figure CN120220156A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the technical field of data recognition, and particularly to a method for collecting electric drive parameters based on a mixer truck platform. Background Art
[0002] Currently, a mixer truck is a special truck used to transport ready-mixed concrete for construction. A cylindrical mixing drum is installed on the truck to carry the mixed concrete. Nowadays, most mixer trucks have switched to electric drive, relying on power batteries and control units to achieve energy conservation, emission reduction, and flexible control.
[0003] In addition, in order to ensure the normal operation of the electric drive during transportation and avoid insufficient mixing of concrete, some electric drive parameter acquisition devices are often used to monitor electric drive parameters. Currently, external physical sensors are commonly used at home and abroad to measure electric parameters. This method has the problems of inconvenient sensor maintenance, easy damage, and inevitably damaging the circuit of the tanker, which is likely to cause immeasurable losses.
[0004] In addition to the above problems, during transportation, the transport vehicle will inevitably experience bumps, resulting in vibration of the display screen and the acquisition device. In response to this problem, anti-shake OCR technology has been proposed to eliminate the error caused by the vibration of the screen during the process of collecting the parameters of the electric drive system on the display screen, and then upload the data to the cloud platform to achieve real-time monitoring.
[0005] The initial OCR technology mainly relied on the template matching method, with low recognition efficiency and limited application scope. The traditionally defined OCR technology can be regarded as a classic pattern recognition problem, mainly including three steps: preprocessing of images, feature extraction, and classification, and has formed a relatively complete system. However, its application scenario is relatively single, and the recognition effect in complex environments is poor. The end-to-end deep learning OCR technology, as a new research hotspot, demonstrates stronger detection and recognition capabilities.
[0006] Traditional OCR technology first uses connected component detection or MSER to locate the text area in the image, and judges whether it is a text connected component through heuristic rules. Then, the recognition algorithm performs feature classification and recognition on the text in the extracted text area. In addition, in order to improve the accuracy of text detection, traditional OCR technology also performs image correction operations, such as horizontal correction and perspective correction, before image input detection. These steps together constitute the core of traditional OCR technology, aiming to accurately extract text information from images. However, it is slow, with complex post-processing, requires loading both the text detection model and the recognition model simultaneously, has complex post-processing, is slow, and the accuracy is affected by the interaction of the two models.
[0007] Obviously, there is an urgent need for a method for collecting electric drive parameters based on a mixer truck platform with high recognition efficiency and accuracy. Summary of the Invention
[0008] In view of this, embodiments of the present disclosure provide a method for collecting electric drive parameters based on a mixer truck platform, which at least partially solves the problems of poor recognition efficiency and accuracy in the prior art.
[0009] Embodiments of the present disclosure provide a method for collecting electric drive parameters based on a mixer truck platform, including:
[0010] Step 1, within a preset duration of the mixer truck running, capture the screen of the display through a camera and perform label annotation to obtain an original data set;
[0011] Step 2, perform anti-shake processing on the original data set to obtain a preprocessed data set;
[0012] Step 3, design an excitation branch structure and input the preprocessed data set into it to obtain a reconstructed image data set;
[0013] Step 4, use the cross pseudo-supervision algorithm and the reconstructed image data set to train a cross hybrid model to obtain a target model;
[0014] Step 5, collect the real-time screen image of the mixer truck and input it into the target model. After obtaining the electric drive parameter recognition result, upload it to the cloud platform, and the cloud platform analyzes it and formulates a macro-control strategy.
[0015] According to a specific implementation manner of the embodiments of the present disclosure, the specific steps of Step 1 include:
[0016] Step 1.1, obtain the screen image of the display captured by the camera, and output the screen image to the industrial computer in the form of a video stream;
[0017] Step 1.2, the industrial computer uses a video processing library to read the video stream, and extracts one frame from the video stream at a fixed interval to form a sampling data set;
[0018] Step 1.3, label the text labels for a preset proportion of the video frame data in the sampling data set according to the data types collected by the sensor, and do not label the remaining part, jointly forming the original data set.
[0019] According to a specific implementation manner of the embodiments of the present disclosure, the specific steps of Step 2 include:
[0020] Step 2.1, use the optical flow learning algorithm to calculate a backward dense warping field for each frame of the original data set;
[0021] Step 2.2, adopt the RAFT algorithm to restore the pixel loss caused by warping through optical flow estimation technology;
[0022] Step 2.3, through the frame fusion technology, mix and distort the color frames in the image space to generate output stable frames, forming a preprocessed dataset.
[0023] According to a specific implementation manner of the embodiment of the present disclosure, the step 3 specifically includes:
[0024] Step 3.1, design an excitation branch structure, where the excitation branch structure includes an encoder, a decoder, a slot attention module, an attention loss, and a CTC loss. The encoder is used to extract features from the image, the decoder is used to predict the final text or background detection box and reconstruct the image, the slot attention module clusters each pixel into the text or background layer based on the semantic characteristics of the text, and the attention loss and the CTC loss are used to locate the position of the target;
[0025] Step 3.2, input all the images in the preprocessed dataset into the excitation branch structure to obtain a reconstructed image dataset with detection boxes.
[0026] According to a specific implementation manner of the embodiment of the present disclosure, the expression of the attention loss is
[0027]
[0028] where α ij is the true attention weight, is the predicted attention weight, and N represents the number of elements in the sequence T;
[0029] The expression of the CTC loss is
[0030]
[0031] where p(T i |X i ) is the probability of the predicted sequence T i when the input sequence X i is given.
[0032] According to a specific implementation manner of the embodiment of the present disclosure, the cross - hybrid model includes a teacher model SVTRv2 and a student model RepSVTR, and the step 4 specifically includes:
[0033] Step 4.1, divide the training of the cross - hybrid model into a supervised branch and a semi - supervised branch. For the supervised branch, the input requirement is the labeled images in the reconstructed image dataset. For the semi - supervised branch, the input requirement is a fixed proportion of unlabeled images and labeled images;
[0034] Step 4.2: According to the inputs of the supervised branch and the semi-supervised branch, the teacher model SVTRv2 and the student model RepSVTR perform parallel inference and output their respective prediction results. The prediction results of the labeled images are constrained by the ground truth labels. For unlabeled images, their prediction results are used as the constraint labels for each other for cross-validation.
[0035] Step 4.3: Calculate and use the focal Tversky loss and the guided cross-entropy loss to control the parameter update of the cross hybrid model according to the prediction results.
[0036] Step 4.4: After the parameter update of the cross hybrid model is completed, perform model pruning, model conversion, and quantization on it to obtain the target model.
[0037] According to a specific implementation manner of the embodiments of the present disclosure, the specific steps of step 5 are as follows:
[0038] Step 5.1: Collect the real-time display screen image of the mixer truck and input it into the target model. After obtaining the electric drive parameter recognition result, store the electric drive parameter recognition result locally and upload it to the cloud platform through the mobile network.
[0039] Step 5.2: After the cloud platform receives the electric drive parameter recognition result, analyze it, generate a scheduling strategy, and feedback it to the transportation personnel.
[0040] The electric drive parameter acquisition solution based on the mixer truck platform in the embodiments of the present disclosure includes: Step 1, within a preset duration of the operation of the mixer truck, capture the screen of the display screen through a camera and perform label annotation to obtain an original data set; Step 2, perform anti-shake processing on the original data set to obtain a preprocessed data set; Step 3, design an excitation branch structure and input the preprocessed data set into it to obtain a reconstructed image data set; Step 4, use the cross pseudo-supervision algorithm and the reconstructed image data set to train a cross hybrid model to obtain a target model; Step 5, collect the real-time display screen image of the mixer truck and input it into the target model. After obtaining the electric drive parameter recognition result, upload it to the cloud platform, and the cloud platform analyzes it and formulates a macro-control strategy.
[0041] The beneficial effects of the embodiments of the present disclosure are as follows: Through the solution of the present disclosure, an anti-shake algorithm is adopted to improve the data quality, an excitation branch with an attention mechanism is used to accelerate the recognition speed, a cross pseudo-supervision model with an advanced backbone is built, and a loss function with clear division of labor is designed to obtain the optimal training result, improving the efficiency and accuracy of electric drive parameter recognition. Description of the Drawings
[0042] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following will briefly introduce the accompanying drawings required for the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0043] Figure 1 It is a schematic flowchart of a method for collecting electric drive parameters based on a mixer truck platform provided by an embodiment of the present disclosure;
[0044] Figure 2 It is a schematic structural diagram of a data collection device provided by an embodiment of the present disclosure;
[0045] Figure 3 It is a schematic flowchart of a method for generating a stable frame by mixing and distorting color frames provided by an embodiment of the present disclosure;
[0046] Figure 4 It is a schematic diagram of a technical solution of a teacher model SVTRv2 provided by an embodiment of the present disclosure;
[0047] Figure 5 It is a structural diagram of an excitation branch provided by an embodiment of the present disclosure;
[0048] Figure 6 It is a schematic diagram of a training branch under a cross-model structure provided by an embodiment of the present disclosure;
[0049] Figure 7 It is a model effect demonstration provided by an embodiment of the present disclosure;
[0050] Figure 8 It is an overall flowchart of a method for collecting electric drive parameters based on a mixer truck platform provided by an embodiment of the present disclosure. Specific Embodiments
[0051] The following will describe the embodiments of the present disclosure in detail with reference to the accompanying drawings.
[0052] The following illustrates the embodiments of the present disclosure through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of them. The present disclosure can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.
[0053] It should be noted that the following describes various aspects of embodiments within the scope of the appended claims. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on this disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement a device and / or practice a method. Additionally, this device and / or method can be implemented using other structures and / or functionality in addition to one or more of the aspects set forth herein.
[0054] It should also be noted that the diagrams provided in the following embodiments only schematically illustrate the basic concept of the present disclosure. The diagrams only show the components related to the present disclosure and are not drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0055] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0056] An embodiment of the present disclosure provides a method for collecting electric drive parameters based on a mixer truck platform. The method can be applied to the process of identifying electric drive parameters of vehicles in scenarios such as industry and transportation.
[0057] See Figure 1 , which is a schematic flowchart of a method for collecting electric drive parameters based on a mixer truck platform provided by an embodiment of the present disclosure. As Figure 1 shown, the method mainly includes the following steps:
[0058] Step 1, within a preset duration of the mixer truck running, capture the screen of the display through a camera and perform label annotation to obtain an original data set;
[0059] Specifically, when implementing, the process of collecting data and performing label annotation can be as follows:
[0060] The deployment diagram of the non-contact electric drive parameter acquisition device based on OCR technology is as Figure 2, mainly including a voltage sensor, a current sensor, a rotational speed sensor, a display screen, a camera, an industrial control computer, and a storage battery. Among them, general contact sensors such as voltage, current, and rotational speed are used as key components of the electric drive system of the mixer truck to collect electric drive parameter information in real time. To achieve interaction with the platform, non-contact sensors such as cameras and industrial control computers are used to obtain the electric drive system parameters on the display screen. The camera captures the display screen image, and OCR technology is applied to read the electric drive parameters on the display screen. Finally, the industrial control computer uploads the electric drive parameters to the platform in real time.
[0061] The camera captures the display screen image, and then outputs the image to the industrial control computer in the form of a video stream. By using a video processing library to read the video file, a frame is extracted from the video every fixed time interval.
[0062] First, clarify the types of data collected by the sensors, such as current, voltage, rotational speed, etc. Labeling tools such as LabelImg and CVAT are selected. Since the subsequent model involves semi-supervised learning, only part of the data is labeled. The specific classification is as follows (the values are for reference only):
[0063] {Current: 5A|10A|20A, Label: Low|Normal|High}
[0064] {Voltage: 200V|220V|240V, Label: Low|Normal|High}
[0065] {Rotational speed: 1000rpm|1500rpm|1800rpm, Label: Low speed|Normal|High speed}.
[0066] Step 2, perform anti-shake processing on the original dataset to obtain a preprocessed dataset;
[0067] Specifically, when collecting the electrical signal data on the display screen, there is a problem of image jitter. To eliminate this influence, an anti-shake algorithm is added to preprocess the original data to improve the image quality:
[0068] Motion estimation and smoothing method: In this patent, an advanced video stabilization algorithm based on optical flow learning is adopted. This method calculates a backward dense warping field F for each frame of the video kt →ks. Here, ks represents the source input frame, and kt represents the target stable output frame. These warping fields can be directly applied to the input video to achieve a stable effect. However, the stable video generated by this method often has irregular boundaries and a large number of missing pixels, resulting in a large amount of cropping of the output video, thus losing a part of the content.
[0069] Optical flow estimation technology: To recover the missing pixels caused by warping, optical flow estimation technology is adopted. Specifically, for each key frame I ksAt time k, the RAFT (Recurrent All-Pairs Field Transforms for Optical Flow) algorithm is used to compute the optical flow {F ns →ks}n∈Ω k from neighboring frames to the key frame. Here, n represents the index of the neighboring frame, and Ω k denotes the set of frames adjacent to the key frame I ks . In this way, the pixels of the neighboring frames are accurately projected onto the target stable frame.
[0070] Frame warping: The neighboring frames {I ns}n∈Ω k are warped to align with the target frame I kt in the virtual camera space. We compute the warping field {F kt →ns}n∈Ω ks from the target frame to the neighboring frames by chaining the warping field F k →ks from the target frame to the key frame and the estimated optical flow {F kt →ns}n∈Ω k from the key frame to the neighboring frames. Using the backward warping technique, the neighboring frame I ns is warped to match the target frame I kt . Meanwhile, the visibility mask {M ns}n∈Ω k of each neighboring frame is computed to identify which pixels are valid or occluded in the source frame.
[0071] Frame blending: After frame alignment, the output stable frame is generated by directly blending the warped color frames in the image space, as Figure 3 shown.
[0072] The following expression represents the process from the input video frame to the output stable video frame:
[0073] StabilizedFrame = F stabilize (F smooth (F flow (Frame t-1 , Frame t )), Frame t )
[0074] where Frame t-1 and Frame t represent consecutive video frames respectively; F flow is the optical flow estimation function that estimates the optical flow between two consecutive frames. F smooth is the motion trajectory smoothing function that is used to smooth the motion trajectories in the optical flow field. F stabilizeIt is a frame warping and content filling function that warps video frames according to the smoothed optical flow field and processes blank areas to generate stable video frames.
[0075] Step 3: Design an excitation branch structure and input the preprocessed dataset into it to obtain a reconstructed image dataset.
[0076] In specific implementation, the end-to-end recognition accuracy of the SVTRv2 algorithm is improved by 2.5% compared with PP-OCRv4, and the inference speed remains the same. This algorithm is used in a knowledge distillation framework. As Figure 4 shown, the teacher model is SVTRv2, and the student is RepSVTR, where RepViT and SVTR are the backbones of the student model. The specific process is as follows:
[0077] Model construction of the excitation branch: Since the SVTRv2 is a lightweight model compared with the RepSVTR model and does not have the knowledge transferred by the teacher, SVTRv2 cannot exhibit the same performance as knowledge distillation. Therefore, before inferring samples of SVTRv2, an external clue is provided to guide it to distinguish text and background. This clue needs to frame the area where the text content is located and is provided by the newly designed excitation branch, as Figure 5 shown.
[0078] Through the encoder-decoder structure and the slot attention mechanism, this branch can help the decoder better focus on different parts of the input data. Especially when a picture contains multiple objects, it can enable the model to better understand and represent each object. The slot is represented by an embedding vector, which captures the semantic information of the slot. Let S be the slot embedding matrix, where each row Sj represents the embedding vector of a slot. There are N slots in total, so S ∈ R N×d , where d is the dimension of the embedding vector. Use the attention mechanism to calculate the weight α of each slot for the feature map ij representing the attention weight between the i-th query and the j-th key, S i represents the query vector, represents the key vector, which is also an element in the input sequence but is related to the j-th position. The superscript T represents the transpose operation, that is, converting a row vector into a column vector. D is a scaling factor to prevent numerical instability problems caused when the inner product of S i and F j is very large or very small. By dividing by D, the numerical value of the attention weight can be kept within a reasonable range.
[0079] Perform weighted summation on the feature map to obtain the feature representation h of each slot si . The formula is as follows:
[0080]
[0081] Design of the incentive branch loss function: In the selection of the loss function, the attention loss and the CTC loss are used. The attention loss enables the model to learn which parts of the input sequence should be focused on during the decoding process, helping the model to more accurately capture key information when generating the output sequence; the CTC loss can directly map each time step of the input sequence to specific elements of the output sequence, that is, it enables the model to learn the alignment between the input and the output, and it is a loss function used to train an end-to-end handwritten text recognition system. The loss function is as follows:
[0082]
[0083] where α ij is the true attention weight, is the predicted attention weight p(T i |X i ) is the probability of the predicted sequence T i given the input sequence X i .
[0084] Then all the images in the preprocessed dataset are input into the incentive branch structure to obtain a reconstructed image dataset with detection boxes.
[0085] Step 4, use the cross pseudo-supervision algorithm and the reconstructed image dataset to train the cross hybrid model to obtain the target model;
[0086] Specifically, for the cross hybrid model SVTRCPS, the cross pseudo-supervision algorithm (CPS) is adopted to achieve the parallel inference of the two models SVTRv2 and RepSVTR:
[0087] Given the excellent performance of SVTRv2 and RepSVTR in the OCR field, this patent proposes a new cross hybrid model SVTRCPS based on them. This model abandons the original knowledge distillation structure and adopts the cross pseudo-supervision algorithm (CPS) to achieve the parallel inference of the two models. Compared with knowledge distillation, the relationship between the models in the SVTRCPS model is closer, and the data can be more effectively interchanged. Unlabeled data can be used, or labeled data can be combined with unlabeled data, which is suitable for scenarios where labeled data is scarce.
[0088] Construction of the SVTRCPS model:
[0089] To avoid getting a good overall score on the validation set but poor performance on the samples of the test set, the training will be divided into two branches. One branch conducts supervised training on the labeled data, and the other branch uses unlabeled data for semi-supervised training, aiming to enhance the model's ability to predict external samples. Figure 6The training process is described as follows: Among them, XL and XU+L represent the labeled input, unlabeled, and labeled input respectively. YL and YS are the label values (true label and soft pseudo-label respectively), while P represents the probability map calculated by the network. (→) represents the forward direction, ( / / on→) represents stopping the gradient, (-→) represents loss supervision, and (-·→) represents masked loss supervision. Starting from Figure 6 It can be learned that for the supervised branch, RepSVTR and SVTRv2 will output two predicted label vectors (the predicted probabilities of the most likely optical characters), and YL can supervise them; while in the semi-supervised branch, since there is no true label, after RepSVTR and SVTRv2 obtain their respective predicted label vectors, they will respectively obtain their unique soft pseudo-labels through the argmax operation, and use these two pseudo-labels as each other's supervision signals. (Use Y2 as the supervision of P1 and Y1 as the supervision of P2).
[0090] Construction of the SVTRCPS Loss Function
[0091] In most training samples, the number of background pixels exceeds the number of text pixels, and the recognition difficulty of some samples is relatively large, resulting in low accuracy of the generated pseudo-labels. To address these issues, SVTRCPS needs to design a loss function to improve the efficiency and performance of the model in recognizing difficult samples. Therefore, this patent incorporates the focal Tversky loss and the cross-entropy loss with online hard example mining loss (also called guided cross-entropy loss). Among them, the focal Tversky loss weights false negatives (FN) and false positives (FP) through the α and β terms. To obtain a high recall rate, more penalties are imposed on FN; the online hard example mining loss only considers the top k highest-loss pixels in the predicted mask, which helps the network not to be overly confident in empty pixels. We limit k to be equal to H×W÷16. Tl has no actual meaning and is just an intermediate quantity. FTL is the Tversky loss, and γ can be used to emphasize or weaken the importance of TP relative to FN and FP. For example, if γ>1, then even if the number of TPs is relatively large, as long as the number of FNs and FPs is large enough, the measure of the overlap degree will be significantly reduced.
[0092] The formula for the focal Tversky loss is as follows:
[0093] FTL=(1 - Tl) γ
[0094]
[0095] In fact, these two loss functions are parameterized variants of the Dice / Cross Entropy loss, which can be adjusted to force the network to pay more attention to the recall score while maintaining good accuracy, thus obtaining better overall results. By specifically focusing on those samples that are difficult for the classifier to correctly identify through the above two losses, the model is forced to learn more robust decision boundaries, thereby improving the ability to handle fuzzy samples. The model can focus more on those samples that are helpful for performance, thus accelerating the training process. In many practical applications, the ratio of positive and negative samples may be very imbalanced. Hard sample mining can help adjust this imbalance by selecting more hard negative samples to balance the training data, so that the classifier does not bias towards the majority class.
[0096] Finally, it is judged by the total supervision loss:
[0097]
[0098] Among them, n and n′ respectively represent the numbers of two types of samples in the dataset, and C represents the number of categories. The first half of the function L is the cost of labeled data, and the second half is the cost of unlabeled data. In the cost of unlabeled data, y′ is the pseudo-label of unlabeled data, which takes the maximum value of the prediction of RepSVTR for unlabeled data. β(t) gradually increases with training based on a Gaussian-type climbing function and is set to 0 during initial training. In addition, techniques such as consistency regularization and entropy minimization are used during training to ensure that the unlabeled data in semi-supervised learning is fully trained and prevent its performance from decreasing rather than increasing compared to fully supervised learning algorithms.
[0099] Lightweight model to optimize the model size and accelerate the inference process
[0100] Model pruning: To meet the requirements of lightweight edge computing models, model pruning is used to reduce the complexity of the neural network. The weight magnitude pruning of the unstructured pruning algorithm is adopted, that is, a threshold θ is set, and those weights close to zero are removed according to whether the absolute value |w| of the weight is less than θ. To promote weight sparsity, an L1 regularization term is added to the loss function.
[0101] w′=(1 - M) * w
[0102] L′ = L + λ * ∑|w′|
[0103] Where M = {0, if |w|≥θ; 1, if |w|<θ}, λ is the regularization coefficient, and L is the loss function.
[0104] Model conversion: The original model can only be used on the Windows system. Therefore, it is necessary to convert the model to ONNX or RKNN to be used on the industrial control computer. Use Paddle2ONNX to convert the PaddlePaddle model format to the ONNX model format. During the process, ensure that the input and output dimensions of the model are consistent with the PaddleOCR model;
[0105] Model quantization: To further reduce the model size and accelerate the inference process, especially on edge devices, quantization work is also required, that is, use onnx runtime or the optional RKNN to convert the weights and activations of the model from floating-point numbers to low-precision integers.
[0106] Step 5, collect the real-time display screen image of the mixer truck and input it into the target model. After obtaining the electric drive parameter recognition result, upload it to the cloud platform, and the cloud platform analyzes it and formulates a macro-control strategy.
[0107] In specific implementation, after the target model recognizes relevant sensor data, it uploads the data to the cloud platform, and the platform analyzes the data and formulates a macro-control strategy:
[0108] After the model outputs, it is necessary to test the accuracy and performance of the final model. Here, code needs to be written to support the inference of the model, including details such as loading the model, image inference, text box display, and text content output. The inspection effect is as Figure 7 shown;
[0109] Data interaction with the platform:
[0110] Data upload: After the OCR model of the industrial control computer recognizes the data, the output data is obtained, and key information such as current, voltage, and speed is selected from it. The data is uploaded to the platform in the format required by the platform through the mobile network. Considering that the network quality is also changing in real time during transportation, in the case of no network, the data can be temporarily stored locally and uploaded in batches once the connection is restored.
[0111] Macro-control: After the platform receives the data, it analyzes and manages it to formulate corresponding scheduling strategies or match existing scheduling strategies, and promptly feedbacks to the transportation personnel whether to continue driving or need to stop for inspection to ensure that the electric drive parameters are normal during the current transportation process. The overall process is as Figure 8 shown.
[0112] The electric drive parameter acquisition method based on the mixer truck platform provided in this embodiment improves the data quality by adopting an anti-shake algorithm, uses an excitation branch with an attention mechanism to accelerate the recognition speed, builds a cross-pseudo-supervision model with an advanced backbone, designs a loss function with clear division of labor to obtain the optimal training result, and combines vision recognition, deep learning, and data analysis technologies to achieve intelligent monitoring and management of mixer trucks. Through accurate data acquisition and label annotation, noise and errors are reduced, the accuracy of image processing is improved, and thus the real-time recognition ability of electric drive parameters is enhanced. At the same time, the cross-pseudo-supervision algorithm improves the generalization ability of the model and ensures the stability and reliability of the system under different environmental conditions. The intelligent analysis and macro-control strategy of the cloud platform make vehicle scheduling more efficient, and fault prediction and preventive maintenance improve operation safety, ultimately improving transportation efficiency and reducing costs.
[0113] It should be understood that each part of the present disclosure can be implemented by hardware, software, firmware, or a combination thereof.
[0114] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present disclosure should be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A method for collecting electric drive parameters based on a mixer truck platform, characterized in that: include: Step 1: During the preset operation time of the mixer truck, the display screen is captured by the camera and labeled to obtain the original data set; Step 2, performing anti-shake processing on the original data set to obtain a preprocessed data set; Step 3, design an excitation branch structure and input the preprocessed data set into it to obtain a reconstructed image data set; Step 4, using the cross pseudo-supervision algorithm and the reconstructed image data set to train the cross hybrid model to obtain the target model; Step 5: Collect the real-time display screen image of the mixer truck and input it into the target model. After obtaining the electric drive parameter identification result, upload it to the cloud platform, which analyzes it and formulates a macro-control strategy.
2. The method according to claim 1, characterized in that The step 1 specifically includes: Step 1.1, obtain the display screen image captured by the camera, and output the display screen image to the industrial computer in the form of a video stream; Step 1.2: The industrial computer uses the video processing library to read the video stream and extracts a frame from the video stream at a fixed time interval to form a sample data set. Step 1.3: annotate a preset proportion of the video frame data in the sampled data set with text labels according to the data type collected by the sensor, and leave the remaining part unlabeled to form the original data set.
3. The method according to claim 2, characterized in that The step 2 specifically includes: Step 2.1, use the optical flow learning algorithm to calculate a backward dense warp field for each frame of the original dataset; Step 2.2, using the RAFT algorithm to restore the missing pixels caused by distortion through optical flow estimation technology; In step 2.3, the distorted color frames are mixed in the image space to generate output stable frames through frame fusion technology to form a preprocessed data set.
4. The method according to claim 3, characterized in that The step 3 specifically includes: Step 3.1, designing an excitation branch structure, wherein the excitation branch structure includes an encoder, a decoder, a slot attention module, an attention loss, and a CTC loss, the encoder is used to extract features from the image, the decoder is used to predict the final text or background detection box and reconstruct the image, the slot attention module clusters each pixel into a text or background layer based on the semantic characteristics of the text, and the attention loss and CTC loss are used to locate the position of the target; In step 3.2, all images in the preprocessed data set are input into the excitation branch structure to obtain a reconstructed image data set with a detection frame.
5. The method according to claim 4, characterized in that The expression of the attention loss is Among them, α ij is the actual attention weight, is the predicted attention weight, N represents the number of elements in column T; The expression of the CTC loss is: Among them, p(T i |X i ) is a given input sequence X i When the prediction sequence T i probability.
6. The method according to claim 5, characterized in that The cross-hybrid model includes a teacher model SVTRv2 and a student model RepSVTR, and step 4 specifically includes: Step 4.1, divide the cross-hybrid model training into a supervised branch and a semi-supervised branch. For the supervised branch, the input requirement is a labeled image in the reconstructed image dataset. For the semi-supervised branch, the input requirement is a fixed ratio of unlabeled images and labeled images. Step 4.2: Based on the inputs of the supervised branch and the semi-supervised branch, the teacher model SVTRv2 and the student model RepSVTR perform inference in parallel and output their respective prediction results. The prediction results of the labeled images are constrained by the true labels. For the unlabeled images, their prediction results are used as the constraint labels of each other for cross-validation. Step 4.3, based on the prediction results, the parameters of the cross-mixing model are updated using the focal Tversky loss and the guided cross entropy loss; Step 4.4, after the parameters of the cross-hybrid model are updated, model pruning, model conversion and quantization are performed to obtain the target model.
7. The method according to claim 6, characterized in that The step 5 specifically includes: Step 5.1, collect the real-time display screen image of the mixer truck and input it into the target model, obtain the electric drive parameter recognition result, store the electric drive parameter recognition result locally and upload it to the cloud platform through the mobile network; Step 5.2: After receiving the electric drive parameter identification results, the cloud platform analyzes them, generates a scheduling strategy and feeds it back to the transportation personnel.
Citation Information
Cited By
Unmanned logistics vehicle driving decision-making method and system based on VLA architecture and distillation learning
CN120762420A