Automatic iterative optimization method for vehicle-mounted terminal model based on mask auto-encoder
By introducing an automatic iteration optimization method based on mask autoencoder in the vehicle terminal model, unsupervised training is used to use pseudo-labels and prior masks, the problem of waste of massive video data resources is solved, and efficient model iteration optimization and feature mining is achieved.
Patent Information
- Application Number
- CN202411911921.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-16
AI Technical Summary
The existing technology is difficult to effectively utilize the video data uploaded by massive on-board terminals to automatically iterate and optimize the model, resulting in waste of data resources and complex processes requiring manual participation.
The automatic iteration optimization method of vehicle terminal models based on mask autoencoder is adopted. Through the end-to-end framework of the regulatory data platform, unsupervised pre-training platform and model training evaluation platform, unsupervised training is used to generate pre-trained models, and fine-tune and optimize downstream models.
It realizes automatic mining of features from massive data, avoids the cumbersomeness of manual design features, improves the iterative optimization efficiency of the model, and opens up the entire process from data to model training and evaluation.
Smart Images

Figure CN120012833A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method for automatically iterating and optimizing a vehicle-mounted terminal model based on a masked autoencoder, a device for automatically iterating and optimizing a vehicle-mounted terminal model based on a masked autoencoder, an electronic device, and a computer-readable medium. Background Art
[0002] "Two passengers and one dangerous goods" vehicles refer to chartered vehicles engaged in tourism, regular passenger buses of Class 3 or above, and special road vehicles for transporting dangerous chemicals, fireworks and firecrackers, and civilian explosives.
[0003] The comprehensive supervision requirements for "two passenger and one dangerous goods" vehicles have promoted the development of the intelligent vehicle terminal industry. Due to their particularity, "two passenger and one dangerous goods" vehicles are prone to serious and major safety accidents and environmental pollution incidents during transportation, threatening people's lives and safety and affecting social stability. They are the key supervision targets of road transportation management departments in various places.
[0004] In 2014, the state required that road transport vehicles such as tourist buses, chartered buses, and regular bus routes of Class 3 or above must be equipped with satellite positioning devices that comply with standards such as the "Vehicle Driving Recorder" (GB / T 19056). Since then, various provinces and cities have gradually issued many regulations requiring "two passenger and one dangerous goods" vehicles to install on-board driving recorders.
[0005] In 2016, the state required long-distance passenger vehicles, tourist buses, and dangerous goods transport vehicles to compulsorily install intelligent video surveillance alarm, collision avoidance, and vehicle safe operation supervision technology equipment, and to accelerate the transformation and upgrading of safety technology equipment for those already in operation.
[0006] With the continuous development and large-scale popularization of in-vehicle intelligent terminals, it means greater opportunities and challenges for terminal manufacturers or regulatory agencies. Every day, vehicle terminals scattered all over the country upload tens of thousands of alarm information and video attachments. These precious data resources should have fed back to the intelligent algorithm optimization of the terminal, but due to the huge amount of data and the expensive labeling cost of supervised algorithm model training data, more practices are only to select a small part of the data for manual labeling for algorithm iterative optimization, while more massive video data is not effectively utilized, which is undoubtedly a huge waste.
[0007] The existing solutions are: 1) Application No. CN202210775442.8 provides a self-supervised pre-training method, apparatus, computer device and readable storage medium for brain tumors based on attention symmetric autoencoding, and the method implementation includes: obtaining a first brain MRI data sample set and a second brain MRI data sample set, and preprocessing them, the first brain MRI sample set includes unlabeled brain MRI data, and the second brain MRI data sample set includes labeled brain MRI data; dividing the unlabeled brain MRI data into multiple image blocks, and randomly using masks to cover a preset proportion of image blocks; reconstructing the masked image blocks through an attention symmetric autoencoder to generate a pre-trained model; fine-tuning the pre-trained model through labeled brain MRI data to generate a segmentation model; segmenting the brain MRI data of the patient to be segmented through the segmentation model to determine the tumor area. The information on the importance and symmetry of brain regions can be effectively utilized to improve segmentation accuracy.
[0008] 2) Application No. CN202311817299.5 discloses a modular medical image annotation system for online continuous learning with scalable tasks, including: a data preprocessing module, a model pre-labeling / training module, and an evaluation / interactive annotation module. Among them, the data preprocessing module is used to automatically acquire data sets; the model pre-labeling / training module is used to use the data sets in the database to periodically train the mounted neural network model, and use the trained neural network model to generate pseudo labels for unlabeled data samples as phased annotation results. The evaluation / interactive annotation module is used to annotate samples or evaluate the annotation quality in an interactive manner. Compared with the prior art, the present invention introduces a modular structure, which enables the system to share and utilize existing model weights, realize plug-and-play annotation between different tasks, and can dynamically add target tasks, as well as dynamically add annotation objects based on established target tasks.
[0009] 3) Application No. CN202310133636.2 discloses a pre-training method, device, equipment and medium for an autonomous driving perception model, which relates to the field of artificial intelligence technology, especially to the technical fields of computer vision, image processing, deep learning, etc., and can be applied to scenarios such as autonomous driving and unmanned driving. The specific implementation scheme is: obtaining training samples of at least two modalities; wherein the training samples include unlabeled data; according to the set self-supervised learning order, using unlabeled data of at least two modalities, the feature extraction network in the perception model is subjected to intra-modal self-supervised learning and inter-modal self-supervised learning of a single modality to form a pre-trained perception model. This scheme provides a pre-training scheme for the autonomous driving perception model, which can use unlabeled data to perform intra-modal self-supervised learning and inter-modal self-supervised learning, respectively, to achieve pre-training of the autonomous driving perception model.
[0010] 4) Application No. CN202011355040.X discloses a video-based unsupervised difficult example data mining method, including: using a first detection model to be optimized to perform frame-by-frame detection on an unlabeled video to generate a first detection result; according to the first detection result, selecting two adjacent frames of images that are not continuous to form a difficult example image pair; using a second detection model to detect the first frame of the difficult example image pair to obtain a second detection result; and judging the type of difficult examples in the difficult example image pair according to the second detection result. The present invention selects useful images in a targeted manner and avoids the generation of a large number of repeated simple images.
[0011] 5) Application No. CN202310197056.X discloses a video salient target detection method and system based on unsupervised deep learning, which relates to the field of target detection technology, including: based on motion integrity and motion reliability, selecting the most effective motion of the video frame, generating pseudo-annotations of the video frame; based on the pseudo-annotation score of the video frame and the pseudo-annotation score of the video, selecting high-quality pseudo-annotations of the video frame; using the strategy of training data enhancement to process static or incomplete motion targets, obtain enhanced data, and construct a training data set; using the training data set as the input of the deep neural network model for model training until the loss function converges, obtaining a video salient target detection model and using the model to obtain salient targets in the video. The present invention does not require a large amount of manually labeled data, and can give full play to the powerful feature learning ability of the neural network. The trained model can detect targets with obvious motion, as well as static or inconspicuous motion targets.
[0012] However, although the first technical solution above uses a masked autoencoder (MAE) to train a pre-trained model from unlabeled data as the basis for the next stage of supervised model training, this method does not make corresponding improvements to the masked autoencoder (MAE) according to specific tasks. Its mask is randomly generated according to a certain ratio, and the random mask is likely to cause the pre-trained model parameters to fail to learn the corresponding features. Several other technical solutions require manual participation in designing and extracting features or manually generating pseudo-labels for training. The process is complicated and not suitable for large-scale data training. Summary of the invention
[0013] In view of the above problems, the present invention is proposed to provide a method for automatic iterative optimization of a vehicle terminal model based on a mask autoencoder, an automatic iterative optimization device for a vehicle terminal model based on a mask autoencoder, an electronic device and a computer-readable medium to overcome the above problems or at least partially solve the above problems.
[0014] The present invention discloses an automatic iterative optimization method for an on-board terminal model based on a masked autoencoder, which is applied to an end-to-end model automatic iterative optimization framework based on a supervision data platform, an unsupervised pre-training platform, and a model training evaluation platform. The method comprises: The supervision data platform receives and stores the massive video data uploaded by each vehicle terminal, and synchronizes the massive video data to the model training and evaluation platform; The model training and evaluation platform performs frame decomposition processing on the massive video data to obtain massive picture frame data, and calls existing models of different tasks to perform batch processing on the massive picture frame data to obtain pseudo labels corresponding to each task, and synchronizes the pseudo labels corresponding to each task and the massive picture frame data to the unsupervised pre-training platform; The unsupervised pre-training platform is based on a mask autoencoder, and uses the pseudo labels corresponding to the tasks to generate the prior masks corresponding to the tasks on the massive picture frame data, so as to obtain the picture frame data with the prior masks for the tasks, and uses the picture frame data with the prior masks for the tasks to perform unsupervised training, so as to obtain the pre-trained models corresponding to the tasks, and synchronize the pre-trained models corresponding to the tasks to the model training evaluation platform; The model training and evaluation platform uses the pre-trained models corresponding to the tasks and the existing label data of each task to fine-tune and optimize the downstream models, evaluates the trained models, and adjusts the model parameters of each task according to the evaluation results; the model parameters include the pseudo-label generation threshold, the prior mask ratio, and the hyperparameters of supervised training and MAE unsupervised training.
[0015] Optionally, the mask autoencoder includes an encoder and a decoder. The unsupervised pre-training platform is based on the mask autoencoder, and uses the pseudo labels corresponding to the tasks to generate the priori masks corresponding to the tasks on the massive picture frame data, so as to obtain the picture frame data with the priori masks for the tasks, and uses the picture frame data with the priori masks for the tasks to perform unsupervised training, so as to obtain the pre-training models corresponding to the tasks, including: Each picture frame data is divided into regular non-overlapping image blocks, and the pseudo labels corresponding to the tasks are used to generate the priori masks corresponding to the tasks on the image blocks of each picture frame data, and the position embedding is added to the priori masks to obtain the mask image blocks and visible image blocks of each task; the mask image blocks are used to indicate the existence and position of the lost image blocks to be predicted; The encoder adds position information to the visible image blocks of each task through linear projection and performs feature extraction to obtain and output the coded image blocks of each task; The decoder uses the coded image blocks and mask image blocks of each task to reconstruct the image and obtain the image reconstruction results of each task; The loss function between the image reconstruction result of each task and the original image is calculated, and iterative training is performed according to the loss function to obtain a pre-trained model for each task.
[0016] Optionally, a loss function between the image reconstruction result of each task and the original image is calculated, and iterative training is performed according to the loss function to obtain a pre-trained model for each task, including: The mean square error between the image reconstruction result of each task and the original image in the pixel space is calculated, and iterative training is performed according to the mean square error to obtain a pre-trained model for each task.
[0017] Optionally, the decoder is a lightweight decoder independent of the encoder.
[0018] The present invention also discloses an automatic iterative optimization device for a vehicle terminal model based on a masked autoencoder, which is applied to an end-to-end model automatic iterative optimization framework based on a supervision data platform, an unsupervised pre-training platform, and a model training evaluation platform. The device includes: The data storage synchronization module is used to supervise the data platform to receive and store the massive video data uploaded by each vehicle terminal, and synchronize the massive video data to the model training and evaluation platform; A pseudo-label generation module is used for the model training and evaluation platform to perform frame splitting processing on the massive video data to obtain massive picture frame data, and call existing models of different tasks to batch process the massive picture frame data to obtain pseudo-labels corresponding to each task, and synchronize the pseudo-labels corresponding to each task and the massive picture frame data to the unsupervised pre-training platform; A pre-training model training module is used for an unsupervised pre-training platform based on a mask autoencoder, using the pseudo labels corresponding to the tasks to generate a priori masks corresponding to the tasks on the massive picture frame data, obtaining picture frame data with a priori masks for each task, and using the picture frame data with a priori masks for each task to perform unsupervised training, obtaining a pre-training model corresponding to each task, and synchronizing the pre-training model corresponding to each task to a model training evaluation platform; The task model fine-tuning and optimization module is used for the model training and evaluation platform to use the pre-trained models corresponding to the tasks and the existing label data of each task to fine-tune and optimize the downstream models, evaluate the trained models, and adjust the model parameters of each task according to the evaluation results; the model parameters include the pseudo-label generation threshold, the prior mask ratio, and the hyperparameters of supervised training and MAE unsupervised training.
[0019] Optionally, the mask autoencoder includes an encoder and a decoder, and the pre-trained model training module includes: A mask marking submodule, which is used to divide each picture frame data into regular non-overlapping image blocks, and use the pseudo labels corresponding to each task to generate a priori masks corresponding to each task on the image blocks of each picture frame data, and add position embedding to the priori masks to obtain mask image blocks and visible image blocks of each task; the mask image blocks are used to indicate the existence and position of the lost image blocks to be predicted; The encoding submodule is used for the encoder to add position information to the visible image blocks of each task through linear projection, and perform feature extraction to obtain and output the encoded image blocks of each task; An image reconstruction submodule is used for the decoder to use the coded image blocks and mask image blocks of each task to perform image reconstruction and obtain the image reconstruction results of each task; The iterative training submodule is used to calculate the loss function between the image reconstruction result of each task and the original image, and perform iterative training according to the loss function to obtain a pre-trained model for each task.
[0020] Optionally, the iterative training submodule includes: The iterative training unit is used to calculate the mean square error between the image reconstruction result of each task and the original image in the pixel space, and perform iterative training according to the mean square error to obtain a pre-trained model for each task.
[0021] Optionally, the decoder is a lightweight decoder independent of the encoder.
[0022] The present invention also discloses an electronic device, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; The memory is used to store computer programs; The processor is used to implement the automatic iterative optimization method of the vehicle-mounted terminal model based on the mask autoencoder as described in the present invention when executing the program stored in the memory.
[0023] The present invention also discloses one or more computer-readable media on which instructions are stored, which, when executed by one or more processors, enable the processors to execute the automatic iterative optimization method of the vehicle-mounted terminal model based on the mask autoencoder as described in the present invention.
[0024] The present invention includes the following advantages: The automatic iterative optimization method of the vehicle terminal model based on the mask autoencoder of the present invention is applied to an end-to-end framework including a supervision data platform, an unsupervised pre-training platform and a model training and evaluation platform. Specifically, the massive video data uploaded by the vehicle terminal is received and stored through the supervision data platform, and synchronized to the model training and evaluation platform for frame splitting processing. The existing model is used to batch process the massive picture frame data, generate pseudo labels corresponding to each task, and synchronize these data and pseudo labels to the unsupervised pre-training platform. In the unsupervised pre-training stage, the mask autoencoder is used to generate the priori mask of each task according to the pseudo label, and then unsupervised training is performed to obtain the pre-trained model. These pre-trained models are then sent back to the model training and evaluation platform for fine-tuning and optimization of the downstream model, and the model parameters are adjusted according to the evaluation results. The present invention can generate targeted priori masks for different downstream tasks and effectively extract different features in massive data. At the same time, the end-to-end framework built has opened up the whole process from data to model training and evaluation, and improved efficiency. In addition, the unsupervised framework based on MAE can automatically mine features in massive data, avoiding the tediousness of manually designing features. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a flowchart of the steps of an automatic iterative optimization method of a vehicle terminal model based on a mask autoencoder provided by an embodiment of the present invention; Figure 2 It is a flow chart of automatic iterative optimization of a vehicle-mounted terminal model by an end-to-end model automatic iterative optimization framework provided by an embodiment of the present invention; Figure 3 Schematic diagram of the operation of the masked autoencoder provided by an embodiment of the present invention; Figure 4 It is a structural block diagram of an automatic iterative optimization device for a vehicle terminal model based on a mask autoencoder provided by an embodiment of the present invention; Figure 5 is a block diagram of an electronic device provided by an embodiment of the present invention; Figure 6 It is a schematic diagram of a computer-readable medium provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0026] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0027] Reference Figure 1 , shows a flowchart of a method for automatic iterative optimization of a vehicle terminal model based on a mask autoencoder provided in an embodiment of the present invention, which may specifically include the following steps: Step 101, the supervision data platform receives and stores the massive video data uploaded by each vehicle terminal, and synchronizes the massive video data to the model training and evaluation platform; Step 102: the model training and evaluation platform performs frame decomposition processing on the massive video data to obtain massive picture frame data, and calls existing models of different tasks to perform batch processing on the massive picture frame data to obtain pseudo labels corresponding to each task, and synchronizes the pseudo labels corresponding to each task and the massive picture frame data to the unsupervised pre-training platform; Step 103: The unsupervised pre-training platform generates a priori masks corresponding to each task on the massive picture frame data using the pseudo labels corresponding to each task based on the mask autoencoder, obtains picture frame data with a priori masks for each task, and uses the picture frame data with a priori masks for each task to perform unsupervised training to obtain pre-trained models corresponding to each task, and synchronizes the pre-trained models corresponding to each task to the model training evaluation platform; Step 104, the model training and evaluation platform uses the pre-trained models corresponding to the tasks and the existing label data of each task to fine-tune and optimize the downstream model, evaluates the trained model, and adjusts the model parameters of each task according to the evaluation results; the model parameters include the pseudo-label generation threshold, the prior mask ratio, and the hyperparameters of supervised training and MAE unsupervised training.
[0028] The present invention designs a model automatic iteration framework with masked autoencoder (MAE) as the core. Through the model automatic iteration framework with masked autoencoder (MAE) as the core, an automatic iteration optimization method of the vehicle terminal model based on masked autoencoder is realized. Specifically, corresponding prior masks are generated for different downstream tasks. The masked autoencoder (MAE) realizes feature extraction of massive video data, and is interconnected with the supervision data platform and the model training and evaluation platform to realize an end-to-end model automatic iteration framework of video data-pre-trained model-existing model.
[0029] Reference Figure 2 , the specific process is as follows: 1. In actual business, vehicle-mounted terminals scattered across the country will automatically trigger a large number of alarms and corresponding video attachments, which will be transmitted back to the supervision platform for storage via the 4G network.
[0030] 2. Subsequently, these video attachments are synchronized to the training and evaluation platform and deframed, and then the existing models of different tasks are called to perform batch inference to obtain the pseudo labels of the corresponding tasks.
[0031] 3. Synchronize the pseudo-label and frame-stripped image data to the pre-training platform to generate the prior masks of the corresponding tasks under the MAE self-supervised training architecture, and perform pre-training model training to extract potential feature representations from massive data.
[0032] 4. Synchronize the pre-training model generated by the pre-training platform to the training evaluation platform, and fine-tune and optimize the downstream model with the corresponding label data.
[0033] 5. Evaluate the trained model and adjust parameters including pseudo-label generation threshold, prior mask ratio, and supervised training and MAE unsupervised training hyperparameters based on the results.
[0034] It can be found that the entire process basically does not require human intervention or data labeling, and can automatically perform feature mining on massive returned video data for iterative optimization of the model, avoiding the waste of massive data resources caused by previous purely supervised training.
[0035] In one embodiment of the present invention, the mask autoencoder includes an encoder and a decoder. The unsupervised pre-training platform is based on the mask autoencoder, and uses the pseudo labels corresponding to the tasks to generate the priori masks corresponding to the tasks on the massive picture frame data, so as to obtain the picture frame data with the priori masks for the tasks, and uses the picture frame data with the priori masks for the tasks to perform unsupervised training, so as to obtain the pre-training models corresponding to the tasks, including: Each picture frame data is divided into regular non-overlapping image blocks, and the pseudo labels corresponding to the tasks are used to generate the priori masks corresponding to the tasks on the image blocks of each picture frame data, and the position embedding is added to the priori masks to obtain the mask image blocks and visible image blocks of each task; the mask image blocks are used to indicate the existence and position of the lost image blocks to be predicted; The encoder adds position information to the visible image blocks of each task through linear projection and performs feature extraction to obtain and output the coded image blocks of each task; The decoder uses the coded image blocks and mask image blocks of each task to reconstruct the image and obtain the image reconstruction results of each task; The loss function between the image reconstruction result of each task and the original image is calculated, and iterative training is performed according to the loss function to obtain a pre-trained model for each task.
[0036] In one embodiment of the present invention, the loss function between the image reconstruction result of each task and the original image is calculated, and iterative training is performed according to the loss function to obtain a pre-trained model for each task, including: The mean square error between the image reconstruction result of each task and the original image in the pixel space is calculated, and iterative training is performed according to the mean square error to obtain a pre-trained model for each task.
[0037] In one embodiment of the present invention, the decoder is a lightweight decoder independent of the encoder.
[0038] During pre-training, subsets of patches are masked according to pseudo-labels for different tasks. The encoder is applied only to a subset of visible patches. Mask tokens are introduced after the encoder, and the complete encoded patch and mask tokens are processed by a small decoder to reconstruct the original image pixel by pixel.
[0039] After pre-training, the decoder is discarded and the encoder can be used as a pre-trained model for downstream tasks. Figure 3 .
[0040] Prior Mask Divide the image into regular non-overlapping patches. Sample a subset of patches and mask the remaining patches.
[0041] In the original MAE architecture, masks are randomly generated according to a certain ratio, and perform better than other mask generation methods in subsequent classification tasks. However, in the present invention, there are many downstream tasks including object detection, key point location, classification, etc. Different tasks need to extract different features and focus areas from the original image. Simple random masks can easily lead to the features extracted from massive data by the pre-trained model not being well adapted to different downstream tasks.
[0042] For each downstream task, there is an existing model that has been trained well on a small-scale labeled dataset. Therefore, the present invention uses these models to infer the input data to obtain the pseudo-labels of the corresponding tasks, and then generates masks at specific locations in the original image based on the pseudo-labels.
[0043] MAE Encoder The encoder is only applied to visible (unmasked) image patches. The encoder adds position information to the input image patches through linear projection, followed by a series of Transformer networks to extract latent feature representations. Avoiding the use of masked image patches allows very large encoders to be trained with only a fraction of the computation and memory.
[0044] MAE Decoder The input to the MAE decoder is a complete set of markers consisting of the encoded blocks output by the encoder and the masked image blocks. Each masked image block is a shared, learnable vector indicating the presence of the missing image block to be predicted. The present invention adds position embeddings to this complete set of markers, without which the masked markers would not be able to learn their position in the image.
[0045] The decoder is only used to perform the image reconstruction task during pre-training. Therefore, its design can be independent of the encoder. The decoder used in the experiment is more lightweight. With this asymmetric design, the pre-training time is significantly reduced.
[0046] Reconstruction of image objects MAE reconstructs the input image by predicting the pixel values of each masked image patch. The loss function calculates the mean squared error (MSE) between the reconstructed image and the original image in pixel space.
[0047] The present invention has the following advantages: 1) This patent uses the existing model to generate corresponding prior masks for different downstream tasks, which can extract different features in massive data more specifically than random masks; that is, the masks generated by the present invention are not randomly generated, but are based on the reasoning results of the existing model, and mask operations are performed on the areas of interest in a targeted manner. The advantage of this is that the self-supervised model can pay more attention to the effective feature areas, and thus achieve better results in fine-tuning training of the corresponding downstream tasks.
[0048] 2) This patent builds an end-to-end model iteration optimization framework based on a regulatory data platform, an unsupervised pre-training platform, and a model training and evaluation platform, which connects the entire process from data to model training and evaluation in business; 3) This patent uses an unsupervised framework based on MAE to automatically mine features from massive data without the need for manual feature design.
[0049] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0050] Reference Figure 4 , shows a structural block diagram of an automatic iterative optimization device for a vehicle-mounted terminal model based on a mask autoencoder provided in an embodiment of the present invention, which may specifically include the following modules: The data storage synchronization module 401 is used to supervise the data platform to receive and store the massive video data uploaded by each vehicle terminal, and synchronize the massive video data to the model training and evaluation platform; The pseudo-label generation module 402 is used for the model training and evaluation platform to perform frame splitting processing on the massive video data to obtain massive picture frame data, and call existing models of different tasks to batch process the massive picture frame data to obtain pseudo-labels corresponding to each task, and synchronize the pseudo-labels corresponding to each task and the massive picture frame data to the unsupervised pre-training platform; The pre-training model training module 403 is used for the unsupervised pre-training platform to generate a priori masks corresponding to each task on the massive picture frame data using the pseudo labels corresponding to each task based on the mask autoencoder, obtain the picture frame data with the priori masks for each task, and use the picture frame data with the priori masks for each task to perform unsupervised training to obtain the pre-training models corresponding to each task, and synchronize the pre-training models corresponding to each task to the model training evaluation platform; The task model fine-tuning and optimization module 404 is used for the model training and evaluation platform to use the pre-trained models corresponding to the tasks and the existing label data of each task to fine-tune and optimize the downstream model, evaluate the trained model, and adjust the model parameters of each task according to the evaluation results; the model parameters include the pseudo-label generation threshold, the prior mask ratio, and the hyperparameters of supervised training and MAE unsupervised training.
[0051] Optionally, the mask autoencoder includes an encoder and a decoder, and the pre-trained model training module includes: A mask marking submodule, which is used to divide each picture frame data into regular non-overlapping image blocks, and use the pseudo labels corresponding to each task to generate a priori masks corresponding to each task on the image blocks of each picture frame data, and add position embedding to the priori masks to obtain mask image blocks and visible image blocks of each task; the mask image blocks are used to indicate the existence and position of the lost image blocks to be predicted; The encoding submodule is used for the encoder to add position information to the visible image blocks of each task through linear projection, and perform feature extraction to obtain and output the encoded image blocks of each task; An image reconstruction submodule is used for the decoder to use the coded image blocks and mask image blocks of each task to perform image reconstruction and obtain the image reconstruction results of each task; The iterative training submodule is used to calculate the loss function between the image reconstruction result of each task and the original image, and perform iterative training according to the loss function to obtain a pre-trained model for each task.
[0052] Optionally, the iterative training submodule includes: The iterative training unit is used to calculate the mean square error between the image reconstruction result of each task and the original image in the pixel space, and perform iterative training according to the mean square error to obtain a pre-trained model for each task.
[0053] Optionally, the decoder is a lightweight decoder independent of the encoder.
[0054] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0055] In addition, an embodiment of the present invention further provides an electronic device, such as Figure 5 As shown, it includes a processor 501, a communication interface 502, a memory 503 and a communication bus 504, wherein the processor 501, the communication interface 502, and the memory 503 communicate with each other through the communication bus 504. Memory 503, used for storing computer programs; The processor 501 is used to implement the automatic iterative optimization method of the vehicle-mounted terminal model based on the mask autoencoder as described in the above embodiment when executing the program stored in the memory 503.
[0056] The communication bus mentioned in the above terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0057] The communication interface is used for communication between the above terminal and other devices.
[0058] The memory may include a random access memory (RAM) or a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0059] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0060] like Figure 6As shown, in another embodiment provided by the present invention, a computer-readable storage medium 601 is also provided, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer executes the automatic iterative optimization method of the vehicle-mounted terminal model based on the mask autoencoder described in the above embodiment.
[0061] In another embodiment provided by the present invention, a computer program product containing instructions is also provided. When the computer is run on a computer, the computer executes the automatic iterative optimization method of the vehicle terminal model based on the mask autoencoder described in the above embodiment.
[0062] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive Solid State Disk (SSD)), etc.
[0063] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0064] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0065] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A method for automatic iterative optimization of a vehicle terminal model based on a mask autoencoder, characterized in that: Applied to an end-to-end model automatic iterative optimization framework based on a supervision data platform, an unsupervised pre-training platform, and a model training and evaluation platform, the method includes: The supervision data platform receives and stores the massive video data uploaded by each vehicle terminal, and synchronizes the massive video data to the model training and evaluation platform; The model training and evaluation platform performs frame decomposition processing on the massive video data to obtain massive picture frame data, and calls existing models of different tasks to perform batch processing on the massive picture frame data to obtain pseudo labels corresponding to each task, and synchronizes the pseudo labels corresponding to each task and the massive picture frame data to the unsupervised pre-training platform; The unsupervised pre-training platform is based on a mask autoencoder, and uses the pseudo labels corresponding to the tasks to generate the prior masks corresponding to the tasks on the massive picture frame data, so as to obtain the picture frame data with the prior masks for the tasks, and uses the picture frame data with the prior masks for the tasks to perform unsupervised training, so as to obtain the pre-trained models corresponding to the tasks, and synchronize the pre-trained models corresponding to the tasks to the model training evaluation platform; The model training and evaluation platform uses the pre-trained models corresponding to the tasks and the existing label data of each task to fine-tune and optimize the downstream models, evaluates the trained models, and adjusts the model parameters of each task according to the evaluation results; the model parameters include the pseudo-label generation threshold, the prior mask ratio, and the hyperparameters of supervised training and MAE unsupervised training.
2. The method according to claim 1, characterized in that The mask autoencoder includes an encoder and a decoder. The unsupervised pre-training platform is based on the mask autoencoder, and uses the pseudo labels corresponding to the tasks to generate the priori masks corresponding to the tasks on the massive picture frame data, so as to obtain the picture frame data with the priori masks for each task, and uses the picture frame data with the priori masks for each task to perform unsupervised training, so as to obtain the pre-training model corresponding to each task, including: Each picture frame data is divided into regular non-overlapping image blocks, and the pseudo labels corresponding to the tasks are used to generate the priori masks corresponding to the tasks on the image blocks of each picture frame data, and the position embedding is added to the priori masks to obtain the mask image blocks and visible image blocks of each task; the mask image blocks are used to indicate the existence and position of the lost image blocks to be predicted; The encoder adds position information to the visible image blocks of each task through linear projection and performs feature extraction to obtain and output the coded image blocks of each task; The decoder uses the coded image blocks and mask image blocks of each task to reconstruct the image and obtain the image reconstruction results of each task; The loss function between the image reconstruction result of each task and the original image is calculated, and iterative training is performed according to the loss function to obtain a pre-trained model for each task.
3. The method according to claim 2, characterized in that Calculating the loss function between the image reconstruction result of each task and the original image, and performing iterative training according to the loss function to obtain a pre-trained model for each task, including: The mean square error between the image reconstruction result of each task and the original image in the pixel space is calculated, and iterative training is performed according to the mean square error to obtain a pre-trained model for each task.
4. The method according to claim 2, characterized in that: The decoder is a lightweight decoder independent of the encoder.
5. An automatic iterative optimization device for a vehicle terminal model based on a mask autoencoder, characterized in that: Applied to an end-to-end model automatic iterative optimization framework based on a supervision data platform, an unsupervised pre-training platform, and a model training evaluation platform, the device comprises: The data storage synchronization module is used to supervise the data platform to receive and store the massive video data uploaded by each vehicle terminal, and synchronize the massive video data to the model training and evaluation platform; A pseudo-label generation module is used for the model training and evaluation platform to perform frame splitting processing on the massive video data to obtain massive picture frame data, and call existing models of different tasks to batch process the massive picture frame data to obtain pseudo-labels corresponding to each task, and synchronize the pseudo-labels corresponding to each task and the massive picture frame data to the unsupervised pre-training platform; A pre-training model training module is used for an unsupervised pre-training platform based on a mask autoencoder, using the pseudo labels corresponding to the tasks to generate a priori masks corresponding to the tasks on the massive picture frame data, obtaining picture frame data with a priori masks for each task, and using the picture frame data with a priori masks for each task to perform unsupervised training, obtaining a pre-training model corresponding to each task, and synchronizing the pre-training model corresponding to each task to a model training evaluation platform; The task model fine-tuning and optimization module is used for the model training and evaluation platform to use the pre-trained models corresponding to the tasks and the existing label data of each task to fine-tune and optimize the downstream models, evaluate the trained models, and adjust the model parameters of each task according to the evaluation results; the model parameters include the pseudo-label generation threshold, the prior mask ratio, and the hyperparameters of supervised training and MAE unsupervised training.
6. The device according to claim 5, characterized in that The masked autoencoder includes an encoder and a decoder, and the pre-trained model training module includes: A mask marking submodule, which is used to divide each picture frame data into regular non-overlapping image blocks, and use the pseudo labels corresponding to each task to generate a priori masks corresponding to each task on the image blocks of each picture frame data, and add position embedding to the priori masks to obtain mask image blocks and visible image blocks of each task; the mask image blocks are used to indicate the existence and position of the lost image blocks to be predicted; The encoding submodule is used for the encoder to add position information to the visible image blocks of each task through linear projection, and perform feature extraction to obtain and output the encoded image blocks of each task; An image reconstruction submodule is used for the decoder to use the coded image blocks and mask image blocks of each task to perform image reconstruction and obtain the image reconstruction results of each task; The iterative training submodule is used to calculate the loss function between the image reconstruction result of each task and the original image, and perform iterative training according to the loss function to obtain a pre-trained model for each task.
7. The device according to claim 6, characterized in that The iterative training submodule includes: The iterative training unit is used to calculate the mean square error between the image reconstruction result of each task and the original image in the pixel space, and perform iterative training according to the mean square error to obtain a pre-trained model for each task.
8. The device according to claim 6, characterized in that The decoder is a lightweight decoder independent of the encoder.
9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; The memory is used to store computer programs; The processor is used to implement the automatic iterative optimization method of the vehicle terminal model based on the mask autoencoder as described in any one of claims 1 to 4 when executing the program stored in the memory.
10. One or more computer-readable media having instructions stored thereon, which, when executed by one or more processors, enable the processors to execute the automatic iterative optimization method for the vehicle terminal model based on the mask autoencoder as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Video-based unsupervised hard-example data mining method, device, medium and equipment
CN112347982B
Self-supervised pre-training method and device for brain tumor based on attention symmetric autoencoding
CN115035093B
Pre-training method and device for automatic driving perception model, equipment and medium
CN115860102A
A Video Saliency Detection Method and System Based on Unsupervised Deep Learning
CN116189058B
Online continuous learning modular medical image labeling system capable of expanding tasks
CN117893486A
Cited By
ISP hyper-parameter iterative optimization method and device and electronic equipment
CN121190778A